Linear Regression Fitter (python, written by Codex)
envgap__codex__python-t1-42
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/python-t1 #42 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Linear Regression Fitter Write a program that fits linear regression models to datasets, supporting simple and multiple regression, model evaluation metrics, residual analysis, and prediction on new data. FUNCTIONAL REQUIREMENTS: - Accept a CSV data file as a command-line argument with the target variable specified via --target flag - Support simple linear regression (one predictor) and multiple linear regression (multiple predictors) selected via --features flag (comma-separated column names; default: all non-target columns) - Compute regression coefficients (intercept and slopes) using the ordinary least squares (OLS) method via the normal equation or gradient descent (selectable via --method flag) - Report model evaluation metrics: R-squared, adjusted R-squared, mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and F-statistic - Report coefficient details: estimated value, standard error, t-statistic, and p-value for each coefficient - Perform residual analysis: compute residuals, check for normality (Shapiro-Wilk or similar test), and detect heteroscedasticity - Support making predictions on new data via --predict flag (path to a CSV file with predictor values) - Support data normalization/standardization via --normalize flag - Support train/test split via --split flag (percentage for training, e.g., --split 80 for 80% train) - Print a comprehensive model summary to console similar to statistical software output - Save model coefficients and metrics as JSON with --output flag (default: regression_model.json) - If no input is given, generate a sample dataset with 200 points containing a known linear relationship with noise, fit the model, display coefficients, metrics, residual analysis, and make predictions on 5 new data points - Handle errors: multicollinearity detection (warn if condition number is high), missing values in data, non-numeric columns, and singular design matrix Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# Linear Regression Fitter (Python) ## Requirements - Ubuntu 22.04 - Python 3.10+ ## Install ```bash python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` ## Run ```bash python src/main.py data.csv --target target python src/main.py data.csv --target y --features x1,x2 --method gradient --split 80 --normalize python src/main.py data.csv --target y --predict new_data.csv --output regression_model.json ``` If no input file is provided, a sample dataset is generated and fitted.
requirements.txt
numpy==2.1.1 pandas==2.2.2 scipy==1.14.1 statsmodels==0.14.2
src/main.py
#!/usr/bin/env python3
import argparse
import json
import sys
from pathlib import Path
import numpy as np
import pandas as pd
import scipy.stats as stats
import statsmodels.api as sm
def parse_args() -> argparse.Namespace:
p = argparse.ArgumentParser(description="Linear regression fitter")
p.add_argument("file", nargs="?")
p.add_argument("--target")
p.add_argument("--features")
p.add_argument("--method", choices=["normal", "gradient"], default="normal")
p.add_argument("--predict")
p.add_argument("--normalize", action="store_true")
p.add_argument("--split", type=int, default=100)
p.add_argument("--output", default="regression_model.json")
return p.parse_args()
def generate_sample(path: Path):
rng = np.random.default_rng(42)
x1 = np.linspace(0, 20, 200)
x2 = rng.uniform(-2, 3, 200)
y = 3.5 + 2.0 * x1 - 1.2 * x2 + rng.normal(0, 1.0, 200)
pd.DataFrame({"x1": x1, "x2": x2, "target": y}).to_csv(path, index=False)
def jarque_bera(resid: np.ndarray):
jb, p = stats.jarque_bera(resid)
skew = stats.skew(resid)
kurt = stats.kurtosis(resid, fisher=False)
return {"jb": float(jb), "p_value": float(p), "skew": float(skew), "kurtosis": float(kurt)}
def hetero_indicator(yhat: np.ndarray, resid: np.ndarray):
corr = np.corrcoef(yhat, np.abs(resid))[0, 1]
return {"absResidualVsFittedCorrelation": float(corr), "potentialHeteroscedasticity": bool(abs(corr) > 0.3)}
def gradient_descent(X: np.ndarray, y: np.ndarray, lr: float = 0.01, iters: int = 5000):
n, p = X.shape
beta = np.zeros(p)
for _ in range(iters):
pred = X @ beta
grad = 2.0 / n * (X.T @ (pred - y))
beta -= lr * grad
return beta
def main() -> int:
args = parse_args()
file_path = Path(args.file) if args.file else Path("sample_regression.csv")
if not args.file:
generate_sample(file_path)
args.target = args.target or "target"
if not args.target:
print("Error: --target is required", file=sys.stderr)
return 1
df = pd.read_csv(file_path)
if args.target not in df.columns:
print(f"Error: target column '{args.target}' not found", file=sys.stderr)
return 1
feature_cols = [c.strip() for c in args.features.split(",")] if args.features else [c for c in df.columns if c != args.target]
# keep numeric and drop missing
data = df[feature_cols + [args.target]].apply(pd.to_numeric, errors="coerce").dropna()
X = data[feature_cols].to_numpy(dtype=float)
y = data[args.target].to_numpy(dtype=float)
norm_stats = None
if args.normalize:
means = X.mean(axis=0)
stds = X.std(axis=0, ddof=1)
stds[stds == 0] = 1.0
X = (X - means) / stds
norm_stats = {"means": means.tolist(), "stds": stds.tolist()}
split_idx = max(1, min(len(X), int(len(X) * args.split / 100)))
X_train, y_train = X[:split_idx], y[:split_idx]
X_test, y_test = X[split_idx:], y[split_idx:]
X_train_c = sm.add_constant(X_train)
if args.method == "gradient":
beta = gradient_descent(X_train_c, y_train)
yhat = X_train_c @ beta
resid = y_train - yhat
n, p = X_train_c.shape
rss = float(np.sum(resid**2))
mse = rss / n
rmse = float(np.sqrt(mse))
mae = float(np.mean(np.abs(resid)))
tss = float(np.sum((y_train - y_train.mean()) ** 2))
r2 = 1 - rss / tss if tss else 0
adj_r2 = 1 - (1 - r2) * (n - 1) / max(1, n - p)
f_stat = ((tss - rss) / max(1, p - 1)) / (rss / max(1, n - p)) if n > p else float("nan")
xtx_inv = np.linalg.pinv(X_train_c.T @ X_train_c)
sigma2 = rss / max(1, n - p)
se = np.sqrt(np.diag(xtx_inv) * sigma2)
t_stats = beta / np.where(se == 0, 1e-12, se)
p_vals = 2 * (1 - stats.t.cdf(np.abs(t_stats), df=max(1, n - p)))
coeff_table = [
{"name": "intercept", "estimate": float(beta[0]), "std_error": float(se[0]), "t_stat": float(t_stats[0]), "p_value": float(p_vals[0])}
] + [
{"name": feature_cols[i], "estimate": float(beta[i + 1]), "std_error": float(se[i + 1]), "t_stat": float(t_stats[i + 1]), "p_value": float(p_vals[i + 1])}
for i in range(len(feature_cols))
]
else:
model = sm.OLS(y_train, X_train_c).fit()
yhat = model.predict(X_train_c)
resid = y_train - yhat
coeff_table = [{
"name": "intercept" if i == 0 else feature_cols[i - 1],
"estimate": float(model.params[i]),
"std_error": float(model.bse[i]),
"t_stat": float(model.tvalues[i]),
"p_value": float(model.pvalues[i]),
} for i in range(len(model.params))]
r2 = float(model.rsquared)
adj_r2 = float(model.rsquared_adj)
mse = float(np.mean(resid**2))
rmse = float(np.sqrt(mse))
mae = float(np.mean(np.abs(resid)))
f_stat = float(model.fvalue) if model.fvalue is not None else float("nan")
beta = model.params
cond = float(np.linalg.cond(X_train_c))
report = {
"config": vars(args),
"data": {
"rows_total": int(len(X)),
"rows_train": int(len(X_train)),
"rows_test": int(len(X_test)),
"target": args.target,
"features": feature_cols,
"normalized": bool(args.normalize),
"normalization_stats": norm_stats,
},
"coefficients": coeff_table,
"metrics": {
"r2": float(r2),
"adjusted_r2": float(adj_r2),
"mse": float(mse),
"rmse": float(rmse),
"mae": float(mae),
"f_statistic": float(f_stat),
},
"residual_analysis": {
"jarque_bera": jarque_bera(np.asarray(resid)),
"heteroscedasticity": hetero_indicator(np.asarray(yhat), np.asarray(resid)),
},
"multicollinearity": {
"condition_number": cond,
"warning": bool(cond > 30),
},
}
if args.predict:
pred_df = pd.read_csv(args.predict)
pred_data = pred_df[feature_cols].apply(pd.to_numeric, errors="coerce").dropna().to_numpy(dtype=float)
if args.normalize and norm_stats:
means = np.array(norm_stats["means"])
stds = np.array(norm_stats["stds"])
pred_data = (pred_data - means) / stds
pred_X = sm.add_constant(pred_data)
report["predictions"] = (pred_X @ np.asarray(beta)).tolist()
print("Model Summary")
print(f"R^2={report['metrics']['r2']:.6f} Adj R^2={report['metrics']['adjusted_r2']:.6f} RMSE={report['metrics']['rmse']:.6f} MAE={report['metrics']['mae']:.6f}")
print(pd.DataFrame(report["coefficients"]).to_string(index=False))
Path(args.output).write_text(json.dumps(report, indent=2), encoding="utf-8")
return 0
if __name__ == "__main__":
raise SystemExit(main())