← All tasks
pythoncodex/python-t1 #42Not a task: already works

Linear Regression Fitter (python, written by Codex)

envgap__codex__python-t1-42

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/python-t1 #42 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Linear Regression Fitter

Write a program that fits linear regression models to datasets, supporting simple and multiple regression, model evaluation metrics, residual analysis, and prediction on new data.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV data file as a command-line argument with the target variable specified via --target flag
- Support simple linear regression (one predictor) and multiple linear regression (multiple predictors) selected via --features flag (comma-separated column names; default: all non-target columns)
- Compute regression coefficients (intercept and slopes) using the ordinary least squares (OLS) method via the normal equation or gradient descent (selectable via --method flag)
- Report model evaluation metrics: R-squared, adjusted R-squared, mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and F-statistic
- Report coefficient details: estimated value, standard error, t-statistic, and p-value for each coefficient
- Perform residual analysis: compute residuals, check for normality (Shapiro-Wilk or similar test), and detect heteroscedasticity
- Support making predictions on new data via --predict flag (path to a CSV file with predictor values)
- Support data normalization/standardization via --normalize flag
- Support train/test split via --split flag (percentage for training, e.g., --split 80 for 80% train)
- Print a comprehensive model summary to console similar to statistical software output
- Save model coefficients and metrics as JSON with --output flag (default: regression_model.json)
- If no input is given, generate a sample dataset with 200 points containing a known linear relationship with noise, fit the model, display coefficients, metrics, residual analysis, and make predictions on 5 new data points
- Handle errors: multicollinearity detection (warn if condition number is high), missing values in data, non-numeric columns, and singular design matrix

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

README.md
# Linear Regression Fitter (Python)

## Requirements
- Ubuntu 22.04
- Python 3.10+

## Install
```bash
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

## Run
```bash
python src/main.py data.csv --target target
python src/main.py data.csv --target y --features x1,x2 --method gradient --split 80 --normalize
python src/main.py data.csv --target y --predict new_data.csv --output regression_model.json
```

If no input file is provided, a sample dataset is generated and fitted.
requirements.txt
numpy==2.1.1
pandas==2.2.2
scipy==1.14.1
statsmodels==0.14.2
src/main.py
#!/usr/bin/env python3
import argparse
import json
import sys
from pathlib import Path

import numpy as np
import pandas as pd
import scipy.stats as stats
import statsmodels.api as sm


def parse_args() -> argparse.Namespace:
    p = argparse.ArgumentParser(description="Linear regression fitter")
    p.add_argument("file", nargs="?")
    p.add_argument("--target")
    p.add_argument("--features")
    p.add_argument("--method", choices=["normal", "gradient"], default="normal")
    p.add_argument("--predict")
    p.add_argument("--normalize", action="store_true")
    p.add_argument("--split", type=int, default=100)
    p.add_argument("--output", default="regression_model.json")
    return p.parse_args()


def generate_sample(path: Path):
    rng = np.random.default_rng(42)
    x1 = np.linspace(0, 20, 200)
    x2 = rng.uniform(-2, 3, 200)
    y = 3.5 + 2.0 * x1 - 1.2 * x2 + rng.normal(0, 1.0, 200)
    pd.DataFrame({"x1": x1, "x2": x2, "target": y}).to_csv(path, index=False)


def jarque_bera(resid: np.ndarray):
    jb, p = stats.jarque_bera(resid)
    skew = stats.skew(resid)
    kurt = stats.kurtosis(resid, fisher=False)
    return {"jb": float(jb), "p_value": float(p), "skew": float(skew), "kurtosis": float(kurt)}


def hetero_indicator(yhat: np.ndarray, resid: np.ndarray):
    corr = np.corrcoef(yhat, np.abs(resid))[0, 1]
    return {"absResidualVsFittedCorrelation": float(corr), "potentialHeteroscedasticity": bool(abs(corr) > 0.3)}


def gradient_descent(X: np.ndarray, y: np.ndarray, lr: float = 0.01, iters: int = 5000):
    n, p = X.shape
    beta = np.zeros(p)
    for _ in range(iters):
        pred = X @ beta
        grad = 2.0 / n * (X.T @ (pred - y))
        beta -= lr * grad
    return beta


def main() -> int:
    args = parse_args()
    file_path = Path(args.file) if args.file else Path("sample_regression.csv")

    if not args.file:
        generate_sample(file_path)
        args.target = args.target or "target"

    if not args.target:
        print("Error: --target is required", file=sys.stderr)
        return 1

    df = pd.read_csv(file_path)
    if args.target not in df.columns:
        print(f"Error: target column '{args.target}' not found", file=sys.stderr)
        return 1

    feature_cols = [c.strip() for c in args.features.split(",")] if args.features else [c for c in df.columns if c != args.target]

    # keep numeric and drop missing
    data = df[feature_cols + [args.target]].apply(pd.to_numeric, errors="coerce").dropna()
    X = data[feature_cols].to_numpy(dtype=float)
    y = data[args.target].to_numpy(dtype=float)

    norm_stats = None
    if args.normalize:
        means = X.mean(axis=0)
        stds = X.std(axis=0, ddof=1)
        stds[stds == 0] = 1.0
        X = (X - means) / stds
        norm_stats = {"means": means.tolist(), "stds": stds.tolist()}

    split_idx = max(1, min(len(X), int(len(X) * args.split / 100)))
    X_train, y_train = X[:split_idx], y[:split_idx]
    X_test, y_test = X[split_idx:], y[split_idx:]

    X_train_c = sm.add_constant(X_train)

    if args.method == "gradient":
        beta = gradient_descent(X_train_c, y_train)
        yhat = X_train_c @ beta
        resid = y_train - yhat

        n, p = X_train_c.shape
        rss = float(np.sum(resid**2))
        mse = rss / n
        rmse = float(np.sqrt(mse))
        mae = float(np.mean(np.abs(resid)))
        tss = float(np.sum((y_train - y_train.mean()) ** 2))
        r2 = 1 - rss / tss if tss else 0
        adj_r2 = 1 - (1 - r2) * (n - 1) / max(1, n - p)
        f_stat = ((tss - rss) / max(1, p - 1)) / (rss / max(1, n - p)) if n > p else float("nan")

        xtx_inv = np.linalg.pinv(X_train_c.T @ X_train_c)
        sigma2 = rss / max(1, n - p)
        se = np.sqrt(np.diag(xtx_inv) * sigma2)
        t_stats = beta / np.where(se == 0, 1e-12, se)
        p_vals = 2 * (1 - stats.t.cdf(np.abs(t_stats), df=max(1, n - p)))

        coeff_table = [
            {"name": "intercept", "estimate": float(beta[0]), "std_error": float(se[0]), "t_stat": float(t_stats[0]), "p_value": float(p_vals[0])}
        ] + [
            {"name": feature_cols[i], "estimate": float(beta[i + 1]), "std_error": float(se[i + 1]), "t_stat": float(t_stats[i + 1]), "p_value": float(p_vals[i + 1])}
            for i in range(len(feature_cols))
        ]
    else:
        model = sm.OLS(y_train, X_train_c).fit()
        yhat = model.predict(X_train_c)
        resid = y_train - yhat
        coeff_table = [{
            "name": "intercept" if i == 0 else feature_cols[i - 1],
            "estimate": float(model.params[i]),
            "std_error": float(model.bse[i]),
            "t_stat": float(model.tvalues[i]),
            "p_value": float(model.pvalues[i]),
        } for i in range(len(model.params))]

        r2 = float(model.rsquared)
        adj_r2 = float(model.rsquared_adj)
        mse = float(np.mean(resid**2))
        rmse = float(np.sqrt(mse))
        mae = float(np.mean(np.abs(resid)))
        f_stat = float(model.fvalue) if model.fvalue is not None else float("nan")
        beta = model.params

    cond = float(np.linalg.cond(X_train_c))

    report = {
        "config": vars(args),
        "data": {
            "rows_total": int(len(X)),
            "rows_train": int(len(X_train)),
            "rows_test": int(len(X_test)),
            "target": args.target,
            "features": feature_cols,
            "normalized": bool(args.normalize),
            "normalization_stats": norm_stats,
        },
        "coefficients": coeff_table,
        "metrics": {
            "r2": float(r2),
            "adjusted_r2": float(adj_r2),
            "mse": float(mse),
            "rmse": float(rmse),
            "mae": float(mae),
            "f_statistic": float(f_stat),
        },
        "residual_analysis": {
            "jarque_bera": jarque_bera(np.asarray(resid)),
            "heteroscedasticity": hetero_indicator(np.asarray(yhat), np.asarray(resid)),
        },
        "multicollinearity": {
            "condition_number": cond,
            "warning": bool(cond > 30),
        },
    }

    if args.predict:
        pred_df = pd.read_csv(args.predict)
        pred_data = pred_df[feature_cols].apply(pd.to_numeric, errors="coerce").dropna().to_numpy(dtype=float)
        if args.normalize and norm_stats:
            means = np.array(norm_stats["means"])
            stds = np.array(norm_stats["stds"])
            pred_data = (pred_data - means) / stds
        pred_X = sm.add_constant(pred_data)
        report["predictions"] = (pred_X @ np.asarray(beta)).tolist()

    print("Model Summary")
    print(f"R^2={report['metrics']['r2']:.6f}  Adj R^2={report['metrics']['adjusted_r2']:.6f}  RMSE={report['metrics']['rmse']:.6f}  MAE={report['metrics']['mae']:.6f}")
    print(pd.DataFrame(report["coefficients"]).to_string(index=False))

    Path(args.output).write_text(json.dumps(report, indent=2), encoding="utf-8")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())