← All tasks
javascriptcodex/javascript-t1 #42Not a task: repair changed code

Linear Regression Fitter (javascript, written by Codex)

envgap__codex__javascript-t1-42

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

Named export jStat not found (CJS/ESM mismatch) + m.svd() is not a function
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
package.json
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/javascript-t1 #42 · read the task the agent was given
Codex wrote this javascript project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Linear Regression Fitter

Write a program that fits linear regression models to datasets, supporting simple and multiple regression, model evaluation metrics, residual analysis, and prediction on new data.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV data file as a command-line argument with the target variable specified via --target flag
- Support simple linear regression (one predictor) and multiple linear regression (multiple predictors) selected via --features flag (comma-separated column names; default: all non-target columns)
- Compute regression coefficients (intercept and slopes) using the ordinary least squares (OLS) method via the normal equation or gradient descent (selectable via --method flag)
- Report model evaluation metrics: R-squared, adjusted R-squared, mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and F-statistic
- Report coefficient details: estimated value, standard error, t-statistic, and p-value for each coefficient
- Perform residual analysis: compute residuals, check for normality (Shapiro-Wilk or similar test), and detect heteroscedasticity
- Support making predictions on new data via --predict flag (path to a CSV file with predictor values)
- Support data normalization/standardization via --normalize flag
- Support train/test split via --split flag (percentage for training, e.g., --split 80 for 80% train)
- Print a comprehensive model summary to console similar to statistical software output
- Save model coefficients and metrics as JSON with --output flag (default: regression_model.json)
- If no input is given, generate a sample dataset with 200 points containing a known linear relationship with noise, fit the model, display coefficients, metrics, residual analysis, and make predictions on 5 new data points
- Handle errors: multicollinearity detection (warn if condition number is high), missing values in data, non-numeric columns, and singular design matrix

Create a complete JavaScript project for a clean Ubuntu 22.04 machine with only Node.js 20+ (LTS) installed. Include:
- Source code
- package.json with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

4 files, exactly as written, before any repair.

package-lock.json
{
  "name": "linear-regression-fitter",
  "version": "1.0.0",
  "lockfileVersion": 3,
  "requires": true,
  "packages": {
    "": {
      "name": "linear-regression-fitter",
      "version": "1.0.0",
      "license": "MIT",
      "dependencies": {
        "jstat": "1.9.6",
        "ml-matrix": "6.12.1"
      },
      "engines": {
        "node": ">=20.0.0"
      }
    },
    "node_modules/is-any-array": {
      "version": "2.0.1",
      "resolved": "https://registry.npmjs.org/is-any-array/-/is-any-array-2.0.1.tgz",
      "integrity": "sha512-UtilS7hLRu++wb/WBAw9bNuP1Eg04Ivn1vERJck8zJthEvXCBEBpGR/33u/xLKWEQf95803oalHrVDptcAvFdQ==",
      "license": "MIT"
    },
    "node_modules/jstat": {
      "version": "1.9.6",
      "resolved": "https://registry.npmjs.org/jstat/-/jstat-1.9.6.tgz",
      "integrity": "sha512-rPBkJbK2TnA8pzs93QcDDPlKcrtZWuuCo2dVR0TFLOJSxhqfWOVCSp8aV3/oSbn+4uY4yw1URtLpHQedtmXfug=="
    },
    "node_modules/ml-array-max": {
      "version": "1.2.4",
      "resolved": "https://registry.npmjs.org/ml-array-max/-/ml-array-max-1.2.4.tgz",
      "integrity": "sha512-BlEeg80jI0tW6WaPyGxf5Sa4sqvcyY6lbSn5Vcv44lp1I2GR6AWojfUvLnGTNsIXrZ8uqWmo8VcG1WpkI2ONMQ==",
      "license": "MIT",
      "dependencies": {
        "is-any-array": "^2.0.0"
      }
    },
    "node_modules/ml-array-min": {
      "version": "1.2.3",
      "resolved": "https://registry.npmjs.org/ml-array-min/-/ml-array-min-1.2.3.tgz",
      "integrity": "sha512-VcZ5f3VZ1iihtrGvgfh/q0XlMobG6GQ8FsNyQXD3T+IlstDv85g8kfV0xUG1QPRO/t21aukaJowDzMTc7j5V6Q==",
      "license": "MIT",
      "dependencies": {
        "is-any-array": "^2.0.0"
      }
    },
    "node_modules/ml-array-rescale": {
      "version": "1.3.7",
      "resolved": "https://registry.npmjs.org/ml-array-rescale/-/ml-array-rescale-1.3.7.tgz",
      "integrity": "sha512-48NGChTouvEo9KBctDfHC3udWnQKNKEWN0ziELvY3KG25GR5cA8K8wNVzracsqSW1QEkAXjTNx+ycgAv06/1mQ==",
      "license": "MIT",
      "dependencies": {
        "is-any-array": "^2.0.0",
        "ml-array-max": "^1.2.4",
        "ml-array-min": "^1.2.3"
      }
    },
    "node_modules/ml-matrix": {
      "version": "6.12.1",
      "resolved": "https://registry.npmjs.org/ml-matrix/-/ml-matrix-6.12.1.tgz",
      "integrity": "sha512-TJ+8eOFdp+INvzR4zAuwBQJznDUfktMtOB6g/hUcGh3rcyjxbz4Te57Pgri8Q9bhSQ7Zys4IYOGhFdnlgeB6Lw==",
      "license": "MIT",
      "dependencies": {
        "is-any-array": "^2.0.1",
        "ml-array-rescale": "^1.3.7"
      }
    }
  }
}
package.json
{
  "name": "linear-regression-fitter",
  "version": "1.0.0",
  "description": "Linear regression with model summary and prediction",
  "type": "module",
  "main": "src/index.js",
  "scripts": { "start": "node src/index.js" },
  "engines": { "node": ">=20.0.0" },
  "dependencies": {
    "jstat": "1.9.6",
    "ml-matrix": "6.12.1"
  },
  "license": "MIT"
}
README.md
# Linear Regression Fitter (JavaScript)

## Requirements
- Ubuntu 22.04
- Node.js 20+

## Install
```bash
npm install
```

## Run
```bash
node src/index.js data.csv --target target
node src/index.js data.csv --target y --features x1,x2 --method gradient --split 80 --normalize
node src/index.js data.csv --target y --predict new_data.csv --output regression_model.json
```

If no input file is provided, a sample dataset is generated and fitted.
src/index.js
import fs from "node:fs";
import path from "node:path";
import process from "node:process";
import { Matrix, inverse } from "ml-matrix";
import { jStat } from "jstat";

function parseArgs(argv) {
  const cfg = {
    file: null,
    target: null,
    features: null,
    method: "normal",
    predict: null,
    normalize: false,
    split: 100,
    output: "regression_model.json"
  };
  const pos = [];
  for (let i = 0; i < argv.length; i += 1) {
    const a = argv[i];
    if (!a.startsWith("--")) {
      pos.push(a);
      continue;
    }
    if (a === "--target") cfg.target = argv[++i];
    else if (a === "--features") cfg.features = argv[++i].split(",").map((x) => x.trim()).filter(Boolean);
    else if (a === "--method") cfg.method = argv[++i];
    else if (a === "--predict") cfg.predict = argv[++i];
    else if (a === "--normalize") cfg.normalize = true;
    else if (a === "--split") cfg.split = Number.parseInt(argv[++i], 10);
    else if (a === "--output") cfg.output = argv[++i];
    else throw new Error(`Unknown option: ${a}`);
  }
  if (pos.length > 0) cfg.file = pos[0];
  return cfg;
}

function parseCsv(file) {
  const lines = fs.readFileSync(file, "utf8").trim().split(/\r?\n/);
  const header = lines[0].split(",").map((h) => h.trim());
  const rows = lines.slice(1).map((line) => line.split(",").map((x) => x.trim()));
  return { header, rows };
}

function toNumericDataset(parsed) {
  const cols = parsed.header;
  const colIdx = Object.fromEntries(cols.map((c, i) => [c, i]));
  return {
    cols,
    colIdx,
    rows: parsed.rows.map((r) => r.map((x) => (x === "" ? null : Number(x))))
  };
}

function selectData(ds, target, features) {
  if (!target) throw new Error("--target is required");
  if (!(target in ds.colIdx)) throw new Error(`Missing target column: ${target}`);
  const featureCols = features && features.length ? features : ds.cols.filter((c) => c !== target);
  featureCols.forEach((f) => { if (!(f in ds.colIdx)) throw new Error(`Missing feature column: ${f}`); });

  const X = [];
  const y = [];
  for (const row of ds.rows) {
    const yt = row[ds.colIdx[target]];
    const xf = featureCols.map((f) => row[ds.colIdx[f]]);
    if (yt === null || xf.some((v) => v === null || Number.isNaN(v)) || Number.isNaN(yt)) continue;
    X.push(xf);
    y.push(yt);
  }
  return { featureCols, X, y };
}

function normalizeCols(X) {
  const n = X.length;
  const p = X[0].length;
  const means = Array(p).fill(0);
  const stds = Array(p).fill(0);

  for (let j = 0; j < p; j += 1) means[j] = X.reduce((s, r) => s + r[j], 0) / n;
  for (let j = 0; j < p; j += 1) {
    const v = X.reduce((s, r) => s + (r[j] - means[j]) ** 2, 0) / Math.max(1, n - 1);
    stds[j] = Math.sqrt(v) || 1;
  }
  const Z = X.map((r) => r.map((v, j) => (v - means[j]) / stds[j]));
  return { Z, means, stds };
}

function trainOLS(Xraw, yraw, method) {
  const n = Xraw.length;
  const p = Xraw[0].length;
  const X = new Matrix(Xraw.map((r) => [1, ...r]));
  const y = Matrix.columnVector(yraw);

  let beta;
  if (method === "gradient") {
    beta = Matrix.zeros(p + 1, 1);
    const lr = 0.01;
    for (let iter = 0; iter < 4000; iter += 1) {
      const pred = X.mmul(beta);
      const grad = X.transpose().mmul(pred.sub(y)).mul(2 / n);
      beta = beta.sub(grad.mul(lr));
    }
  } else {
    const XtX = X.transpose().mmul(X);
    beta = inverse(XtX).mmul(X.transpose()).mmul(y);
  }

  const yhat = X.mmul(beta);
  const resid = y.sub(yhat);

  const ymean = yraw.reduce((s, v) => s + v, 0) / n;
  const rss = resid.transpose().mmul(resid).get(0, 0);
  const tss = yraw.reduce((s, v) => s + (v - ymean) ** 2, 0);
  const r2 = 1 - rss / tss;
  const adjR2 = 1 - ((1 - r2) * (n - 1)) / Math.max(1, n - p - 1);
  const mse = rss / n;
  const rmse = Math.sqrt(mse);
  const mae = resid.to1DArray().reduce((s, v) => s + Math.abs(v), 0) / n;

  const fStat = ((tss - rss) / p) / (rss / Math.max(1, n - p - 1));

  const XtXInv = inverse(X.transpose().mmul(X));
  const sigma2 = rss / Math.max(1, n - p - 1);
  const se = XtXInv.diagonal().map((d) => Math.sqrt(Math.abs(d * sigma2)));
  const coeff = beta.to1DArray();
  const tStats = coeff.map((c, i) => c / (se[i] || 1e-12));
  const dof = Math.max(1, n - p - 1);
  const pValues = tStats.map((t) => 2 * (1 - jStat.studentt.cdf(Math.abs(t), dof)));

  return {
    coefficients: coeff,
    standardErrors: se,
    tStats,
    pValues,
    residuals: resid.to1DArray(),
    predictions: yhat.to1DArray(),
    metrics: { r2, adjR2, mse, rmse, mae, fStat },
    rss,
    n,
    p
  };
}

function jarqueBera(resid) {
  const n = resid.length;
  const mean = resid.reduce((s, v) => s + v, 0) / n;
  const centered = resid.map((r) => r - mean);
  const m2 = centered.reduce((s, v) => s + v ** 2, 0) / n;
  const m3 = centered.reduce((s, v) => s + v ** 3, 0) / n;
  const m4 = centered.reduce((s, v) => s + v ** 4, 0) / n;
  const skew = m3 / Math.pow(m2, 1.5);
  const kurt = m4 / (m2 * m2);
  const jb = (n / 6) * (skew ** 2 + ((kurt - 3) ** 2) / 4);
  return { jb, skew, kurt };
}

function heteroIndicator(pred, resid) {
  const n = pred.length;
  const absR = resid.map((v) => Math.abs(v));
  const meanX = pred.reduce((s, v) => s + v, 0) / n;
  const meanY = absR.reduce((s, v) => s + v, 0) / n;
  let num = 0; let dx = 0; let dy = 0;
  for (let i = 0; i < n; i += 1) {
    const a = pred[i] - meanX;
    const b = absR[i] - meanY;
    num += a * b; dx += a * a; dy += b * b;
  }
  const corr = num / Math.sqrt((dx || 1e-12) * (dy || 1e-12));
  return { absResidualVsFittedCorrelation: corr, potentialHeteroscedasticity: Math.abs(corr) > 0.3 };
}

function conditionNumber(X) {
  const m = new Matrix(X);
  const sv = m.svd().diagonal;
  if (!sv.length || sv[sv.length - 1] === 0) return Infinity;
  return sv[0] / sv[sv.length - 1];
}

function splitData(X, y, trainPct) {
  const n = X.length;
  const trainN = Math.max(1, Math.min(n, Math.floor((trainPct / 100) * n)));
  return {
    Xtrain: X.slice(0, trainN),
    ytrain: y.slice(0, trainN),
    Xtest: X.slice(trainN),
    ytest: y.slice(trainN)
  };
}

function predictWithModel(file, featureCols, normStats, coeff) {
  const parsed = toNumericDataset(parseCsv(file));
  const X = [];
  for (const row of parsed.rows) {
    const vals = featureCols.map((f) => row[parsed.colIdx[f]]);
    if (vals.some((v) => v === null || Number.isNaN(v))) continue;
    X.push(vals);
  }
  let data = X;
  if (normStats) data = X.map((r) => r.map((v, j) => (v - normStats.means[j]) / normStats.stds[j]));
  return data.map((r) => coeff[0] + r.reduce((s, v, j) => s + v * coeff[j + 1], 0));
}

function generateSample(file) {
  const lines = ["x1,x2,target"];
  for (let i = 0; i < 200; i += 1) {
    const x1 = i / 10;
    const x2 = (i % 25) / 5;
    const noise = (Math.random() - 0.5) * 2;
    const y = 3.5 + 2.0 * x1 - 1.2 * x2 + noise;
    lines.push(`${x1.toFixed(4)},${x2.toFixed(4)},${y.toFixed(4)}`);
  }
  fs.writeFileSync(file, `${lines.join("\n")}\n`, "utf8");
}

function main() {
  try {
    const cfg = parseArgs(process.argv.slice(2));
    if (!cfg.file) {
      cfg.file = path.resolve("sample_regression.csv");
      cfg.target = "target";
      generateSample(cfg.file);
    }

    const ds = toNumericDataset(parseCsv(cfg.file));
    const selected = selectData(ds, cfg.target, cfg.features);

    let X = selected.X;
    const y = selected.y;

    let normStats = null;
    if (cfg.normalize) {
      const n = normalizeCols(X);
      X = n.Z;
      normStats = { means: n.means, stds: n.stds };
    }

    const split = splitData(X, y, cfg.split);
    const model = trainOLS(split.Xtrain, split.ytrain, cfg.method);

    const jb = jarqueBera(model.residuals);
    const hetero = heteroIndicator(model.predictions, model.residuals);
    const cond = conditionNumber(split.Xtrain);

    const coeffRows = [
      { name: "intercept", estimate: model.coefficients[0], stdError: model.standardErrors[0], tStat: model.tStats[0], pValue: model.pValues[0] },
      ...selected.featureCols.map((f, i) => ({ name: f, estimate: model.coefficients[i + 1], stdError: model.standardErrors[i + 1], tStat: model.tStats[i + 1], pValue: model.pValues[i + 1] }))
    ];

    const report = {
      config: cfg,
      data: { rowsUsed: split.Xtrain.length, rowsTotal: X.length, features: selected.featureCols, target: cfg.target },
      coefficients: coeffRows,
      metrics: model.metrics,
      residualAnalysis: { jarqueBera: jb, heteroscedasticity: hetero },
      multicollinearity: { conditionNumber: cond, warning: cond > 30 },
      trainTest: { trainRows: split.Xtrain.length, testRows: split.Xtest.length }
    };

    if (cfg.predict) report.predictions = predictWithModel(cfg.predict, selected.featureCols, normStats, model.coefficients);

    console.log("Model Summary");
    console.log(`R^2=${model.metrics.r2.toFixed(6)}  Adj R^2=${model.metrics.adjR2.toFixed(6)}  RMSE=${model.metrics.rmse.toFixed(6)}  MAE=${model.metrics.mae.toFixed(6)}`);
    console.table(coeffRows.map((r) => ({ Coefficient: r.name, Estimate: r.estimate, StdErr: r.stdError, T: r.tStat, P: r.pValue })));

    fs.writeFileSync(cfg.output, JSON.stringify(report, null, 2), "utf8");
  } catch (err) {
    console.error(`Error: ${err.message}`);
    process.exit(1);
  }
}

main();