Linear Regression Fitter (java, written by Codex)
envgap__codex__java-t1-42
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
NullPointerException: Map.of() with null values for predict/normalization + missing imports for SVD
Not a benchmark task.
- Its repair changed source code, so it is not an environment task.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/java-t1 #42 · read the task the agent was given
Codex wrote this java project from the task below. It does not run on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Linear Regression Fitter Write a program that fits linear regression models to datasets, supporting simple and multiple regression, model evaluation metrics, residual analysis, and prediction on new data. FUNCTIONAL REQUIREMENTS: - Accept a CSV data file as a command-line argument with the target variable specified via --target flag - Support simple linear regression (one predictor) and multiple linear regression (multiple predictors) selected via --features flag (comma-separated column names; default: all non-target columns) - Compute regression coefficients (intercept and slopes) using the ordinary least squares (OLS) method via the normal equation or gradient descent (selectable via --method flag) - Report model evaluation metrics: R-squared, adjusted R-squared, mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and F-statistic - Report coefficient details: estimated value, standard error, t-statistic, and p-value for each coefficient - Perform residual analysis: compute residuals, check for normality (Shapiro-Wilk or similar test), and detect heteroscedasticity - Support making predictions on new data via --predict flag (path to a CSV file with predictor values) - Support data normalization/standardization via --normalize flag - Support train/test split via --split flag (percentage for training, e.g., --split 80 for 80% train) - Print a comprehensive model summary to console similar to statistical software output - Save model coefficients and metrics as JSON with --output flag (default: regression_model.json) - If no input is given, generate a sample dataset with 200 points containing a known linear relationship with noise, fit the model, display coefficients, metrics, residual analysis, and make predictions on 5 new data points - Handle errors: multicollinearity detection (warn if condition number is high), missing values in data, non-numeric columns, and singular design matrix Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>org.tmlr</groupId>
<artifactId>linear-regression-fitter</artifactId>
<version>1.0.0</version>
<properties>
<maven.compiler.source>17</maven.compiler.source>
<maven.compiler.target>17</maven.compiler.target>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
</properties>
<dependencies>
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-math3</artifactId>
<version>3.6.1</version>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-assembly-plugin</artifactId>
<version>3.7.1</version>
<configuration>
<archive>
<manifest>
<mainClass>LinearRegressionFitter</mainClass>
</manifest>
</archive>
<descriptorRefs>
<descriptorRef>jar-with-dependencies</descriptorRef>
</descriptorRefs>
</configuration>
</plugin>
</plugins>
</build>
</project>
README.md
# Linear Regression Fitter (Java) ## Requirements - Ubuntu 22.04 - JDK 17+ - Maven 3.8+ ## Build ```bash mvn -q -DskipTests package assembly:single ``` ## Run ```bash java -cp target/linear-regression-fitter-1.0.0-jar-with-dependencies.jar LinearRegressionFitter data.csv --target target java -cp target/linear-regression-fitter-1.0.0-jar-with-dependencies.jar LinearRegressionFitter data.csv --target y --features x1,x2 --method gradient --split 80 --normalize java -cp target/linear-regression-fitter-1.0.0-jar-with-dependencies.jar LinearRegressionFitter data.csv --target y --predict new_data.csv --output regression_model.json ``` If no input file is provided, a sample dataset is generated and fitted.
src/main/java/LinearRegressionFitter.java
import org.apache.commons.math3.distribution.TDistribution;
import org.apache.commons.math3.stat.regression.OLSMultipleLinearRegression;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.time.Instant;
import java.util.*;
public class LinearRegressionFitter {
private static class Config {
String file;
String target;
List<String> features;
String method = "normal";
String predict;
boolean normalize;
int split = 100;
String output = "regression_model.json";
}
public static void main(String[] args) {
try {
Config cfg = parseArgs(args);
if (cfg.file == null) {
cfg.file = "sample_regression.csv";
cfg.target = cfg.target == null ? "target" : cfg.target;
generateSample(Paths.get(cfg.file));
}
if (cfg.target == null) throw new IllegalArgumentException("--target is required");
CsvData csv = readCsv(Paths.get(cfg.file));
Selection sel = selectNumeric(csv, cfg.target, cfg.features);
NormStats norm = null;
if (cfg.normalize) {
norm = normalize(sel.X);
sel.X = norm.Z;
}
SplitData split = split(sel.X, sel.y, cfg.split);
FitResult fit = "gradient".equals(cfg.method) ? fitGradient(split.Xtrain, split.ytrain) : fitOLS(split.Xtrain, split.ytrain);
Map<String, Object> report = new LinkedHashMap<>();
report.put("generatedAt", Instant.now().toString());
report.put("config", Map.of(
"file", cfg.file,
"target", cfg.target,
"features", sel.featureCols,
"method", cfg.method,
"normalize", cfg.normalize,
"split", cfg.split,
"predict", cfg.predict,
"output", cfg.output
));
report.put("data", Map.of(
"rowsTotal", sel.X.length,
"rowsTrain", split.Xtrain.length,
"rowsTest", split.Xtest.length,
"normalization", norm == null ? null : Map.of("means", norm.means, "stds", norm.stds)
));
List<Map<String, Object>> coeff = new ArrayList<>();
coeff.add(coeffRow("intercept", fit.beta[0], fit.stdErr[0], fit.tStats[0], fit.pValues[0]));
for (int i = 0; i < sel.featureCols.size(); i++) {
coeff.add(coeffRow(sel.featureCols.get(i), fit.beta[i + 1], fit.stdErr[i + 1], fit.tStats[i + 1], fit.pValues[i + 1]));
}
report.put("coefficients", coeff);
report.put("metrics", Map.of(
"r2", fit.r2,
"adjustedR2", fit.adjR2,
"mse", fit.mse,
"rmse", fit.rmse,
"mae", fit.mae,
"fStatistic", fit.fStat
));
report.put("residualAnalysis", Map.of(
"normality", jarqueBera(fit.residuals),
"heteroscedasticity", heteroIndicator(fit.predictions, fit.residuals)
));
double cond = conditionNumber(split.Xtrain);
report.put("multicollinearity", Map.of("conditionNumber", cond, "warning", cond > 30));
if (cfg.predict != null) {
CsvData predCsv = readCsv(Paths.get(cfg.predict));
double[][] predX = extractPredictors(predCsv, sel.featureCols);
if (cfg.normalize && norm != null) predX = applyNorm(predX, norm);
double[] predVals = predict(predX, fit.beta);
report.put("predictions", predVals);
}
System.out.printf("R^2=%.6f Adj R^2=%.6f RMSE=%.6f MAE=%.6f%n", fit.r2, fit.adjR2, fit.rmse, fit.mae);
for (Map<String, Object> row : coeff) {
System.out.printf("%-12s est=% .6f se=% .6f t=% .6f p=% .6f%n", row.get("name"), row.get("estimate"), row.get("stdError"), row.get("tStat"), row.get("pValue"));
}
Files.writeString(Paths.get(cfg.output), toJson(report), StandardCharsets.UTF_8);
} catch (Exception ex) {
System.err.println("Error: " + ex.getMessage());
System.exit(1);
}
}
private static Config parseArgs(String[] args) {
Config cfg = new Config();
List<String> pos = new ArrayList<>();
for (int i = 0; i < args.length; i++) {
String a = args[i];
if (!a.startsWith("--")) { pos.add(a); continue; }
switch (a) {
case "--target" -> cfg.target = args[++i];
case "--features" -> cfg.features = Arrays.stream(args[++i].split(",")).map(String::trim).filter(s -> !s.isEmpty()).toList();
case "--method" -> cfg.method = args[++i];
case "--predict" -> cfg.predict = args[++i];
case "--normalize" -> cfg.normalize = true;
case "--split" -> cfg.split = Integer.parseInt(args[++i]);
case "--output" -> cfg.output = args[++i];
default -> throw new IllegalArgumentException("Unknown option: " + a);
}
}
if (!pos.isEmpty()) cfg.file = pos.get(0);
return cfg;
}
private record CsvData(List<String> header, List<List<String>> rows) {}
private static CsvData readCsv(Path p) throws IOException {
List<String> lines = Files.readAllLines(p, StandardCharsets.UTF_8);
List<String> header = Arrays.stream(lines.get(0).split(",")).map(String::trim).toList();
List<List<String>> rows = new ArrayList<>();
for (int i = 1; i < lines.size(); i++) rows.add(Arrays.stream(lines.get(i).split(",")).map(String::trim).toList());
return new CsvData(header, rows);
}
private static class Selection {
List<String> featureCols;
double[][] X;
double[] y;
}
private static Selection selectNumeric(CsvData csv, String target, List<String> features) {
Map<String, Integer> idx = new LinkedHashMap<>();
for (int i = 0; i < csv.header.size(); i++) idx.put(csv.header.get(i), i);
if (!idx.containsKey(target)) throw new IllegalArgumentException("Missing target column: " + target);
List<String> fcols = features != null && !features.isEmpty() ? features : csv.header.stream().filter(c -> !c.equals(target)).toList();
for (String f : fcols) if (!idx.containsKey(f)) throw new IllegalArgumentException("Missing feature column: " + f);
List<double[]> X = new ArrayList<>();
List<Double> y = new ArrayList<>();
for (List<String> row : csv.rows) {
try {
double yt = Double.parseDouble(row.get(idx.get(target)));
double[] xr = new double[fcols.size()];
for (int i = 0; i < fcols.size(); i++) xr[i] = Double.parseDouble(row.get(idx.get(fcols.get(i))));
X.add(xr);
y.add(yt);
} catch (Exception ignored) {
}
}
Selection s = new Selection();
s.featureCols = fcols;
s.X = X.toArray(new double[0][]);
s.y = y.stream().mapToDouble(Double::doubleValue).toArray();
return s;
}
private record NormStats(double[] means, double[] stds, double[][] Z) {}
private static NormStats normalize(double[][] X) {
int n = X.length, p = X[0].length;
double[] means = new double[p];
double[] stds = new double[p];
for (int j = 0; j < p; j++) {
for (int i = 0; i < n; i++) means[j] += X[i][j];
means[j] /= n;
for (int i = 0; i < n; i++) stds[j] += Math.pow(X[i][j] - means[j], 2);
stds[j] = Math.sqrt(stds[j] / Math.max(1, n - 1));
if (stds[j] == 0) stds[j] = 1;
}
double[][] Z = new double[n][p];
for (int i = 0; i < n; i++) for (int j = 0; j < p; j++) Z[i][j] = (X[i][j] - means[j]) / stds[j];
return new NormStats(means, stds, Z);
}
private static double[][] applyNorm(double[][] X, NormStats n) {
double[][] Z = new double[X.length][X[0].length];
for (int i = 0; i < X.length; i++) for (int j = 0; j < X[0].length; j++) Z[i][j] = (X[i][j] - n.means[j]) / n.stds[j];
return Z;
}
private record SplitData(double[][] Xtrain, double[] ytrain, double[][] Xtest, double[] ytest) {}
private static SplitData split(double[][] X, double[] y, int trainPct) {
int n = X.length;
int trainN = Math.max(1, Math.min(n, (int) Math.floor((trainPct / 100.0) * n)));
return new SplitData(Arrays.copyOfRange(X, 0, trainN), Arrays.copyOfRange(y, 0, trainN), Arrays.copyOfRange(X, trainN, n), Arrays.copyOfRange(y, trainN, n));
}
private static class FitResult {
double[] beta;
double[] stdErr;
double[] tStats;
double[] pValues;
double[] residuals;
double[] predictions;
double r2, adjR2, mse, rmse, mae, fStat;
}
private static FitResult fitOLS(double[][] X, double[] y) {
OLSMultipleLinearRegression reg = new OLSMultipleLinearRegression();
reg.setNoIntercept(false);
reg.newSampleData(y, X);
double[] beta = reg.estimateRegressionParameters();
double[] se = reg.estimateRegressionParametersStandardErrors();
double[] resid = reg.estimateResiduals();
double[] yhat = predict(X, beta);
int n = y.length;
int p = X[0].length;
double rss = dot(resid, resid);
double ymean = Arrays.stream(y).average().orElse(0);
double tss = 0;
for (double v : y) tss += (v - ymean) * (v - ymean);
double r2 = 1 - rss / tss;
double adj = 1 - ((1 - r2) * (n - 1)) / Math.max(1, n - p - 1);
double mse = rss / n;
double rmse = Math.sqrt(mse);
double mae = Arrays.stream(resid).map(Math::abs).average().orElse(0);
double fStat = ((tss - rss) / p) / (rss / Math.max(1, n - p - 1));
double[] tvals = new double[beta.length];
double[] pvals = new double[beta.length];
TDistribution td = new TDistribution(Math.max(1, n - p - 1));
for (int i = 0; i < beta.length; i++) {
tvals[i] = beta[i] / (se[i] == 0 ? 1e-12 : se[i]);
pvals[i] = 2 * (1 - td.cumulativeProbability(Math.abs(tvals[i])));
}
FitResult f = new FitResult();
f.beta = beta; f.stdErr = se; f.tStats = tvals; f.pValues = pvals;
f.residuals = resid; f.predictions = yhat;
f.r2 = r2; f.adjR2 = adj; f.mse = mse; f.rmse = rmse; f.mae = mae; f.fStat = fStat;
return f;
}
private static FitResult fitGradient(double[][] X, double[] y) {
int n = X.length;
int p = X[0].length + 1;
double[][] Xc = withIntercept(X);
double[] beta = new double[p];
double lr = 0.01;
for (int iter = 0; iter < 5000; iter++) {
double[] pred = matVec(Xc, beta);
double[] grad = new double[p];
for (int j = 0; j < p; j++) {
double s = 0;
for (int i = 0; i < n; i++) s += Xc[i][j] * (pred[i] - y[i]);
grad[j] = 2.0 * s / n;
}
for (int j = 0; j < p; j++) beta[j] -= lr * grad[j];
}
// Use OLS post-fit stats around gradient solution approximation
OLSMultipleLinearRegression reg = new OLSMultipleLinearRegression();
reg.setNoIntercept(false);
reg.newSampleData(y, X);
double[] se = reg.estimateRegressionParametersStandardErrors();
double[] pred = matVec(Xc, beta);
double[] resid = new double[n];
for (int i = 0; i < n; i++) resid[i] = y[i] - pred[i];
double rss = dot(resid, resid);
double ymean = Arrays.stream(y).average().orElse(0);
double tss = 0;
for (double v : y) tss += (v - ymean) * (v - ymean);
double r2 = 1 - rss / tss;
double adj = 1 - ((1 - r2) * (n - 1)) / Math.max(1, n - p);
double mse = rss / n;
double rmse = Math.sqrt(mse);
double mae = Arrays.stream(resid).map(Math::abs).average().orElse(0);
double fStat = ((tss - rss) / Math.max(1, p - 1)) / (rss / Math.max(1, n - p));
double[] tvals = new double[p];
double[] pvals = new double[p];
TDistribution td = new TDistribution(Math.max(1, n - p));
for (int i = 0; i < p; i++) {
tvals[i] = beta[i] / (se[i] == 0 ? 1e-12 : se[i]);
pvals[i] = 2 * (1 - td.cumulativeProbability(Math.abs(tvals[i])));
}
FitResult f = new FitResult();
f.beta = beta; f.stdErr = se; f.tStats = tvals; f.pValues = pvals;
f.residuals = resid; f.predictions = pred;
f.r2 = r2; f.adjR2 = adj; f.mse = mse; f.rmse = rmse; f.mae = mae; f.fStat = fStat;
return f;
}
private static double[][] extractPredictors(CsvData csv, List<String> featureCols) {
Map<String, Integer> idx = new LinkedHashMap<>();
for (int i = 0; i < csv.header.size(); i++) idx.put(csv.header.get(i), i);
List<double[]> out = new ArrayList<>();
for (List<String> row : csv.rows) {
try {
double[] r = new double[featureCols.size()];
for (int j = 0; j < featureCols.size(); j++) r[j] = Double.parseDouble(row.get(idx.get(featureCols.get(j))));
out.add(r);
} catch (Exception ignored) {}
}
return out.toArray(new double[0][]);
}
private static double[] predict(double[][] X, double[] beta) { return matVec(withIntercept(X), beta); }
private static double[][] withIntercept(double[][] X) {
double[][] out = new double[X.length][X[0].length + 1];
for (int i = 0; i < X.length; i++) {
out[i][0] = 1;
System.arraycopy(X[i], 0, out[i], 1, X[0].length);
}
return out;
}
private static double[] matVec(double[][] A, double[] x) {
double[] y = new double[A.length];
for (int i = 0; i < A.length; i++) {
double s = 0;
for (int j = 0; j < A[0].length; j++) s += A[i][j] * x[j];
y[i] = s;
}
return y;
}
private static double dot(double[] a, double[] b) { double s = 0; for (int i = 0; i < a.length; i++) s += a[i] * b[i]; return s; }
private static Map<String, Object> coeffRow(String name, double est, double se, double t, double p) {
return Map.of("name", name, "estimate", est, "stdError", se, "tStat", t, "pValue", p);
}
private static Map<String, Object> jarqueBera(double[] r) {
int n = r.length;
double mean = Arrays.stream(r).average().orElse(0);
double m2 = 0, m3 = 0, m4 = 0;
for (double v : r) {
double d = v - mean;
m2 += d * d;
m3 += d * d * d;
m4 += d * d * d * d;
}
m2 /= n; m3 /= n; m4 /= n;
double skew = m3 / Math.pow(m2, 1.5);
double kurt = m4 / (m2 * m2);
double jb = (n / 6.0) * (skew * skew + Math.pow(kurt - 3, 2) / 4.0);
return Map.of("jb", jb, "skew", skew, "kurtosis", kurt);
}
private static Map<String, Object> heteroIndicator(double[] yhat, double[] resid) {
int n = yhat.length;
double[] abs = Arrays.stream(resid).map(Math::abs).toArray();
double mx = Arrays.stream(yhat).average().orElse(0);
double my = Arrays.stream(abs).average().orElse(0);
double num = 0, dx = 0, dy = 0;
for (int i = 0; i < n; i++) {
double a = yhat[i] - mx, b = abs[i] - my;
num += a * b; dx += a * a; dy += b * b;
}
double corr = num / Math.sqrt((dx == 0 ? 1e-12 : dx) * (dy == 0 ? 1e-12 : dy));
return Map.of("absResidualVsFittedCorrelation", corr, "potentialHeteroscedasticity", Math.abs(corr) > 0.3);
}
private static double conditionNumber(double[][] X) {
RealMatrix m = MatrixUtils.createRealMatrix(withIntercept(X));
SingularValueDecomposition svd = new SingularValueDecomposition(m);
double[] s = svd.getSingularValues();
if (s.length == 0 || s[s.length - 1] == 0) return Double.POSITIVE_INFINITY;
return s[0] / s[s.length - 1];
}
private static void generateSample(Path p) throws IOException {
Random rnd = new Random(42);
StringBuilder sb = new StringBuilder("x1,x2,target\n");
for (int i = 0; i < 200; i++) {
double x1 = i / 10.0;
double x2 = (i % 25) / 5.0;
double noise = (rnd.nextDouble() - 0.5) * 2.0;
double y = 3.5 + 2.0 * x1 - 1.2 * x2 + noise;
sb.append(x1).append(',').append(x2).append(',').append(y).append('\n');
}
Files.writeString(p, sb.toString(), StandardCharsets.UTF_8);
}
private static String toJson(Object obj) {
if (obj == null) return "null";
if (obj instanceof String s) return '"' + s.replace("\\", "\\\\").replace("\"", "\\\"") + '"';
if (obj instanceof Number || obj instanceof Boolean) return obj.toString();
if (obj instanceof double[] arr) {
StringBuilder sb = new StringBuilder("[");
for (int i = 0; i < arr.length; i++) {
if (i > 0) sb.append(',');
sb.append(arr[i]);
}
sb.append(']');
return sb.toString();
}
if (obj instanceof Map<?, ?> m) {
StringBuilder sb = new StringBuilder("{");
boolean first = true;
for (Map.Entry<?, ?> e : m.entrySet()) {
if (!first) sb.append(',');
first = false;
sb.append(toJson(String.valueOf(e.getKey()))).append(':').append(toJson(e.getValue()));
}
sb.append('}');
return sb.toString();
}
if (obj instanceof Iterable<?> it) {
StringBuilder sb = new StringBuilder("[");
boolean first = true;
for (Object x : it) {
if (!first) sb.append(',');
first = false;
sb.append(toJson(x));
}
sb.append(']');
return sb.toString();
}
return toJson(String.valueOf(obj));
}
}