← All tasks
javacodex/java-t1 #42Not a task: repair changed code

Linear Regression Fitter (java, written by Codex)

envgap__codex__java-t1-42

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

NullPointerException: Map.of() with null values for predict/normalization + missing imports for SVD
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
pom.xml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/java-t1 #42 · read the task the agent was given
Codex wrote this java project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Linear Regression Fitter

Write a program that fits linear regression models to datasets, supporting simple and multiple regression, model evaluation metrics, residual analysis, and prediction on new data.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV data file as a command-line argument with the target variable specified via --target flag
- Support simple linear regression (one predictor) and multiple linear regression (multiple predictors) selected via --features flag (comma-separated column names; default: all non-target columns)
- Compute regression coefficients (intercept and slopes) using the ordinary least squares (OLS) method via the normal equation or gradient descent (selectable via --method flag)
- Report model evaluation metrics: R-squared, adjusted R-squared, mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and F-statistic
- Report coefficient details: estimated value, standard error, t-statistic, and p-value for each coefficient
- Perform residual analysis: compute residuals, check for normality (Shapiro-Wilk or similar test), and detect heteroscedasticity
- Support making predictions on new data via --predict flag (path to a CSV file with predictor values)
- Support data normalization/standardization via --normalize flag
- Support train/test split via --split flag (percentage for training, e.g., --split 80 for 80% train)
- Print a comprehensive model summary to console similar to statistical software output
- Save model coefficients and metrics as JSON with --output flag (default: regression_model.json)
- If no input is given, generate a sample dataset with 200 points containing a known linear relationship with noise, fit the model, display coefficients, metrics, residual analysis, and make predictions on 5 new data points
- Handle errors: multicollinearity detection (warn if condition number is high), missing values in data, non-numeric columns, and singular design matrix

Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include:
- Source code
- pom.xml with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
  <modelVersion>4.0.0</modelVersion>
  <groupId>org.tmlr</groupId>
  <artifactId>linear-regression-fitter</artifactId>
  <version>1.0.0</version>
  <properties>
    <maven.compiler.source>17</maven.compiler.source>
    <maven.compiler.target>17</maven.compiler.target>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
  </properties>
  <dependencies>
    <dependency>
      <groupId>org.apache.commons</groupId>
      <artifactId>commons-math3</artifactId>
      <version>3.6.1</version>
    </dependency>
  </dependencies>
  <build>
    <plugins>
      <plugin>
        <groupId>org.apache.maven.plugins</groupId>
        <artifactId>maven-assembly-plugin</artifactId>
        <version>3.7.1</version>
        <configuration>
          <archive>
            <manifest>
              <mainClass>LinearRegressionFitter</mainClass>
            </manifest>
          </archive>
          <descriptorRefs>
            <descriptorRef>jar-with-dependencies</descriptorRef>
          </descriptorRefs>
        </configuration>
      </plugin>
    </plugins>
  </build>
</project>
README.md
# Linear Regression Fitter (Java)

## Requirements
- Ubuntu 22.04
- JDK 17+
- Maven 3.8+

## Build
```bash
mvn -q -DskipTests package assembly:single
```

## Run
```bash
java -cp target/linear-regression-fitter-1.0.0-jar-with-dependencies.jar LinearRegressionFitter data.csv --target target
java -cp target/linear-regression-fitter-1.0.0-jar-with-dependencies.jar LinearRegressionFitter data.csv --target y --features x1,x2 --method gradient --split 80 --normalize
java -cp target/linear-regression-fitter-1.0.0-jar-with-dependencies.jar LinearRegressionFitter data.csv --target y --predict new_data.csv --output regression_model.json
```

If no input file is provided, a sample dataset is generated and fitted.
src/main/java/LinearRegressionFitter.java
import org.apache.commons.math3.distribution.TDistribution;
import org.apache.commons.math3.stat.regression.OLSMultipleLinearRegression;

import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.time.Instant;
import java.util.*;

public class LinearRegressionFitter {
    private static class Config {
        String file;
        String target;
        List<String> features;
        String method = "normal";
        String predict;
        boolean normalize;
        int split = 100;
        String output = "regression_model.json";
    }

    public static void main(String[] args) {
        try {
            Config cfg = parseArgs(args);
            if (cfg.file == null) {
                cfg.file = "sample_regression.csv";
                cfg.target = cfg.target == null ? "target" : cfg.target;
                generateSample(Paths.get(cfg.file));
            }
            if (cfg.target == null) throw new IllegalArgumentException("--target is required");

            CsvData csv = readCsv(Paths.get(cfg.file));
            Selection sel = selectNumeric(csv, cfg.target, cfg.features);

            NormStats norm = null;
            if (cfg.normalize) {
                norm = normalize(sel.X);
                sel.X = norm.Z;
            }

            SplitData split = split(sel.X, sel.y, cfg.split);

            FitResult fit = "gradient".equals(cfg.method) ? fitGradient(split.Xtrain, split.ytrain) : fitOLS(split.Xtrain, split.ytrain);

            Map<String, Object> report = new LinkedHashMap<>();
            report.put("generatedAt", Instant.now().toString());
            report.put("config", Map.of(
                "file", cfg.file,
                "target", cfg.target,
                "features", sel.featureCols,
                "method", cfg.method,
                "normalize", cfg.normalize,
                "split", cfg.split,
                "predict", cfg.predict,
                "output", cfg.output
            ));
            report.put("data", Map.of(
                "rowsTotal", sel.X.length,
                "rowsTrain", split.Xtrain.length,
                "rowsTest", split.Xtest.length,
                "normalization", norm == null ? null : Map.of("means", norm.means, "stds", norm.stds)
            ));

            List<Map<String, Object>> coeff = new ArrayList<>();
            coeff.add(coeffRow("intercept", fit.beta[0], fit.stdErr[0], fit.tStats[0], fit.pValues[0]));
            for (int i = 0; i < sel.featureCols.size(); i++) {
                coeff.add(coeffRow(sel.featureCols.get(i), fit.beta[i + 1], fit.stdErr[i + 1], fit.tStats[i + 1], fit.pValues[i + 1]));
            }
            report.put("coefficients", coeff);

            report.put("metrics", Map.of(
                "r2", fit.r2,
                "adjustedR2", fit.adjR2,
                "mse", fit.mse,
                "rmse", fit.rmse,
                "mae", fit.mae,
                "fStatistic", fit.fStat
            ));

            report.put("residualAnalysis", Map.of(
                "normality", jarqueBera(fit.residuals),
                "heteroscedasticity", heteroIndicator(fit.predictions, fit.residuals)
            ));

            double cond = conditionNumber(split.Xtrain);
            report.put("multicollinearity", Map.of("conditionNumber", cond, "warning", cond > 30));

            if (cfg.predict != null) {
                CsvData predCsv = readCsv(Paths.get(cfg.predict));
                double[][] predX = extractPredictors(predCsv, sel.featureCols);
                if (cfg.normalize && norm != null) predX = applyNorm(predX, norm);
                double[] predVals = predict(predX, fit.beta);
                report.put("predictions", predVals);
            }

            System.out.printf("R^2=%.6f Adj R^2=%.6f RMSE=%.6f MAE=%.6f%n", fit.r2, fit.adjR2, fit.rmse, fit.mae);
            for (Map<String, Object> row : coeff) {
                System.out.printf("%-12s est=% .6f se=% .6f t=% .6f p=% .6f%n", row.get("name"), row.get("estimate"), row.get("stdError"), row.get("tStat"), row.get("pValue"));
            }

            Files.writeString(Paths.get(cfg.output), toJson(report), StandardCharsets.UTF_8);
        } catch (Exception ex) {
            System.err.println("Error: " + ex.getMessage());
            System.exit(1);
        }
    }

    private static Config parseArgs(String[] args) {
        Config cfg = new Config();
        List<String> pos = new ArrayList<>();
        for (int i = 0; i < args.length; i++) {
            String a = args[i];
            if (!a.startsWith("--")) { pos.add(a); continue; }
            switch (a) {
                case "--target" -> cfg.target = args[++i];
                case "--features" -> cfg.features = Arrays.stream(args[++i].split(",")).map(String::trim).filter(s -> !s.isEmpty()).toList();
                case "--method" -> cfg.method = args[++i];
                case "--predict" -> cfg.predict = args[++i];
                case "--normalize" -> cfg.normalize = true;
                case "--split" -> cfg.split = Integer.parseInt(args[++i]);
                case "--output" -> cfg.output = args[++i];
                default -> throw new IllegalArgumentException("Unknown option: " + a);
            }
        }
        if (!pos.isEmpty()) cfg.file = pos.get(0);
        return cfg;
    }

    private record CsvData(List<String> header, List<List<String>> rows) {}
    private static CsvData readCsv(Path p) throws IOException {
        List<String> lines = Files.readAllLines(p, StandardCharsets.UTF_8);
        List<String> header = Arrays.stream(lines.get(0).split(",")).map(String::trim).toList();
        List<List<String>> rows = new ArrayList<>();
        for (int i = 1; i < lines.size(); i++) rows.add(Arrays.stream(lines.get(i).split(",")).map(String::trim).toList());
        return new CsvData(header, rows);
    }

    private static class Selection {
        List<String> featureCols;
        double[][] X;
        double[] y;
    }

    private static Selection selectNumeric(CsvData csv, String target, List<String> features) {
        Map<String, Integer> idx = new LinkedHashMap<>();
        for (int i = 0; i < csv.header.size(); i++) idx.put(csv.header.get(i), i);
        if (!idx.containsKey(target)) throw new IllegalArgumentException("Missing target column: " + target);

        List<String> fcols = features != null && !features.isEmpty() ? features : csv.header.stream().filter(c -> !c.equals(target)).toList();
        for (String f : fcols) if (!idx.containsKey(f)) throw new IllegalArgumentException("Missing feature column: " + f);

        List<double[]> X = new ArrayList<>();
        List<Double> y = new ArrayList<>();
        for (List<String> row : csv.rows) {
            try {
                double yt = Double.parseDouble(row.get(idx.get(target)));
                double[] xr = new double[fcols.size()];
                for (int i = 0; i < fcols.size(); i++) xr[i] = Double.parseDouble(row.get(idx.get(fcols.get(i))));
                X.add(xr);
                y.add(yt);
            } catch (Exception ignored) {
            }
        }

        Selection s = new Selection();
        s.featureCols = fcols;
        s.X = X.toArray(new double[0][]);
        s.y = y.stream().mapToDouble(Double::doubleValue).toArray();
        return s;
    }

    private record NormStats(double[] means, double[] stds, double[][] Z) {}

    private static NormStats normalize(double[][] X) {
        int n = X.length, p = X[0].length;
        double[] means = new double[p];
        double[] stds = new double[p];
        for (int j = 0; j < p; j++) {
            for (int i = 0; i < n; i++) means[j] += X[i][j];
            means[j] /= n;
            for (int i = 0; i < n; i++) stds[j] += Math.pow(X[i][j] - means[j], 2);
            stds[j] = Math.sqrt(stds[j] / Math.max(1, n - 1));
            if (stds[j] == 0) stds[j] = 1;
        }
        double[][] Z = new double[n][p];
        for (int i = 0; i < n; i++) for (int j = 0; j < p; j++) Z[i][j] = (X[i][j] - means[j]) / stds[j];
        return new NormStats(means, stds, Z);
    }

    private static double[][] applyNorm(double[][] X, NormStats n) {
        double[][] Z = new double[X.length][X[0].length];
        for (int i = 0; i < X.length; i++) for (int j = 0; j < X[0].length; j++) Z[i][j] = (X[i][j] - n.means[j]) / n.stds[j];
        return Z;
    }

    private record SplitData(double[][] Xtrain, double[] ytrain, double[][] Xtest, double[] ytest) {}
    private static SplitData split(double[][] X, double[] y, int trainPct) {
        int n = X.length;
        int trainN = Math.max(1, Math.min(n, (int) Math.floor((trainPct / 100.0) * n)));
        return new SplitData(Arrays.copyOfRange(X, 0, trainN), Arrays.copyOfRange(y, 0, trainN), Arrays.copyOfRange(X, trainN, n), Arrays.copyOfRange(y, trainN, n));
    }

    private static class FitResult {
        double[] beta;
        double[] stdErr;
        double[] tStats;
        double[] pValues;
        double[] residuals;
        double[] predictions;
        double r2, adjR2, mse, rmse, mae, fStat;
    }

    private static FitResult fitOLS(double[][] X, double[] y) {
        OLSMultipleLinearRegression reg = new OLSMultipleLinearRegression();
        reg.setNoIntercept(false);
        reg.newSampleData(y, X);

        double[] beta = reg.estimateRegressionParameters();
        double[] se = reg.estimateRegressionParametersStandardErrors();
        double[] resid = reg.estimateResiduals();
        double[] yhat = predict(X, beta);

        int n = y.length;
        int p = X[0].length;
        double rss = dot(resid, resid);
        double ymean = Arrays.stream(y).average().orElse(0);
        double tss = 0;
        for (double v : y) tss += (v - ymean) * (v - ymean);
        double r2 = 1 - rss / tss;
        double adj = 1 - ((1 - r2) * (n - 1)) / Math.max(1, n - p - 1);
        double mse = rss / n;
        double rmse = Math.sqrt(mse);
        double mae = Arrays.stream(resid).map(Math::abs).average().orElse(0);
        double fStat = ((tss - rss) / p) / (rss / Math.max(1, n - p - 1));

        double[] tvals = new double[beta.length];
        double[] pvals = new double[beta.length];
        TDistribution td = new TDistribution(Math.max(1, n - p - 1));
        for (int i = 0; i < beta.length; i++) {
            tvals[i] = beta[i] / (se[i] == 0 ? 1e-12 : se[i]);
            pvals[i] = 2 * (1 - td.cumulativeProbability(Math.abs(tvals[i])));
        }

        FitResult f = new FitResult();
        f.beta = beta; f.stdErr = se; f.tStats = tvals; f.pValues = pvals;
        f.residuals = resid; f.predictions = yhat;
        f.r2 = r2; f.adjR2 = adj; f.mse = mse; f.rmse = rmse; f.mae = mae; f.fStat = fStat;
        return f;
    }

    private static FitResult fitGradient(double[][] X, double[] y) {
        int n = X.length;
        int p = X[0].length + 1;
        double[][] Xc = withIntercept(X);
        double[] beta = new double[p];
        double lr = 0.01;

        for (int iter = 0; iter < 5000; iter++) {
            double[] pred = matVec(Xc, beta);
            double[] grad = new double[p];
            for (int j = 0; j < p; j++) {
                double s = 0;
                for (int i = 0; i < n; i++) s += Xc[i][j] * (pred[i] - y[i]);
                grad[j] = 2.0 * s / n;
            }
            for (int j = 0; j < p; j++) beta[j] -= lr * grad[j];
        }

        // Use OLS post-fit stats around gradient solution approximation
        OLSMultipleLinearRegression reg = new OLSMultipleLinearRegression();
        reg.setNoIntercept(false);
        reg.newSampleData(y, X);

        double[] se = reg.estimateRegressionParametersStandardErrors();
        double[] pred = matVec(Xc, beta);
        double[] resid = new double[n];
        for (int i = 0; i < n; i++) resid[i] = y[i] - pred[i];

        double rss = dot(resid, resid);
        double ymean = Arrays.stream(y).average().orElse(0);
        double tss = 0;
        for (double v : y) tss += (v - ymean) * (v - ymean);
        double r2 = 1 - rss / tss;
        double adj = 1 - ((1 - r2) * (n - 1)) / Math.max(1, n - p);
        double mse = rss / n;
        double rmse = Math.sqrt(mse);
        double mae = Arrays.stream(resid).map(Math::abs).average().orElse(0);
        double fStat = ((tss - rss) / Math.max(1, p - 1)) / (rss / Math.max(1, n - p));

        double[] tvals = new double[p];
        double[] pvals = new double[p];
        TDistribution td = new TDistribution(Math.max(1, n - p));
        for (int i = 0; i < p; i++) {
            tvals[i] = beta[i] / (se[i] == 0 ? 1e-12 : se[i]);
            pvals[i] = 2 * (1 - td.cumulativeProbability(Math.abs(tvals[i])));
        }

        FitResult f = new FitResult();
        f.beta = beta; f.stdErr = se; f.tStats = tvals; f.pValues = pvals;
        f.residuals = resid; f.predictions = pred;
        f.r2 = r2; f.adjR2 = adj; f.mse = mse; f.rmse = rmse; f.mae = mae; f.fStat = fStat;
        return f;
    }

    private static double[][] extractPredictors(CsvData csv, List<String> featureCols) {
        Map<String, Integer> idx = new LinkedHashMap<>();
        for (int i = 0; i < csv.header.size(); i++) idx.put(csv.header.get(i), i);
        List<double[]> out = new ArrayList<>();
        for (List<String> row : csv.rows) {
            try {
                double[] r = new double[featureCols.size()];
                for (int j = 0; j < featureCols.size(); j++) r[j] = Double.parseDouble(row.get(idx.get(featureCols.get(j))));
                out.add(r);
            } catch (Exception ignored) {}
        }
        return out.toArray(new double[0][]);
    }

    private static double[] predict(double[][] X, double[] beta) { return matVec(withIntercept(X), beta); }

    private static double[][] withIntercept(double[][] X) {
        double[][] out = new double[X.length][X[0].length + 1];
        for (int i = 0; i < X.length; i++) {
            out[i][0] = 1;
            System.arraycopy(X[i], 0, out[i], 1, X[0].length);
        }
        return out;
    }

    private static double[] matVec(double[][] A, double[] x) {
        double[] y = new double[A.length];
        for (int i = 0; i < A.length; i++) {
            double s = 0;
            for (int j = 0; j < A[0].length; j++) s += A[i][j] * x[j];
            y[i] = s;
        }
        return y;
    }

    private static double dot(double[] a, double[] b) { double s = 0; for (int i = 0; i < a.length; i++) s += a[i] * b[i]; return s; }

    private static Map<String, Object> coeffRow(String name, double est, double se, double t, double p) {
        return Map.of("name", name, "estimate", est, "stdError", se, "tStat", t, "pValue", p);
    }

    private static Map<String, Object> jarqueBera(double[] r) {
        int n = r.length;
        double mean = Arrays.stream(r).average().orElse(0);
        double m2 = 0, m3 = 0, m4 = 0;
        for (double v : r) {
            double d = v - mean;
            m2 += d * d;
            m3 += d * d * d;
            m4 += d * d * d * d;
        }
        m2 /= n; m3 /= n; m4 /= n;
        double skew = m3 / Math.pow(m2, 1.5);
        double kurt = m4 / (m2 * m2);
        double jb = (n / 6.0) * (skew * skew + Math.pow(kurt - 3, 2) / 4.0);
        return Map.of("jb", jb, "skew", skew, "kurtosis", kurt);
    }

    private static Map<String, Object> heteroIndicator(double[] yhat, double[] resid) {
        int n = yhat.length;
        double[] abs = Arrays.stream(resid).map(Math::abs).toArray();
        double mx = Arrays.stream(yhat).average().orElse(0);
        double my = Arrays.stream(abs).average().orElse(0);
        double num = 0, dx = 0, dy = 0;
        for (int i = 0; i < n; i++) {
            double a = yhat[i] - mx, b = abs[i] - my;
            num += a * b; dx += a * a; dy += b * b;
        }
        double corr = num / Math.sqrt((dx == 0 ? 1e-12 : dx) * (dy == 0 ? 1e-12 : dy));
        return Map.of("absResidualVsFittedCorrelation", corr, "potentialHeteroscedasticity", Math.abs(corr) > 0.3);
    }

    private static double conditionNumber(double[][] X) {
        RealMatrix m = MatrixUtils.createRealMatrix(withIntercept(X));
        SingularValueDecomposition svd = new SingularValueDecomposition(m);
        double[] s = svd.getSingularValues();
        if (s.length == 0 || s[s.length - 1] == 0) return Double.POSITIVE_INFINITY;
        return s[0] / s[s.length - 1];
    }

    private static void generateSample(Path p) throws IOException {
        Random rnd = new Random(42);
        StringBuilder sb = new StringBuilder("x1,x2,target\n");
        for (int i = 0; i < 200; i++) {
            double x1 = i / 10.0;
            double x2 = (i % 25) / 5.0;
            double noise = (rnd.nextDouble() - 0.5) * 2.0;
            double y = 3.5 + 2.0 * x1 - 1.2 * x2 + noise;
            sb.append(x1).append(',').append(x2).append(',').append(y).append('\n');
        }
        Files.writeString(p, sb.toString(), StandardCharsets.UTF_8);
    }

    private static String toJson(Object obj) {
        if (obj == null) return "null";
        if (obj instanceof String s) return '"' + s.replace("\\", "\\\\").replace("\"", "\\\"") + '"';
        if (obj instanceof Number || obj instanceof Boolean) return obj.toString();
        if (obj instanceof double[] arr) {
            StringBuilder sb = new StringBuilder("[");
            for (int i = 0; i < arr.length; i++) {
                if (i > 0) sb.append(',');
                sb.append(arr[i]);
            }
            sb.append(']');
            return sb.toString();
        }
        if (obj instanceof Map<?, ?> m) {
            StringBuilder sb = new StringBuilder("{");
            boolean first = true;
            for (Map.Entry<?, ?> e : m.entrySet()) {
                if (!first) sb.append(',');
                first = false;
                sb.append(toJson(String.valueOf(e.getKey()))).append(':').append(toJson(e.getValue()));
            }
            sb.append('}');
            return sb.toString();
        }
        if (obj instanceof Iterable<?> it) {
            StringBuilder sb = new StringBuilder("[");
            boolean first = true;
            for (Object x : it) {
                if (!first) sb.append(',');
                first = false;
                sb.append(toJson(x));
            }
            sb.append(']');
            return sb.toString();
        }
        return toJson(String.valueOf(obj));
    }
}