← All tasks
cppclaude-code/cpp-t1 #42Not a task: already works

Linear Regression Fitter (cpp, written by Claude Code)

envgap__claude-code__cpp-t1-42

Written by a coding agent; not on GitHubWritten 2026-02-28

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/cpp-t1 #42 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Linear Regression Fitter

Write a program that fits linear regression models to datasets, supporting simple and multiple regression, model evaluation metrics, residual analysis, and prediction on new data.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV data file as a command-line argument with the target variable specified via --target flag
- Support simple linear regression (one predictor) and multiple linear regression (multiple predictors) selected via --features flag (comma-separated column names; default: all non-target columns)
- Compute regression coefficients (intercept and slopes) using the ordinary least squares (OLS) method via the normal equation or gradient descent (selectable via --method flag)
- Report model evaluation metrics: R-squared, adjusted R-squared, mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and F-statistic
- Report coefficient details: estimated value, standard error, t-statistic, and p-value for each coefficient
- Perform residual analysis: compute residuals, check for normality (Shapiro-Wilk or similar test), and detect heteroscedasticity
- Support making predictions on new data via --predict flag (path to a CSV file with predictor values)
- Support data normalization/standardization via --normalize flag
- Support train/test split via --split flag (percentage for training, e.g., --split 80 for 80% train)
- Print a comprehensive model summary to console similar to statistical software output
- Save model coefficients and metrics as JSON with --output flag (default: regression_model.json)
- If no input is given, generate a sample dataset with 200 points containing a known linear relationship with noise, fit the model, display coefficients, metrics, residual analysis, and make predictions on 5 new data points
- Handle errors: multicollinearity detection (warn if condition number is high), missing values in data, non-numeric columns, and singular design matrix

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.14)
project(LinearRegressionFitter VERSION 1.0.0 LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)

include(FetchContent)

# Fetch Eigen
FetchContent_Declare(
    eigen
    GIT_REPOSITORY https://gitlab.com/libeigen/eigen.git
    GIT_TAG 3.4.0
)
FetchContent_MakeAvailable(eigen)

# Fetch matplotplusplus
FetchContent_Declare(
    matplotplusplus
    GIT_REPOSITORY https://github.com/alandefreitas/matplotplusplus.git
    GIT_TAG v1.2.1
)
FetchContent_MakeAvailable(matplotplusplus)

add_executable(linear_regression main.cpp)

target_link_libraries(linear_regression
    PRIVATE
        Eigen3::Eigen
        matplot
)

install(TARGETS linear_regression DESTINATION bin)
main.cpp
/**
 * Linear Regression Fitter
 * Fits OLS linear regression with metrics (R-squared, MSE, RMSE),
 * residual analysis, and prediction capabilities.
 *
 * Dependencies: Eigen (linear algebra), matplotplusplus (visualization)
 */

#include <Eigen/Dense>
#include <matplot/matplot.h>

#include <iostream>
#include <fstream>
#include <vector>
#include <string>
#include <cmath>
#include <iomanip>
#include <sstream>
#include <random>
#include <algorithm>
#include <numeric>

using Eigen::MatrixXd;
using Eigen::VectorXd;

class LinearRegressionFitter {
private:
    VectorXd coefficients_;
    double intercept_;
    std::vector<std::string> featureNames_;

public:
    LinearRegressionFitter() : intercept_(0.0) {}

    void fit(const MatrixXd& X, const VectorXd& y, const std::vector<std::string>& names) {
        featureNames_ = names;

        // Add intercept column (column of ones)
        MatrixXd Xa(X.rows(), X.cols() + 1);
        Xa << MatrixXd::Ones(X.rows(), 1), X;

        // OLS: beta = (X^T X)^{-1} X^T y
        VectorXd beta = (Xa.transpose() * Xa).ldlt().solve(Xa.transpose() * y);

        intercept_ = beta(0);
        coefficients_ = beta.tail(X.cols());
    }

    VectorXd predict(const MatrixXd& X) const {
        return (X * coefficients_).array() + intercept_;
    }

    struct Metrics {
        double r_squared;
        double mse;
        double rmse;
    };

    Metrics computeMetrics(const VectorXd& yTrue, const VectorXd& yPred) const {
        double yMean = yTrue.mean();
        VectorXd residuals = yTrue - yPred;
        double ssRes = residuals.squaredNorm();
        double ssTot = (yTrue.array() - yMean).matrix().squaredNorm();
        double mse = ssRes / yTrue.size();

        return {1.0 - ssRes / ssTot, mse, std::sqrt(mse)};
    }

    struct ResidualStats {
        double mean, std, min, max, median, skewness, kurtosis;
    };

    ResidualStats residualAnalysis(const VectorXd& yTrue, const VectorXd& yPred) const {
        VectorXd residuals = yTrue - yPred;
        int n = residuals.size();

        double mean = residuals.mean();
        double variance = (residuals.array() - mean).square().sum() / (n - 1);
        double stdDev = std::sqrt(variance);

        std::vector<double> sorted(residuals.data(), residuals.data() + n);
        std::sort(sorted.begin(), sorted.end());

        double median = (n % 2 == 0)
            ? (sorted[n / 2 - 1] + sorted[n / 2]) / 2.0
            : sorted[n / 2];

        double m3 = 0, m4 = 0;
        for (int i = 0; i < n; i++) {
            double z = (residuals(i) - mean) / stdDev;
            m3 += z * z * z;
            m4 += z * z * z * z;
        }
        m3 /= n;
        m4 = m4 / n - 3.0;

        return {mean, stdDev, sorted.front(), sorted.back(), median, m3, m4};
    }

    void printCoefficients() const {
        std::cout << std::endl << std::string(45, '=') << std::endl;
        std::cout << "  Model Coefficients" << std::endl;
        std::cout << std::string(45, '=') << std::endl;
        std::cout << std::fixed << std::setprecision(6);
        std::cout << "  Intercept : " << intercept_ << std::endl;
        for (int i = 0; i < coefficients_.size(); i++) {
            std::string name = (i < static_cast<int>(featureNames_.size()))
                ? featureNames_[i] : "x" + std::to_string(i + 1);
            std::cout << "  " << std::left << std::setw(12) << name
                      << ": " << coefficients_(i) << std::endl;
        }
        std::cout << std::string(45, '=') << std::endl;
    }

    void printMetrics(const Metrics& m, const std::string& label) const {
        std::cout << std::endl << std::string(45, '=') << std::endl;
        std::cout << "  " << label << " Set Metrics" << std::endl;
        std::cout << std::string(45, '=') << std::endl;
        std::cout << std::fixed << std::setprecision(6);
        std::cout << "  R-squared : " << m.r_squared << std::endl;
        std::cout << "  MSE       : " << m.mse << std::endl;
        std::cout << "  RMSE      : " << m.rmse << std::endl;
        std::cout << std::string(45, '=') << std::endl;
    }

    void printResidualAnalysis(const ResidualStats& s) const {
        std::cout << std::endl << std::string(45, '=') << std::endl;
        std::cout << "  Residual Analysis" << std::endl;
        std::cout << std::string(45, '=') << std::endl;
        std::cout << std::fixed << std::setprecision(6);
        std::cout << "  Mean     : " << s.mean << std::endl;
        std::cout << "  Std Dev  : " << s.std << std::endl;
        std::cout << "  Min      : " << s.min << std::endl;
        std::cout << "  Max      : " << s.max << std::endl;
        std::cout << "  Median   : " << s.median << std::endl;
        std::cout << "  Skewness : " << s.skewness << std::endl;
        std::cout << "  Kurtosis : " << s.kurtosis << std::endl;
        std::cout << std::string(45, '=') << std::endl;
    }

    /**
     * Plot residuals vs predicted values and save to file using matplotplusplus.
     */
    void plotResiduals(const VectorXd& yTrue, const VectorXd& yPred,
                       const std::string& outputPath = "residuals.png") const {
        VectorXd residuals = yTrue - yPred;

        std::vector<double> predVec(yPred.data(), yPred.data() + yPred.size());
        std::vector<double> resVec(residuals.data(), residuals.data() + residuals.size());

        auto fig = matplot::figure(true);
        fig->size(800, 600);

        // Residuals vs Predicted scatter plot
        matplot::subplot(1, 2, 0);
        matplot::scatter(predVec, resVec);
        matplot::xlabel("Predicted Values");
        matplot::ylabel("Residuals");
        matplot::title("Residuals vs Predicted");
        matplot::hold(matplot::on);
        std::vector<double> zeroLine(predVec.size(), 0.0);
        auto sortedPred = predVec;
        std::sort(sortedPred.begin(), sortedPred.end());
        std::vector<double> zeros(sortedPred.size(), 0.0);
        matplot::plot(sortedPred, zeros)->color("red").line_style("--");
        matplot::hold(matplot::off);

        // Residual histogram
        matplot::subplot(1, 2, 1);
        matplot::hist(resVec, 20);
        matplot::xlabel("Residual");
        matplot::ylabel("Frequency");
        matplot::title("Residual Distribution");

        matplot::save(outputPath);
        std::cout << "\nResidual plots saved to: " << outputPath << std::endl;
    }

    /**
     * Plot actual vs predicted values.
     */
    void plotActualVsPredicted(const VectorXd& yTrue, const VectorXd& yPred,
                                const std::string& outputPath = "actual_vs_predicted.png") const {
        std::vector<double> trueVec(yTrue.data(), yTrue.data() + yTrue.size());
        std::vector<double> predVec(yPred.data(), yPred.data() + yPred.size());

        auto fig = matplot::figure(true);
        fig->size(600, 600);

        matplot::scatter(trueVec, predVec);
        matplot::xlabel("Actual Values");
        matplot::ylabel("Predicted Values");
        matplot::title("Actual vs Predicted");

        // Perfect prediction line
        double minVal = *std::min_element(trueVec.begin(), trueVec.end());
        double maxVal = *std::max_element(trueVec.begin(), trueVec.end());
        matplot::hold(matplot::on);
        matplot::plot({minVal, maxVal}, {minVal, maxVal})->color("red").line_style("--");
        matplot::hold(matplot::off);

        matplot::save(outputPath);
        std::cout << "Actual vs Predicted plot saved to: " << outputPath << std::endl;
    }

    std::string toJson(const Metrics& m, const ResidualStats& rs) const {
        std::ostringstream oss;
        oss << std::fixed << std::setprecision(6);
        oss << "{\n  \"intercept\": " << intercept_ << ",\n  \"coefficients\": {";
        for (int i = 0; i < coefficients_.size(); i++) {
            std::string name = (i < static_cast<int>(featureNames_.size()))
                ? featureNames_[i] : "x" + std::to_string(i + 1);
            if (i > 0) oss << ",";
            oss << "\n    \"" << name << "\": " << coefficients_(i);
        }
        oss << "\n  },\n  \"metrics\": {\n"
            << "    \"r_squared\": " << m.r_squared << ",\n"
            << "    \"mse\": " << m.mse << ",\n"
            << "    \"rmse\": " << m.rmse << "\n  },\n"
            << "  \"residual_analysis\": {\n"
            << "    \"mean\": " << rs.mean << ",\n"
            << "    \"std\": " << rs.std << ",\n"
            << "    \"min\": " << rs.min << ",\n"
            << "    \"max\": " << rs.max << ",\n"
            << "    \"median\": " << rs.median << ",\n"
            << "    \"skewness\": " << rs.skewness << ",\n"
            << "    \"kurtosis\": " << rs.kurtosis << "\n  }\n}";
        return oss.str();
    }

    double getIntercept() const { return intercept_; }
    const VectorXd& getCoefficients() const { return coefficients_; }
};

struct DataSet {
    MatrixXd X;
    VectorXd y;
    std::vector<std::string> featureNames;
};

DataSet generateSyntheticData(int nSamples, int nFeatures, double noise, unsigned seed) {
    std::mt19937 rng(seed);
    std::normal_distribution<double> normal(0.0, 1.0);

    std::vector<double> trueCoefs(nFeatures);
    for (auto& c : trueCoefs) c = normal(rng) * 5.0;

    MatrixXd X(nSamples, nFeatures);
    VectorXd y(nSamples);

    for (int i = 0; i < nSamples; i++) {
        double yi = 15.0;
        for (int j = 0; j < nFeatures; j++) {
            X(i, j) = normal(rng) * 10.0;
            yi += trueCoefs[j] * X(i, j);
        }
        yi += normal(rng) * noise;
        y(i) = yi;
    }

    std::vector<std::string> names;
    for (int j = 0; j < nFeatures; j++) {
        names.push_back("feature_" + std::to_string(j + 1));
    }
    return {X, y, names};
}

DataSet loadFromFile(const std::string& filepath) {
    std::ifstream file(filepath);
    if (!file.is_open()) {
        throw std::runtime_error("Cannot open file: " + filepath);
    }

    std::string headerLine;
    std::getline(file, headerLine);

    std::vector<std::string> headers;
    std::istringstream hss(headerLine);
    std::string token;
    while (std::getline(hss, token, ',')) {
        // Trim whitespace
        token.erase(0, token.find_first_not_of(" \t\r\n"));
        token.erase(token.find_last_not_of(" \t\r\n") + 1);
        headers.push_back(token);
    }

    std::vector<std::vector<double>> rows;
    std::string line;
    while (std::getline(file, line)) {
        if (line.empty()) continue;
        std::istringstream lss(line);
        std::vector<double> row;
        std::string val;
        while (std::getline(lss, val, ',')) {
            row.push_back(std::stod(val));
        }
        rows.push_back(row);
    }

    int nFeatures = static_cast<int>(headers.size()) - 1;
    int nRows = static_cast<int>(rows.size());
    MatrixXd X(nRows, nFeatures);
    VectorXd y(nRows);

    for (int i = 0; i < nRows; i++) {
        for (int j = 0; j < nFeatures; j++) {
            X(i, j) = rows[i][j];
        }
        y(i) = rows[i][nFeatures];
    }

    std::vector<std::string> names(headers.begin(), headers.begin() + nFeatures);
    return {X, y, names};
}

void trainTestSplit(const DataSet& ds, double testRatio, unsigned seed,
                    MatrixXd& xTrain, VectorXd& yTrain,
                    MatrixXd& xTest, VectorXd& yTest) {
    std::mt19937 rng(seed);
    std::uniform_real_distribution<double> dist(0.0, 1.0);

    std::vector<int> trainIdx, testIdx;
    for (int i = 0; i < ds.X.rows(); i++) {
        if (dist(rng) < testRatio) testIdx.push_back(i);
        else trainIdx.push_back(i);
    }

    xTrain.resize(trainIdx.size(), ds.X.cols());
    yTrain.resize(trainIdx.size());
    xTest.resize(testIdx.size(), ds.X.cols());
    yTest.resize(testIdx.size());

    for (size_t i = 0; i < trainIdx.size(); i++) {
        xTrain.row(i) = ds.X.row(trainIdx[i]);
        yTrain(i) = ds.y(trainIdx[i]);
    }
    for (size_t i = 0; i < testIdx.size(); i++) {
        xTest.row(i) = ds.X.row(testIdx[i]);
        yTest(i) = ds.y(testIdx[i]);
    }
}

int main(int argc, char* argv[]) {
    std::cout << std::string(55, '=') << std::endl;
    std::cout << "  Linear Regression Fitter (Eigen + matplotplusplus)" << std::endl;
    std::cout << std::string(55, '=') << std::endl;

    DataSet ds;
    if (argc > 1) {
        try {
            ds = loadFromFile(argv[1]);
            std::cout << "Loaded dataset with " << ds.X.rows()
                      << " rows and " << (ds.X.cols() + 1) << " columns." << std::endl;
        } catch (const std::exception& e) {
            std::cerr << "Error loading file: " << e.what() << std::endl;
            return 1;
        }
    } else {
        std::cout << "\nNo CSV file provided. Using synthetic data.\n" << std::endl;
        ds = generateSyntheticData(200, 3, 10.0, 42);
    }

    MatrixXd xTrain, xTest;
    VectorXd yTrain, yTest;
    trainTestSplit(ds, 0.2, 42, xTrain, yTrain, xTest, yTest);

    std::cout << "Training samples: " << xTrain.rows() << std::endl;
    std::cout << "Test samples    : " << xTest.rows() << std::endl;

    LinearRegressionFitter lr;
    lr.fit(xTrain, yTrain, ds.featureNames);
    lr.printCoefficients();

    VectorXd yTrainPred = lr.predict(xTrain);
    auto trainMetrics = lr.computeMetrics(yTrain, yTrainPred);
    lr.printMetrics(trainMetrics, "Training");

    VectorXd yTestPred = lr.predict(xTest);
    auto testMetrics = lr.computeMetrics(yTest, yTestPred);
    lr.printMetrics(testMetrics, "Test");

    auto residualStats = lr.residualAnalysis(yTest, yTestPred);
    lr.printResidualAnalysis(residualStats);

    // Generate plots using matplotplusplus
    lr.plotResiduals(yTest, yTestPred, "residuals.png");
    lr.plotActualVsPredicted(yTest, yTestPred, "actual_vs_predicted.png");

    std::cout << std::endl << std::string(45, '=') << std::endl;
    std::cout << "  Prediction Example" << std::endl;
    std::cout << std::string(45, '=') << std::endl;
    int displayCount = std::min(5, static_cast<int>(xTest.rows()));
    for (int i = 0; i < displayCount; i++) {
        std::cout << std::fixed << std::setprecision(4);
        std::cout << "  Sample " << (i + 1) << ": actual=" << yTest(i)
                  << ", predicted=" << yTestPred(i) << std::endl;
    }
    std::cout << std::string(45, '=') << std::endl;

    std::cout << "\nJSON Output:" << std::endl;
    std::cout << lr.toJson(testMetrics, residualStats) << std::endl;

    return 0;
}
README.md
# Linear Regression Fitter (C++ - Eigen + matplotplusplus)

Fits OLS linear regression models using the Eigen linear algebra library with visualization via matplotplusplus.

## Dependencies

- **Eigen**: C++ template library for linear algebra (matrix operations, LDLT decomposition)
- **matplotplusplus**: C++ graphics library for MATLAB-like plotting and visualization

## Building

```bash
mkdir build && cd build
cmake ..
make
```

## Usage

```bash
# Run with synthetic data (generates plots automatically)
./linear_regression

# Run with a CSV file (last column is target)
./linear_regression data.csv
```

## Features

- OLS linear regression via Eigen's LDLT decomposition
- CSV data loading with manual parsing
- Synthetic data generation with configurable noise
- Train/test split for model evaluation
- Regression metrics: R-squared, MSE, RMSE
- Residual analysis: mean, std, min, max, median, skewness, kurtosis
- Residual plots (scatter + histogram) via matplotplusplus
- Actual vs Predicted scatter plot with perfect-prediction line
- JSON output of model parameters and metrics