← All tasks
javascriptgemini/javascript-t1 #1Not a task: already works

CSV Statistical Analyzer (javascript, written by Gemini Code Assist)

envgap__gemini__javascript-t1-1

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
package.json
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

gemini/javascript-t1 #1 · read the task the agent was given
Gemini Code Assist wrote this javascript project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: CSV Statistical Analyzer

Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Auto-detect which columns are numeric vs categorical
- For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count
- Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column
- For each categorical column compute: unique count, most frequent value, and top 10 value frequencies
- Print a formatted summary table to the console with aligned columns
- Save the complete analysis to report.json including all stats, outlier details, and column type classifications
- If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it
- Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas

Create a complete JavaScript project for a clean Ubuntu 22.04 machine with only Node.js 20+ (LTS) installed. Include:
- Source code
- package.json with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

csv_analyzer.js
const fs = require('fs');
const csv = require('csv-parser');
const ss = require('simple-statistics');
const { format } = require('fast-csv');

const filePath = process.argv[2] || 'sample.csv';

async function generateSampleCsv(path) {
    console.log(`Generating sample CSV: ${path}`);
    const stream = format({ headers: true });
    const writableStream = fs.createWriteStream(path);
    stream.pipe(writableStream);

    const depts = ['HR', 'Engineering', 'Sales', 'Marketing'];
    for (let i = 1; i <= 200; i++) {
        stream.write({
            id: i,
            age: Math.floor(Math.random() * 62) + 18,
            salary: (50000 + (Math.random() - 0.5) * 20000).toFixed(2),
            score: (Math.random() * 100).toFixed(2),
            department: depts[Math.floor(Math.random() * depts.length)]
        });
    }
    stream.end();
    return new Promise(resolve => writableStream.on('finish', resolve));
}

async function analyzeCsv(path) {
    const results = [];
    fs.createReadStream(path)
        .pipe(csv())
        .on('data', (data) => results.push(data))
        .on('end', () => {
            if (results.length === 0) {
                console.log("CSV is empty.");
                return;
            }

            const headers = Object.keys(results[0]);
            const report = { statistics: {} };

            console.log(`${'COLUMN'.padEnd(20)} | ${'TYPE'.padEnd(10)} | ${'MEAN/MODE'.padEnd(15)} | ${'STD/FREQ'.padEnd(15)}`);
            console.log('-'.repeat(70));

            headers.forEach(header => {
                const values = results.map(row => row[header]).filter(v => v !== '');
                const numericValues = values.map(v => parseFloat(v)).filter(v => !isNaN(v));

                if (numericValues.length > 0 && numericValues.length === values.length) {
                    // Numeric Analysis
                    const mean = ss.mean(numericValues);
                    const median = ss.median(numericValues);
                    const std = ss.standardDeviation(numericValues);
                    const q1 = ss.quantile(numericValues, 0.25);
                    const q3 = ss.quantile(numericValues, 0.75);
                    const iqr = q3 - q1;
                    const outliers = numericValues.filter(v => v < q1 - 1.5 * iqr || v > q3 + 1.5 * iqr);

                    report.statistics[header] = {
                        type: 'numeric',
                        mean,
                        median,
                        std,
                        min: ss.min(numericValues),
                        max: ss.max(numericValues),
                        outliers
                    };

                    console.log(`${header.padEnd(20)} | ${'Numeric'.padEnd(10)} | ${mean.toFixed(2).padEnd(15)} | ${std.toFixed(2).padEnd(15)}`);
                } else {
                    // Categorical Analysis
                    const freqMap = {};
                    values.forEach(v => freqMap[v] = (freqMap[v] || 0) + 1);
                    
                    const mode = Object.keys(freqMap).reduce((a, b) => freqMap[a] > freqMap[b] ? a : b, 'N/A');
                    const freq = freqMap[mode] || 0;

                    report.statistics[header] = {
                        type: 'categorical',
                        mode,
                        unique_count: Object.keys(freqMap).length,
                        top_10: Object.entries(freqMap).sort((a, b) => b[1] - a[1]).slice(0, 10)
                    };

                    console.log(`${header.padEnd(20)} | ${'Categorical'.padEnd(10)} | ${mode.padEnd(15)} | ${freq.toString().padEnd(15)}`);
                }
            });

            fs.writeFileSync('report.json', JSON.stringify(report, null, 4));
            console.log("Report saved to report.json");
        });
}

(async () => {
    if (!fs.existsSync(filePath)) {
        await generateSampleCsv(filePath);
    }
    analyzeCsv(filePath);
})();
package.json
{
  "name": "csv-statistical-analyzer",
  "version": "1.0.0",
  "description": "Comprehensive statistical analysis of CSV files",
  "main": "csv_analyzer.js",
  "engines": {
    "node": ">=20.0.0"
  },
  "dependencies": {
    "csv-parser": "3.0.0",
    "simple-statistics": "7.8.3",
    "fast-csv": "5.0.0"
  },
  "scripts": {
    "start": "node csv_analyzer.js"
  }
}
README.md
# CSV Statistical Analyzer (JavaScript)

A comprehensive tool for statistical analysis of CSV data.

## Setup Instructions

1. Ensure Node.js 20+ (LTS) is installed.
2. Install dependencies:
   ```bash
   npm install
   ```

## Run Commands

- Run with a specific CSV:
  ```bash
  node csv_analyzer.js data.csv
  ```
- Run with generated sample data:
  ```bash
  node csv_analyzer.js
  ```

## Output
- Console: Formatted summary table.
- File: `report.json` containing detailed statistics and outliers.