CSV Statistical Analyzer (javascript, written by Gemini Code Assist)
envgap__gemini__javascript-t1-1
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
package.json- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
gemini/javascript-t1 #1 · read the task the agent was given
Gemini Code Assist wrote this javascript project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete JavaScript project for a clean Ubuntu 22.04 machine with only Node.js 20+ (LTS) installed. Include: - Source code - package.json with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
csv_analyzer.js
const fs = require('fs');
const csv = require('csv-parser');
const ss = require('simple-statistics');
const { format } = require('fast-csv');
const filePath = process.argv[2] || 'sample.csv';
async function generateSampleCsv(path) {
console.log(`Generating sample CSV: ${path}`);
const stream = format({ headers: true });
const writableStream = fs.createWriteStream(path);
stream.pipe(writableStream);
const depts = ['HR', 'Engineering', 'Sales', 'Marketing'];
for (let i = 1; i <= 200; i++) {
stream.write({
id: i,
age: Math.floor(Math.random() * 62) + 18,
salary: (50000 + (Math.random() - 0.5) * 20000).toFixed(2),
score: (Math.random() * 100).toFixed(2),
department: depts[Math.floor(Math.random() * depts.length)]
});
}
stream.end();
return new Promise(resolve => writableStream.on('finish', resolve));
}
async function analyzeCsv(path) {
const results = [];
fs.createReadStream(path)
.pipe(csv())
.on('data', (data) => results.push(data))
.on('end', () => {
if (results.length === 0) {
console.log("CSV is empty.");
return;
}
const headers = Object.keys(results[0]);
const report = { statistics: {} };
console.log(`${'COLUMN'.padEnd(20)} | ${'TYPE'.padEnd(10)} | ${'MEAN/MODE'.padEnd(15)} | ${'STD/FREQ'.padEnd(15)}`);
console.log('-'.repeat(70));
headers.forEach(header => {
const values = results.map(row => row[header]).filter(v => v !== '');
const numericValues = values.map(v => parseFloat(v)).filter(v => !isNaN(v));
if (numericValues.length > 0 && numericValues.length === values.length) {
// Numeric Analysis
const mean = ss.mean(numericValues);
const median = ss.median(numericValues);
const std = ss.standardDeviation(numericValues);
const q1 = ss.quantile(numericValues, 0.25);
const q3 = ss.quantile(numericValues, 0.75);
const iqr = q3 - q1;
const outliers = numericValues.filter(v => v < q1 - 1.5 * iqr || v > q3 + 1.5 * iqr);
report.statistics[header] = {
type: 'numeric',
mean,
median,
std,
min: ss.min(numericValues),
max: ss.max(numericValues),
outliers
};
console.log(`${header.padEnd(20)} | ${'Numeric'.padEnd(10)} | ${mean.toFixed(2).padEnd(15)} | ${std.toFixed(2).padEnd(15)}`);
} else {
// Categorical Analysis
const freqMap = {};
values.forEach(v => freqMap[v] = (freqMap[v] || 0) + 1);
const mode = Object.keys(freqMap).reduce((a, b) => freqMap[a] > freqMap[b] ? a : b, 'N/A');
const freq = freqMap[mode] || 0;
report.statistics[header] = {
type: 'categorical',
mode,
unique_count: Object.keys(freqMap).length,
top_10: Object.entries(freqMap).sort((a, b) => b[1] - a[1]).slice(0, 10)
};
console.log(`${header.padEnd(20)} | ${'Categorical'.padEnd(10)} | ${mode.padEnd(15)} | ${freq.toString().padEnd(15)}`);
}
});
fs.writeFileSync('report.json', JSON.stringify(report, null, 4));
console.log("Report saved to report.json");
});
}
(async () => {
if (!fs.existsSync(filePath)) {
await generateSampleCsv(filePath);
}
analyzeCsv(filePath);
})();
package.json
{
"name": "csv-statistical-analyzer",
"version": "1.0.0",
"description": "Comprehensive statistical analysis of CSV files",
"main": "csv_analyzer.js",
"engines": {
"node": ">=20.0.0"
},
"dependencies": {
"csv-parser": "3.0.0",
"simple-statistics": "7.8.3",
"fast-csv": "5.0.0"
},
"scripts": {
"start": "node csv_analyzer.js"
}
}
README.md
# CSV Statistical Analyzer (JavaScript) A comprehensive tool for statistical analysis of CSV data. ## Setup Instructions 1. Ensure Node.js 20+ (LTS) is installed. 2. Install dependencies: ```bash npm install ``` ## Run Commands - Run with a specific CSV: ```bash node csv_analyzer.js data.csv ``` - Run with generated sample data: ```bash node csv_analyzer.js ``` ## Output - Console: Formatted summary table. - File: `report.json` containing detailed statistics and outliers.