CSV Statistical Analyzer (java, written by Gemini Code Assist)
envgap__gemini__java-t1-1
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
Unclosed string literal (broken newlines in printf)
Not a benchmark task.
- Its repair changed source code, so it is not an environment task.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
gemini/java-t1 #1 · read the task the agent was given
Gemini Code Assist wrote this java project from the task below. It does not run on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>com.csvanalyzer</groupId>
<artifactId>csv-statistical-analyzer</artifactId>
<version>1.0-SNAPSHOT</version>
<properties>
<maven.compiler.source>17</maven.compiler.source>
<maven.compiler.target>17</maven.compiler.target>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
</properties>
<dependencies>
<!-- OpenCSV for parsing -->
<dependency>
<groupId>com.opencsv</groupId>
<artifactId>opencsv</artifactId>
<version>5.9</version>
</dependency>
<!-- Apache Commons Math for statistics -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-math3</artifactId>
<version>3.6.1</version>
</dependency>
<!-- Jackson for JSON output -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-databind</artifactId>
<version>2.16.1</version>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-shade-plugin</artifactId>
<version>3.5.1</version>
<executions>
<execution>
<phase>package</phase>
<goals>
<goal>shade</goal>
</goals>
<configuration>
<transformers>
<transformer implementation="org.apache.maven.plugins.shade.resource.ManifestResourceTransformer">
<mainClass>com.csvanalyzer.CsvStatisticalAnalyzer</mainClass>
</transformer>
</transformers>
</configuration>
</execution>
</executions>
</plugin>
</plugins>
</build>
</project>
README.md
# CSV Statistical Analyzer (Java) A comprehensive tool for statistical analysis of CSV data. ## Setup Instructions 1. Ensure JDK 17+ and Maven are installed. 2. Build the project: ```bash mvn clean package ``` ## Run Commands - Run with a specific CSV: ```bash java -jar target/csv-statistical-analyzer-1.0-SNAPSHOT.jar data.csv ``` - Run with generated sample data: ```bash java -jar target/csv-statistical-analyzer-1.0-SNAPSHOT.jar ``` ## Output - Console: Formatted summary table. - File: `report.json` containing detailed statistics and outliers.
src/main/java/com/csvanalyzer/CsvStatisticalAnalyzer.java
package com.csvanalyzer;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ArrayNode;
import com.fasterxml.jackson.databind.node.ObjectNode;
import com.opencsv.CSVReader;
import com.opencsv.CSVWriter;
import org.apache.commons.math3.stat.descriptive.DescriptiveStatistics;
import java.io.File;
import java.io.FileReader;
import java.io.FileWriter;
import java.util.*;
public class CsvStatisticalAnalyzer {
public static void main(String[] args) {
String filePath = args.length > 0 ? args[0] : "sample.csv";
File file = new File(filePath);
if (!file.exists()) {
generateSampleCsv(filePath);
}
analyzeCsv(filePath);
}
private static void generateSampleCsv(String filePath) {
System.out.println("Generating sample CSV: " + filePath);
try (CSVWriter writer = new CSVWriter(new FileWriter(filePath))) {
writer.writeNext(new String[]{"id", "age", "salary", "score", "department"});
Random rand = new Random(42);
String[] departments = {"HR", "Engineering", "Sales", "Marketing"};
for (int i = 1; i <= 200; i++) {
writer.writeNext(new String[]{
String.valueOf(i),
String.valueOf(18 + rand.nextInt(62)),
String.valueOf(30000 + rand.nextGaussian() * 10000),
String.valueOf(rand.nextDouble() * 100),
departments[rand.nextInt(departments.length)]
});
}
} catch (Exception e) {
e.printStackTrace();
}
}
private static void analyzeCsv(String filePath) {
ObjectMapper mapper = new ObjectMapper();
ObjectNode report = mapper.createObjectNode();
ObjectNode statisticsNode = report.putObject("statistics");
try (CSVReader reader = new CSVReader(new FileReader(filePath))) {
String[] header = reader.readNext();
if (header == null) return;
List<String[]> rows = reader.readAll();
Map<String, List<String>> columnData = new HashMap<>();
for (String h : header) columnData.put(h, new ArrayList<>());
for (String[] row : rows) {
for (int i = 0; i < header.length; i++) {
if (i < row.length) columnData.get(header[i]).add(row[i]);
}
}
System.out.printf("%-20s | %-10s | %-15s | %-15s
", "COLUMN", "TYPE", "MEAN/MODE", "STD/FREQ");
System.out.println("-".repeat(70));
for (String col : header) {
List<String> values = columnData.get(col);
boolean isNumeric = isColumnNumeric(values);
if (isNumeric) {
DescriptiveStatistics stats = new DescriptiveStatistics();
for (String v : values) {
try {
if (!v.isEmpty()) stats.addValue(Double.parseDouble(v));
} catch (NumberFormatException ignored) {}
}
ObjectNode colStats = statisticsNode.putObject(col);
colStats.put("type", "numeric");
colStats.put("mean", stats.getMean());
colStats.put("median", stats.getPercentile(50));
colStats.put("std", stats.getStandardDeviation());
colStats.put("count", stats.getN());
double q1 = stats.getPercentile(25);
double q3 = stats.getPercentile(75);
double iqr = q3 - q1;
ArrayNode outliers = colStats.putArray("outliers");
for (double v : stats.getValues()) {
if (v < q1 - 1.5 * iqr || v > q3 + 1.5 * iqr) outliers.add(v);
}
System.out.printf("%-20s | %-10s | %-15.2f | %-15.2f
", col, "Numeric", stats.getMean(), stats.getStandardDeviation());
} else {
Map<String, Integer> freqMap = new HashMap<>();
for (String v : values) freqMap.put(v, freqMap.getOrDefault(v, 0) + 1);
String mode = freqMap.entrySet().stream().max(Map.Entry.comparingByValue()).map(Map.Entry::getKey).orElse("N/A");
int freq = freqMap.getOrDefault(mode, 0);
ObjectNode colStats = statisticsNode.putObject(col);
colStats.put("type", "categorical");
colStats.put("mode", mode);
colStats.put("unique_count", freqMap.size());
System.out.printf("%-20s | %-10s | %-15s | %-15d
", col, "Categorical", mode, freq);
}
}
mapper.writerWithDefaultPrettyPrinter().writeValue(new File("report.json"), report);
System.out.println("Report saved to report.json");
} catch (Exception e) {
e.printStackTrace();
}
}
private static boolean isColumnNumeric(List<String> values) {
int count = 0;
for (String v : values) {
if (v == null || v.isEmpty()) continue;
try {
Double.parseDouble(v);
count++;
} catch (NumberFormatException e) {
return false;
}
}
return count > 0;
}
}