CSV Statistical Analyzer (java, written by Claude Code)
envgap__claude-code__java-t1-1
Written by a coding agent; not on GitHubWritten 2026-02-27
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
claude-code/java-t1 #1 · read the task the agent was given
Claude Code wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
4 files, exactly as written, before any repair.
generation_metadata.json
{
"project": "p_01",
"language": "java",
"trial": 1,
"agent": "claude",
"model": "sonnet",
"elapsed_seconds": 307.76,
"exit_code": 0,
"files_created": [
"pom.xml",
"README.md"
],
"timestamp": "2026-02-27T04:40:14Z",
"prompt_file": "java_prompts/p_01.md"
}pom.xml
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0
http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>com.csvanalyzer</groupId>
<artifactId>csv-statistical-analyzer</artifactId>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>CSV Statistical Analyzer</name>
<properties>
<maven.compiler.source>17</maven.compiler.source>
<maven.compiler.target>17</maven.compiler.target>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
</properties>
<dependencies>
<!-- CSV parsing: handles quoted fields, missing values, malformed rows -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-csv</artifactId>
<version>1.10.0</version>
</dependency>
<!-- JSON serialization for report.json output -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-databind</artifactId>
<version>2.15.2</version>
</dependency>
<!-- Transitive dependency of jackson-databind -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-annotations</artifactId>
<version>2.15.2</version>
</dependency>
<!-- Transitive dependency of jackson-databind -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-core</artifactId>
<version>2.15.2</version>
</dependency>
</dependencies>
<build>
<plugins>
<!-- Compiler plugin pinned to exact version -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.11.0</version>
<configuration>
<source>17</source>
<target>17</target>
</configuration>
</plugin>
<!-- Assembly plugin to build a fat JAR with all dependencies -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-assembly-plugin</artifactId>
<version>3.6.0</version>
<configuration>
<archive>
<manifest>
<mainClass>com.csvanalyzer.CsvStatisticalAnalyzer</mainClass>
</manifest>
</archive>
<descriptorRefs>
<descriptorRef>jar-with-dependencies</descriptorRef>
</descriptorRefs>
<finalName>csv-analyzer</finalName>
<appendAssemblyId>false</appendAssemblyId>
</configuration>
<executions>
<execution>
<id>make-assembly</id>
<phase>package</phase>
<goals>
<goal>single</goal>
</goals>
</execution>
</executions>
</plugin>
</plugins>
</build>
</project>
README.md
# CSV Statistical Analyzer
A Java command-line tool that reads a CSV file, auto-detects column types,
computes comprehensive statistics for every column, detects outliers via the
IQR method, prints a formatted console report, and saves a complete
`report.json`.
---
## Requirements
| Component | Minimum version |
|-----------|----------------|
| Ubuntu | 22.04 LTS |
| JDK | 17 |
| Maven | 3.6 |
Check your installed versions:
```bash
java -version
mvn -version
```
If Maven is not installed:
```bash
sudo apt-get update && sudo apt-get install -y maven
```
---
## Project Layout
```
.
├── pom.xml
├── README.md
└── src/
└── main/
└── java/
└── com/
└── csvanalyzer/
└── CsvStatisticalAnalyzer.java
```
All build outputs land in `target/`.
---
## Dependencies
All versions are pinned in `pom.xml`.
| Artifact | Version | Purpose |
|----------|---------|---------|
| `org.apache.commons:commons-csv` | 1.10.0 | Robust CSV parsing — handles quoted fields that contain commas, optional headers, missing cells, and malformed rows |
| `com.fasterxml.jackson.core:jackson-databind` | 2.15.2 | Serialises the analysis result to pretty-printed JSON (`report.json`) |
| `com.fasterxml.jackson.core:jackson-annotations` | 2.15.2 | Transitive dependency of `jackson-databind` |
| `com.fasterxml.jackson.core:jackson-core` | 2.15.2 | Transitive dependency of `jackson-databind`; low-level JSON streaming |
The `maven-assembly-plugin` (3.6.0) bundles all dependencies into a single
fat JAR (`target/csv-analyzer.jar`) for easy deployment.
---
## Build Steps
```bash
# 1. Clone / navigate to the project root (directory containing pom.xml)
cd /path/to/project
# 2. Compile and package (creates target/csv-analyzer.jar)
mvn clean package
# If Maven needs to download dependencies for the first time this may take
# a minute on a fresh machine.
```
A successful build ends with:
```
[INFO] BUILD SUCCESS
```
---
## Run Commands
### Analyze your own CSV file
```bash
java -jar target/csv-analyzer.jar /path/to/your/data.csv
```
### Generate a built-in sample (200 rows, 5 numeric + 2 categorical columns)
```bash
java -jar target/csv-analyzer.jar
```
This writes `sample_data.csv` in the current directory, analyzes it, and
saves `report.json`.
---
## Expected Output
### Console
```
==========================================================================================
CSV STATISTICAL ANALYSIS REPORT
==========================================================================================
File : sample_data.csv
Timestamp : 2024-06-01T12:34:56.123
Rows : 202
Columns : 7 (5 numeric, 2 categorical)
NUMERIC COLUMNS — SUMMARY
------------------------------------------------------------------------------------------
Column | Count | Mean | Median | Std Dev | Min | Max | Missing | Outliers
-----------|-------|------------|------------|------------|------------|------------|---------|--------
age | 190 | 38.74 | 39.00 | 13.80 | 18.00 | 130.00 | 12 | 1
salary | 197 | 115432.81 | 115263.47 | 49012.33 | 30021.55 | 4500000.00 | 5 | 1
score | 194 | 50.43 | 51.20 | 28.97 | 0.10 | 155.00 | 8 | 1
height_cm | 200 | 174.82 | 174.91 | 14.51 | 150.03 | 199.97 | 2 | 0
weight_kg | 189 | 84.99 | 84.73 | 20.17 | 50.01 | 310.00 | 13 | 1
NUMERIC COLUMNS — PERCENTILES & IQR FENCES
------------------------------------------------------------------------------------------
Column | P25 | P50 (Med) | P75 | IQR | Lower Fence | Upper Fence
...
OUTLIER DETAILS (IQR method, first 20 shown per column)
------------------------------------------------------------------------------------------
age (1 outlier ): 130.00
salary (1 outlier ): 4500000.00
score (1 outlier ): 155.00
weight_kg (1 outlier ): 310.00
CATEGORICAL COLUMNS — SUMMARY
------------------------------------------------------------------------------------------
Column | Count | Unique | Most Frequent | Freq. | Missing
-----------|-------|--------|---------------|--------|--------
department | 202 | 5 | Engineering | 46 | 0
city | 202 | 9 | Tokyo | 32 | 0
CATEGORICAL COLUMNS — TOP VALUE FREQUENCIES
------------------------------------------------------------------------------------------
department:
Engineering 46
Sales 42
Finance 40
Marketing 38
HR 36
city:
Tokyo 32
...
==========================================================================================
Full JSON report saved to: report.json
```
### report.json (excerpt)
```json
{
"analysisTimestamp" : "2024-06-01T12:34:56.123",
"inputFile" : "sample_data.csv",
"totalRows" : 202,
"totalColumns" : 7,
"numericColumns" : 5,
"categoricalColumns" : 2,
"columns" : {
"age" : {
"type" : "NUMERIC",
"count" : 190,
"missing" : 12,
"mean" : 38.74,
"median" : 39.0,
"stdDev" : 13.8,
"variance" : 190.44,
"min" : 18.0,
"max" : 130.0,
"p25" : 27.0,
"p50" : 39.0,
"p75" : 51.0,
"iqr" : 24.0,
"lowerFence" : -9.0,
"upperFence" : 87.0,
"outlierCount" : 1,
"outliers" : [ 130.0 ]
},
"department" : {
"type" : "CATEGORICAL",
"count" : 202,
"missing" : 0,
"uniqueCount" : 5,
"mostFrequent" : "Engineering",
"mostFrequentCount" : 46,
"topFrequencies" : {
"Engineering" : 46,
"Sales" : 42,
"Finance" : 40,
"Marketing" : 38,
"HR" : 36
}
}
}
}
```
---
## Edge Cases Handled
| Scenario | Behaviour |
|----------|-----------|
| No CLI argument | Generates `sample_data.csv` (200 rows) then analyzes it |
| Empty file | Prints a warning and exits cleanly |
| Header-only file | Reports 0 rows; all stats are N/A |
| All values missing in a column | Stats show `null` in JSON, `N/A` in console |
| Single data row | Variance / std dev reported as `0` |
| Quoted fields containing commas | Parsed correctly by Apache Commons CSV |
| Mixed types in a column (≥ 30 % non-numeric) | Column treated as categorical |
| Malformed / short rows | Missing cells treated as empty strings |
---
## How Column Type Detection Works
For each column the analyzer collects all non-empty cell values and attempts
to parse each as a `double`. If **70 % or more** parse successfully the
column is classified as **NUMERIC**; otherwise it is **CATEGORICAL**.
This threshold tolerates occasional stray text in otherwise numeric columns
(e.g. `"N/A"`, `"?"`) while still catching genuinely mixed-type columns.
---
## Outlier Detection (IQR Method)
For every numeric column:
1. Compute **Q1** (25th percentile) and **Q3** (75th percentile).
2. **IQR** = Q3 − Q1
3. **Lower fence** = Q1 − 1.5 × IQR
4. **Upper fence** = Q3 + 1.5 × IQR
5. Any value outside `[lower fence, upper fence]` is flagged as an outlier.
The full list of outlier values is included in `report.json`; the console
shows the first 20 per column.
src/main/java/com/csvanalyzer/CsvStatisticalAnalyzer.java
package com.csvanalyzer;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.SerializationFeature;
import org.apache.commons.csv.CSVFormat;
import org.apache.commons.csv.CSVParser;
import org.apache.commons.csv.CSVPrinter;
import org.apache.commons.csv.CSVRecord;
import java.io.*;
import java.nio.charset.StandardCharsets;
import java.time.LocalDateTime;
import java.time.format.DateTimeFormatter;
import java.util.*;
import java.util.stream.Collectors;
/**
* CSV Statistical Analyzer
*
* Reads a CSV file, auto-detects numeric vs categorical columns, computes
* comprehensive statistics, detects outliers via IQR, prints a formatted
* console report, and writes a complete report.json.
*
* Usage:
* java -jar csv-analyzer.jar [path/to/file.csv]
*
* If no argument is given, a sample 200-row CSV is generated first.
*/
public class CsvStatisticalAnalyzer {
private static final String SAMPLE_CSV = "sample_data.csv";
private static final String OUTPUT_JSON = "report.json";
/** A column is treated as numeric if at least this fraction of its
* non-empty values parse successfully as a double. */
private static final double NUMERIC_THRESHOLD = 0.70;
/** IQR multiplier for outlier fences (Tukey's method). */
private static final double IQR_MULTIPLIER = 1.5;
// =========================================================================
// Data classes
// =========================================================================
public static class NumericStats {
public String columnName;
public int count;
public int missing;
public double mean;
public double median;
public double stdDev;
public double variance;
public double min;
public double max;
public double p25;
public double p50;
public double p75;
public double iqr;
public double lowerFence;
public double upperFence;
public int outlierCount;
public List<Double> outliers;
}
public static class CategoricalStats {
public String columnName;
public int count;
public int missing;
public int uniqueCount;
public String mostFrequent;
public int mostFrequentCount;
public Map<String, Integer> topFrequencies;
}
public static class AnalysisResult {
public String analysisTimestamp;
public String inputFile;
public int totalRows;
public int totalColumns;
public int numericColumnCount;
public int categoricalColumnCount;
public List<NumericStats> numericStats;
public List<CategoricalStats> categoricalStats;
}
// =========================================================================
// Entry point
// =========================================================================
public static void main(String[] args) throws Exception {
String csvFilePath;
if (args.length == 0) {
System.out.println("No input file provided. Generating sample CSV with 200 rows...");
csvFilePath = generateSampleCsv();
System.out.println("Sample CSV generated: " + csvFilePath + "\n");
} else {
csvFilePath = args[0];
}
File csvFile = new File(csvFilePath);
if (!csvFile.exists()) {
System.err.println("Error: File not found: " + csvFilePath);
System.exit(1);
}
System.out.println("Analyzing: " + csvFilePath);
AnalysisResult result = analyzeCsv(csvFilePath);
if (result == null) {
System.err.println("Analysis produced no result (file may be empty).");
System.exit(1);
}
printConsoleReport(result);
saveJsonReport(result);
System.out.println("Full JSON report saved to: " + OUTPUT_JSON);
}
// =========================================================================
// CSV parsing and analysis
// =========================================================================
public static AnalysisResult analyzeCsv(String filePath) throws Exception {
File file = new File(filePath);
// Empty file guard
if (file.length() == 0) {
System.err.println("Warning: File is empty — nothing to analyze.");
return null;
}
// --- Parse the file ---------------------------------------------------
List<String> headers = new ArrayList<>();
Map<String, List<String>> rawData = new LinkedHashMap<>();
int totalRows;
CSVFormat fmt = CSVFormat.DEFAULT.builder()
.setHeader()
.setSkipHeaderRecord(true)
.setIgnoreEmptyLines(true)
.setTrim(true)
.build();
try (Reader reader = new InputStreamReader(
new FileInputStream(file), StandardCharsets.UTF_8);
CSVParser parser = new CSVParser(reader, fmt)) {
headers.addAll(parser.getHeaderNames());
if (headers.isEmpty()) {
System.err.println("Warning: No column headers found.");
return null;
}
for (String h : headers) {
rawData.put(h, new ArrayList<>());
}
int rowIdx = 0;
for (CSVRecord record : parser) {
rowIdx++;
for (String h : headers) {
String val = "";
try {
val = record.get(h);
} catch (Exception ignored) {
// Malformed / short row — treat cell as missing
}
rawData.get(h).add(val == null ? "" : val.trim());
}
}
totalRows = rowIdx;
}
if (totalRows == 0) {
System.err.println("Warning: Header-only file — no data rows found.");
}
// --- Classify and analyse each column ---------------------------------
AnalysisResult result = new AnalysisResult();
result.analysisTimestamp = LocalDateTime.now()
.format(DateTimeFormatter.ISO_LOCAL_DATE_TIME);
result.inputFile = filePath;
result.totalRows = totalRows;
result.totalColumns = headers.size();
result.numericStats = new ArrayList<>();
result.categoricalStats = new ArrayList<>();
for (String header : headers) {
List<String> values = rawData.get(header);
List<String> nonEmpty = values.stream()
.filter(v -> !v.isEmpty())
.collect(Collectors.toList());
long numericCount = nonEmpty.stream()
.filter(CsvStatisticalAnalyzer::isNumeric)
.count();
boolean isNumeric = !nonEmpty.isEmpty()
&& (double) numericCount / nonEmpty.size() >= NUMERIC_THRESHOLD;
if (isNumeric) {
result.numericStats.add(computeNumericStats(header, values));
} else {
result.categoricalStats.add(computeCategoricalStats(header, values));
}
}
result.numericColumnCount = result.numericStats.size();
result.categoricalColumnCount = result.categoricalStats.size();
return result;
}
private static boolean isNumeric(String s) {
if (s == null || s.isEmpty()) return false;
try {
Double.parseDouble(s);
return true;
} catch (NumberFormatException e) {
return false;
}
}
// =========================================================================
// Numeric statistics
// =========================================================================
private static NumericStats computeNumericStats(String name, List<String> raw) {
List<Double> values = raw.stream()
.filter(v -> v != null && !v.isEmpty() && isNumeric(v))
.map(Double::parseDouble)
.collect(Collectors.toList());
NumericStats s = new NumericStats();
s.columnName = name;
s.count = values.size();
s.missing = raw.size() - values.size();
if (values.isEmpty()) {
// All values missing — fill with NaN sentinels
s.mean = s.median = s.stdDev = s.variance = Double.NaN;
s.min = s.max = Double.NaN;
s.p25 = s.p50 = s.p75 = s.iqr = Double.NaN;
s.lowerFence = s.upperFence = Double.NaN;
s.outlierCount = 0;
s.outliers = Collections.emptyList();
return s;
}
List<Double> sorted = new ArrayList<>(values);
Collections.sort(sorted);
int n = sorted.size();
// Central tendency
double sum = values.stream().mapToDouble(Double::doubleValue).sum();
s.mean = sum / n;
s.p25 = percentile(sorted, 25.0);
s.p50 = percentile(sorted, 50.0);
s.p75 = percentile(sorted, 75.0);
s.median = s.p50;
s.min = sorted.get(0);
s.max = sorted.get(n - 1);
s.iqr = s.p75 - s.p25;
// Dispersion (sample variance, n-1)
if (n > 1) {
double sumSqDiff = values.stream()
.mapToDouble(v -> (v - s.mean) * (v - s.mean))
.sum();
s.variance = sumSqDiff / (n - 1);
s.stdDev = Math.sqrt(s.variance);
} else {
s.variance = 0.0;
s.stdDev = 0.0;
}
// IQR outlier fences
s.lowerFence = s.p25 - IQR_MULTIPLIER * s.iqr;
s.upperFence = s.p75 + IQR_MULTIPLIER * s.iqr;
final double lo = s.lowerFence;
final double hi = s.upperFence;
s.outliers = values.stream()
.filter(v -> v < lo || v > hi)
.sorted()
.collect(Collectors.toList());
s.outlierCount = s.outliers.size();
return s;
}
/** Linear-interpolation percentile on a sorted list. */
private static double percentile(List<Double> sorted, double p) {
if (sorted.isEmpty()) return Double.NaN;
if (sorted.size() == 1) return sorted.get(0);
double idx = (p / 100.0) * (sorted.size() - 1);
int lower = (int) Math.floor(idx);
int upper = (int) Math.ceil(idx);
if (lower == upper) return sorted.get(lower);
double frac = idx - lower;
return sorted.get(lower) * (1.0 - frac) + sorted.get(upper) * frac;
}
// =========================================================================
// Categorical statistics
// =========================================================================
private static CategoricalStats computeCategoricalStats(
String name, List<String> raw) {
CategoricalStats s = new CategoricalStats();
s.columnName = name;
s.missing = (int) raw.stream()
.filter(v -> v == null || v.isEmpty()).count();
s.count = raw.size() - s.missing;
Map<String, Integer> freq = new LinkedHashMap<>();
for (String v : raw) {
if (v != null && !v.isEmpty()) {
freq.merge(v, 1, Integer::sum);
}
}
s.uniqueCount = freq.size();
// Sort by frequency desc, then alphabetically for stable output
List<Map.Entry<String, Integer>> ranked = freq.entrySet().stream()
.sorted(Map.Entry.<String, Integer>comparingByValue().reversed()
.thenComparing(Map.Entry.comparingByKey()))
.collect(Collectors.toList());
if (!ranked.isEmpty()) {
s.mostFrequent = ranked.get(0).getKey();
s.mostFrequentCount = ranked.get(0).getValue();
} else {
s.mostFrequent = null;
s.mostFrequentCount = 0;
}
s.topFrequencies = new LinkedHashMap<>();
ranked.stream().limit(10)
.forEach(e -> s.topFrequencies.put(e.getKey(), e.getValue()));
return s;
}
// =========================================================================
// Console report
// =========================================================================
private static void printConsoleReport(AnalysisResult r) {
final String SEP = "=".repeat(90);
final String sep = "-".repeat(90);
System.out.println(SEP);
System.out.println(" CSV STATISTICAL ANALYSIS REPORT");
System.out.println(SEP);
System.out.printf(" File : %s%n", r.inputFile);
System.out.printf(" Timestamp : %s%n", r.analysisTimestamp);
System.out.printf(" Rows : %d%n", r.totalRows);
System.out.printf(" Columns : %d (%d numeric, %d categorical)%n",
r.totalColumns, r.numericColumnCount, r.categoricalColumnCount);
System.out.println();
// --- Numeric summary table -------------------------------------------
if (!r.numericStats.isEmpty()) {
int nw = r.numericStats.stream()
.mapToInt(s -> s.columnName.length()).max().orElse(6);
nw = Math.max(nw, 6);
System.out.println("NUMERIC COLUMNS — SUMMARY");
System.out.println(sep);
String h1 = String.format(
"%-" + nw + "s | %6s | %10s | %10s | %10s | %10s | %10s | %7s | %8s",
"Column", "Count", "Mean", "Median", "Std Dev",
"Min", "Max", "Missing", "Outliers");
System.out.println(h1);
System.out.println("-".repeat(h1.length()));
for (NumericStats s : r.numericStats) {
System.out.printf(
"%-" + nw + "s | %6d | %10s | %10s | %10s | %10s | %10s | %7d | %8d%n",
s.columnName, s.count,
fmt(s.mean), fmt(s.median), fmt(s.stdDev),
fmt(s.min), fmt(s.max),
s.missing, s.outlierCount);
}
System.out.println();
// Percentile table
System.out.println("NUMERIC COLUMNS — PERCENTILES & IQR FENCES");
System.out.println(sep);
String h2 = String.format(
"%-" + nw + "s | %10s | %10s | %10s | %10s | %12s | %12s",
"Column", "P25", "P50 (Med)", "P75", "IQR",
"Lower Fence", "Upper Fence");
System.out.println(h2);
System.out.println("-".repeat(h2.length()));
for (NumericStats s : r.numericStats) {
System.out.printf(
"%-" + nw + "s | %10s | %10s | %10s | %10s | %12s | %12s%n",
s.columnName,
fmt(s.p25), fmt(s.p50), fmt(s.p75), fmt(s.iqr),
fmt(s.lowerFence), fmt(s.upperFence));
}
System.out.println();
// Outlier details
boolean anyOutliers = r.numericStats.stream()
.anyMatch(s -> s.outlierCount > 0);
if (anyOutliers) {
System.out.println("OUTLIER DETAILS (IQR method, first 20 shown per column)");
System.out.println(sep);
for (NumericStats s : r.numericStats) {
if (s.outlierCount > 0) {
String vals = s.outliers.stream()
.limit(20)
.map(CsvStatisticalAnalyzer::fmt)
.collect(Collectors.joining(", "));
if (s.outlierCount > 20) vals += ", ...";
System.out.printf(" %-" + nw + "s (%d outlier%s): %s%n",
s.columnName, s.outlierCount,
s.outlierCount == 1 ? "" : "s", vals);
}
}
System.out.println();
}
}
// --- Categorical summary table ----------------------------------------
if (!r.categoricalStats.isEmpty()) {
int nw = r.categoricalStats.stream()
.mapToInt(s -> s.columnName.length()).max().orElse(6);
nw = Math.max(nw, 6);
int mfw = r.categoricalStats.stream()
.mapToInt(s -> s.mostFrequent != null ? s.mostFrequent.length() : 3)
.max().orElse(13);
mfw = Math.max(mfw, 13);
System.out.println("CATEGORICAL COLUMNS — SUMMARY");
System.out.println(sep);
String h3 = String.format(
"%-" + nw + "s | %6s | %7s | %-" + mfw + "s | %5s | %7s",
"Column", "Count", "Unique", "Most Frequent", "Freq.", "Missing");
System.out.println(h3);
System.out.println("-".repeat(h3.length()));
for (CategoricalStats s : r.categoricalStats) {
System.out.printf(
"%-" + nw + "s | %6d | %7d | %-" + mfw + "s | %5d | %7d%n",
s.columnName, s.count, s.uniqueCount,
s.mostFrequent != null ? s.mostFrequent : "N/A",
s.mostFrequentCount, s.missing);
}
System.out.println();
System.out.println("CATEGORICAL COLUMNS — TOP VALUE FREQUENCIES");
System.out.println(sep);
for (CategoricalStats s : r.categoricalStats) {
System.out.printf(" %s:%n", s.columnName);
s.topFrequencies.forEach((k, v) ->
System.out.printf(" %-25s %d%n", k, v));
System.out.println();
}
}
System.out.println(SEP);
System.out.println();
}
/** Format a double for console display (2 dp, or "N/A" for NaN/Inf). */
private static String fmt(double v) {
if (Double.isNaN(v)) return "N/A";
if (Double.isInfinite(v)) return v > 0 ? "+Inf" : "-Inf";
return String.format("%.2f", v);
}
// =========================================================================
// JSON report
// =========================================================================
private static void saveJsonReport(AnalysisResult result) throws Exception {
ObjectMapper mapper = new ObjectMapper();
mapper.enable(SerializationFeature.INDENT_OUTPUT);
Map<String, Object> report = new LinkedHashMap<>();
report.put("analysisTimestamp", result.analysisTimestamp);
report.put("inputFile", result.inputFile);
report.put("totalRows", result.totalRows);
report.put("totalColumns", result.totalColumns);
report.put("numericColumns", result.numericColumnCount);
report.put("categoricalColumns", result.categoricalColumnCount);
Map<String, Object> columns = new LinkedHashMap<>();
for (NumericStats s : result.numericStats) {
Map<String, Object> col = new LinkedHashMap<>();
col.put("type", "NUMERIC");
col.put("count", s.count);
col.put("missing", s.missing);
col.put("mean", safe(s.mean));
col.put("median", safe(s.median));
col.put("stdDev", safe(s.stdDev));
col.put("variance", safe(s.variance));
col.put("min", safe(s.min));
col.put("max", safe(s.max));
col.put("p25", safe(s.p25));
col.put("p50", safe(s.p50));
col.put("p75", safe(s.p75));
col.put("iqr", safe(s.iqr));
col.put("lowerFence", safe(s.lowerFence));
col.put("upperFence", safe(s.upperFence));
col.put("outlierCount", s.outlierCount);
col.put("outliers", s.outliers.stream()
.map(CsvStatisticalAnalyzer::safe)
.collect(Collectors.toList()));
columns.put(s.columnName, col);
}
for (CategoricalStats s : result.categoricalStats) {
Map<String, Object> col = new LinkedHashMap<>();
col.put("type", "CATEGORICAL");
col.put("count", s.count);
col.put("missing", s.missing);
col.put("uniqueCount", s.uniqueCount);
col.put("mostFrequent", s.mostFrequent);
col.put("mostFrequentCount", s.mostFrequentCount);
col.put("topFrequencies", s.topFrequencies);
columns.put(s.columnName, col);
}
report.put("columns", columns);
mapper.writeValue(new File(OUTPUT_JSON), report);
}
/**
* Convert a double to a JSON-safe Object.
* NaN / Infinity become null; finite values are rounded to 6 dp.
*/
private static Object safe(double v) {
if (Double.isNaN(v) || Double.isInfinite(v)) return null;
return Math.round(v * 1_000_000.0) / 1_000_000.0;
}
// =========================================================================
// Sample CSV generator
// =========================================================================
private static String generateSampleCsv() throws Exception {
Random rng = new Random(42);
String[] departments = {"Engineering", "Marketing", "Sales", "HR", "Finance"};
String[] cities = {
"New York", "London", "Berlin", "Tokyo", "Sydney",
"Paris", "Toronto", "Portland, OR", "Austin, TX"
};
List<String[]> rows = new ArrayList<>();
rows.add(new String[]{"age", "salary", "score", "height_cm",
"weight_kg", "department", "city"});
for (int i = 0; i < 200; i++) {
// Randomly missing values
String age = rng.nextDouble() < 0.05 ? ""
: String.valueOf(18 + rng.nextInt(48));
String salary = rng.nextDouble() < 0.03 ? ""
: String.format("%.2f", 30_000 + rng.nextDouble() * 170_000);
String score = rng.nextDouble() < 0.04 ? ""
: String.format("%.1f", rng.nextDouble() * 100);
String height = rng.nextDouble() < 0.02 ? ""
: String.format("%.1f", 150 + rng.nextDouble() * 50);
String weight = rng.nextDouble() < 0.06 ? ""
: String.format("%.1f", 50 + rng.nextDouble() * 70);
String dept = departments[rng.nextInt(departments.length)];
String city = cities[rng.nextInt(cities.length)];
// Inject obvious outliers for testing the IQR detector
if (i == 10) age = "130"; // extreme outlier
if (i == 50) salary = "4500000.00"; // extreme outlier
if (i == 100) score = "155.0"; // extreme outlier
if (i == 150) weight = "310.0"; // extreme outlier
// Row 199: simulate all numeric fields missing
if (i == 199) {
rows.add(new String[]{"", "", "", "", "", dept, city});
continue;
}
rows.add(new String[]{age, salary, score, height, weight, dept, city});
}
// Extra rows with quoted commas already in the cities array above
// ("Portland, OR" and "Austin, TX") — Apache Commons CSV will quote them.
CSVFormat outFmt = CSVFormat.DEFAULT.builder().build();
try (Writer writer = new FileWriter(SAMPLE_CSV);
CSVPrinter printer = new CSVPrinter(writer, outFmt)) {
for (String[] row : rows) {
printer.printRecord((Object[]) row);
}
}
return SAMPLE_CSV;
}
}