CSV Statistical Analyzer (java, written by Codex)
envgap__codex__java-t1-1
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/java-t1 #1 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>tmlr.codex_generated.p01</groupId>
<artifactId>csv-statistical-analyzer</artifactId>
<version>1.0.0</version>
<name>CSV Statistical Analyzer</name>
<description>Robust CSV analyzer for numeric and categorical statistics.</description>
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<maven.compiler.release>17</maven.compiler.release>
</properties>
<dependencies>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.13.0</version>
</plugin>
<plugin>
<groupId>org.codehaus.mojo</groupId>
<artifactId>exec-maven-plugin</artifactId>
<version>3.5.0</version>
<configuration>
<mainClass>CsvStatisticalAnalyzer</mainClass>
</configuration>
</plugin>
</plugins>
</build>
</project>
README.md
# CSV Statistical Analyzer (Java) A robust Java CLI that reads CSV data, auto-detects numeric/categorical columns, computes comprehensive statistics, detects IQR outliers, prints aligned console tables, and writes a full `report.json`. ## Features - Accept CSV file path as a command-line argument. - If no path is provided, generate `sample.csv` (250 rows, 5 numeric + 2 categorical columns) and analyze it. - Auto-detect numeric vs categorical columns from data. - Numeric statistics per column: - mean, median, standard deviation, variance - min, max - percentiles (25th, 50th, 75th) - non-missing value count - outliers via IQR method (`Q1 - 1.5*IQR`, `Q3 + 1.5*IQR`) - Categorical statistics per column: - unique count - most frequent value - top 10 value frequencies - Handles messy real-world CSV data: - missing values - mixed types - malformed rows (length mismatch; padded or truncated and recorded) - empty files and header-only files - quoted fields containing commas - Writes machine-readable JSON report to `report.json`. ## Project Structure - `src/main/java/CsvStatisticalAnalyzer.java` - full analyzer implementation. - `pom.xml` - Maven project file with pinned plugin versions and no external dependencies. ## Requirements - Ubuntu 22.04 - Java 17+ (JDK) - Optional: Maven 3.9+ (only if using Maven commands) ## Dependencies (Pinned) This implementation uses Java standard library only. - Direct runtime dependencies: none - Transitive runtime dependencies: none `pom.xml` includes pinned plugin versions for reproducible Maven builds. ## Build and Run ### Option A: Plain JDK (no Maven) ```bash cd TMLR/code_generation/codex_generated/p_01/java mkdir -p out javac -d out src/main/java/CsvStatisticalAnalyzer.java java -cp out CsvStatisticalAnalyzer ./your_file.csv ``` Run with no input to auto-generate sample data: ```bash java -cp out CsvStatisticalAnalyzer ``` ### Option B: Maven ```bash cd TMLR/code_generation/codex_generated/p_01/java mvn -q compile mvn -q exec:java -Dexec.args="./your_file.csv" ``` Run with no input file: ```bash mvn -q exec:java ``` ## Output - Console report: aligned summary tables for numeric and categorical stats and outlier details. - `report.json`: full structured analysis with: - metadata - column type classifications - numeric and categorical column stats - malformed row records - warning messages - If no input file is provided: - `sample.csv` is generated in the working directory. ## Statistical Notes - Variance and standard deviation are sample-based (`n - 1`) when `n > 1`. - Percentiles use linear interpolation on sorted values. - Numeric type detection classifies a column as numeric when at least 80% of non-missing values parse as valid numbers.
src/main/java/CsvStatisticalAnalyzer.java
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Collections;
import java.util.HashMap;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Objects;
import java.util.Random;
import java.util.Set;
import java.util.regex.Pattern;
public class CsvStatisticalAnalyzer {
private static final int DEFAULT_SAMPLE_ROWS = 250;
private static final Set<String> MISSING_TOKENS = new HashSet<>(Arrays.asList(
"", "na", "n/a", "null", "none", "nan", "undefined", "-"
));
private static final Pattern NUMBER_PATTERN = Pattern.compile("^[-+]?(?:\\d+(?:\\.\\d+)?|\\.\\d+)(?:[eE][-+]?\\d+)?$");
private static final Pattern COMMA_NUMBER_PATTERN = Pattern.compile("^[-+]?\\d{1,3}(?:,\\d{3})+(?:\\.\\d+)?$");
public static void main(String[] args) {
try {
run(args);
} catch (Exception ex) {
System.err.println("Unexpected failure while analyzing CSV data.");
ex.printStackTrace(System.err);
System.exit(1);
}
}
private static void run(String[] args) throws IOException {
Path inputPath;
if (args.length == 0) {
inputPath = Paths.get("sample.csv").toAbsolutePath().normalize();
generateSampleCsv(inputPath, DEFAULT_SAMPLE_ROWS);
System.out.println("No input file provided. Generated sample dataset at " + inputPath);
} else {
inputPath = Paths.get(args[0]).toAbsolutePath().normalize();
}
String content;
try {
content = Files.readString(inputPath, StandardCharsets.UTF_8);
} catch (IOException ex) {
System.err.println("Failed to read input file: " + inputPath);
System.err.println(ex.getMessage());
System.exit(1);
return;
}
Map<String, Object> report = buildBaseReport(inputPath.toString());
@SuppressWarnings("unchecked")
Map<String, Object> metadata = (Map<String, Object>) report.get("metadata");
@SuppressWarnings("unchecked")
List<String> warnings = (List<String>) report.get("warnings");
if (content.trim().isEmpty()) {
warnings.add("The input file is empty. No rows were available for analysis.");
} else {
List<List<String>> parsedRows = parseCsv(content);
NormalizedCsv normalized = normalizeRows(parsedRows);
metadata.put("totalColumns", normalized.headers.size());
metadata.put("sourceRowCount", normalized.sourceRowCount);
metadata.put("dataRowCount", normalized.dataRows.size());
metadata.put("malformedRowCount", normalized.malformedRows.size());
metadata.put("skippedBlankRowCount", normalized.skippedBlankRows);
report.put("malformedRows", normalized.malformedRows);
if (normalized.headers.isEmpty()) {
warnings.add("No header row was detected.");
} else if (normalized.dataRows.isEmpty()) {
warnings.add("Header detected, but no data rows were found.");
}
AnalysisResult analyzed = analyzeData(normalized.headers, normalized.dataRows);
report.put("columnClassifications", analyzed.columnClassifications);
report.put("numericColumns", analyzed.numericColumns);
report.put("categoricalColumns", analyzed.categoricalColumns);
}
summarizeForConsole(report);
Path outputPath = Paths.get("report.json").toAbsolutePath().normalize();
Files.writeString(outputPath, toJson(report) + System.lineSeparator(), StandardCharsets.UTF_8);
System.out.println();
System.out.println("Saved JSON report to " + outputPath);
}
private static Map<String, Object> buildBaseReport(String inputFile) {
Map<String, Object> report = new LinkedHashMap<>();
Map<String, Object> metadata = new LinkedHashMap<>();
metadata.put("generatedAt", Instant.now().toString());
metadata.put("inputFile", inputFile);
metadata.put("totalColumns", 0);
metadata.put("sourceRowCount", 0);
metadata.put("dataRowCount", 0);
metadata.put("malformedRowCount", 0);
metadata.put("skippedBlankRowCount", 0);
report.put("metadata", metadata);
report.put("columnClassifications", new LinkedHashMap<String, Object>());
report.put("numericColumns", new LinkedHashMap<String, Object>());
report.put("categoricalColumns", new LinkedHashMap<String, Object>());
report.put("malformedRows", new ArrayList<>());
report.put("warnings", new ArrayList<String>());
return report;
}
private static List<List<String>> parseCsv(String content) {
List<List<String>> rows = new ArrayList<>();
List<String> currentRow = new ArrayList<>();
StringBuilder currentField = new StringBuilder();
boolean inQuotes = false;
for (int i = 0; i < content.length(); i++) {
char ch = content.charAt(i);
if (inQuotes) {
if (ch == '"') {
if (i + 1 < content.length() && content.charAt(i + 1) == '"') {
currentField.append('"');
i++;
} else {
inQuotes = false;
}
} else {
currentField.append(ch);
}
continue;
}
if (ch == '"') {
if (currentField.length() == 0) {
inQuotes = true;
} else {
currentField.append(ch);
}
continue;
}
if (ch == ',') {
currentRow.add(currentField.toString());
currentField.setLength(0);
continue;
}
if (ch == '\n') {
currentRow.add(currentField.toString());
currentField.setLength(0);
rows.add(currentRow);
currentRow = new ArrayList<>();
continue;
}
if (ch == '\r') {
currentRow.add(currentField.toString());
currentField.setLength(0);
rows.add(currentRow);
currentRow = new ArrayList<>();
if (i + 1 < content.length() && content.charAt(i + 1) == '\n') {
i++;
}
continue;
}
currentField.append(ch);
}
if (currentField.length() > 0 || !currentRow.isEmpty() || content.endsWith(",")) {
currentRow.add(currentField.toString());
rows.add(currentRow);
}
return rows;
}
private static NormalizedCsv normalizeRows(List<List<String>> rows) {
if (rows.isEmpty()) {
return new NormalizedCsv(Collections.emptyList(), Collections.emptyList(), Collections.emptyList(), 0, 0);
}
List<String> headers = uniqueHeaders(rows.get(0));
int expectedColumns = headers.size();
List<RowData> dataRows = new ArrayList<>();
List<Map<String, Object>> malformedRows = new ArrayList<>();
int skippedBlankRows = 0;
for (int i = 1; i < rows.size(); i++) {
List<String> current = rows.get(i);
boolean allBlank = current.isEmpty();
if (!allBlank) {
allBlank = true;
for (String value : current) {
if (!isMissing(value)) {
allBlank = false;
break;
}
}
}
if (allBlank) {
skippedBlankRows++;
continue;
}
if (current.size() != expectedColumns) {
Map<String, Object> malformed = new LinkedHashMap<>();
malformed.put("rowNumber", i + 1);
malformed.put("expectedColumns", expectedColumns);
malformed.put("actualColumns", current.size());
malformed.put("action", current.size() < expectedColumns ? "padded_with_missing" : "truncated_extra_columns");
malformedRows.add(malformed);
}
List<String> normalized = new ArrayList<>(current.subList(0, Math.min(expectedColumns, current.size())));
while (normalized.size() < expectedColumns) {
normalized.add("");
}
dataRows.add(new RowData(i + 1, normalized));
}
return new NormalizedCsv(headers, dataRows, malformedRows, skippedBlankRows, rows.size());
}
private static List<String> uniqueHeaders(List<String> rawHeaders) {
List<String> result = new ArrayList<>();
Map<String, Integer> seen = new HashMap<>();
for (int i = 0; i < rawHeaders.size(); i++) {
String cleaned = Objects.toString(rawHeaders.get(i), "").replace("\uFEFF", "").trim();
if (cleaned.isEmpty()) {
cleaned = "column_" + (i + 1);
}
int count = seen.getOrDefault(cleaned, 0);
seen.put(cleaned, count + 1);
result.add(count == 0 ? cleaned : cleaned + "_" + (count + 1));
}
return result;
}
private static AnalysisResult analyzeData(List<String> headers, List<RowData> dataRows) {
Map<String, List<ColumnEntry>> entriesByColumn = new LinkedHashMap<>();
for (String header : headers) {
entriesByColumn.put(header, new ArrayList<>());
}
for (RowData row : dataRows) {
for (int i = 0; i < headers.size(); i++) {
entriesByColumn.get(headers.get(i)).add(new ColumnEntry(row.rowNumber, row.values.get(i)));
}
}
Map<String, Object> columnClassifications = new LinkedHashMap<>();
Map<String, Object> numericColumns = new LinkedHashMap<>();
Map<String, Object> categoricalColumns = new LinkedHashMap<>();
for (Map.Entry<String, List<ColumnEntry>> entry : entriesByColumn.entrySet()) {
String header = entry.getKey();
List<ColumnEntry> entries = entry.getValue();
List<String> values = new ArrayList<>();
for (ColumnEntry item : entries) {
values.add(item.value);
}
TypeInfo typeInfo = detectColumnType(values);
Map<String, Object> classification = new LinkedHashMap<>();
classification.put("type", typeInfo.type);
classification.put("nonMissingCount", typeInfo.nonMissingCount);
classification.put("numericValueCount", typeInfo.numericCount);
classification.put("numericRatio", round(typeInfo.numericRatio, 4));
classification.put("allValuesMissing", typeInfo.allValuesMissing);
columnClassifications.put(header, classification);
if ("numeric".equals(typeInfo.type)) {
numericColumns.put(header, analyzeNumericColumn(header, entries, typeInfo));
} else {
categoricalColumns.put(header, analyzeCategoricalColumn(header, entries, typeInfo));
}
}
return new AnalysisResult(columnClassifications, numericColumns, categoricalColumns);
}
private static TypeInfo detectColumnType(List<String> values) {
int nonMissingCount = 0;
int numericCount = 0;
for (String value : values) {
if (isMissing(value)) {
continue;
}
nonMissingCount++;
Double parsed = parsePotentialNumber(value);
if (parsed != null) {
numericCount++;
}
}
double ratio = nonMissingCount == 0 ? 0.0 : (double) numericCount / nonMissingCount;
String type = nonMissingCount > 0 && ratio >= 0.8 ? "numeric" : "categorical";
return new TypeInfo(type, nonMissingCount, numericCount, ratio, nonMissingCount == 0);
}
private static Map<String, Object> analyzeNumericColumn(String name, List<ColumnEntry> entries, TypeInfo typeInfo) {
int missingCount = 0;
int invalidCount = 0;
List<NumericEntry> numericEntries = new ArrayList<>();
for (ColumnEntry entry : entries) {
if (isMissing(entry.value)) {
missingCount++;
continue;
}
Double parsed = parsePotentialNumber(entry.value);
if (parsed != null) {
numericEntries.add(new NumericEntry(entry.rowNumber, parsed));
} else {
invalidCount++;
}
}
List<Double> numericValues = new ArrayList<>();
for (NumericEntry entry : numericEntries) {
numericValues.add(entry.value);
}
Collections.sort(numericValues);
Map<String, Object> result = new LinkedHashMap<>();
result.put("column", name);
result.put("type", "numeric");
result.put("typeConfidence", round(typeInfo.numericRatio, 4));
if (numericValues.isEmpty()) {
result.put("nonMissingValueCount", 0);
result.put("missingValueCount", missingCount);
result.put("invalidNumericValueCount", invalidCount);
result.put("mean", null);
result.put("median", null);
result.put("standardDeviation", null);
result.put("variance", null);
result.put("min", null);
result.put("max", null);
Map<String, Object> percentiles = new LinkedHashMap<>();
percentiles.put("p25", null);
percentiles.put("p50", null);
percentiles.put("p75", null);
result.put("percentiles", percentiles);
Map<String, Object> outliers = new LinkedHashMap<>();
outliers.put("lowerBound", null);
outliers.put("upperBound", null);
outliers.put("iqr", null);
outliers.put("count", 0);
outliers.put("values", new ArrayList<>());
result.put("outliers", outliers);
return result;
}
int n = numericValues.size();
double sum = 0.0;
for (Double value : numericValues) {
sum += value;
}
double mean = sum / n;
double q1 = percentile(numericValues, 25);
double q2 = percentile(numericValues, 50);
double q3 = percentile(numericValues, 75);
double variance = 0.0;
if (n > 1) {
for (Double value : numericValues) {
double diff = value - mean;
variance += diff * diff;
}
variance /= (n - 1);
}
double stdDev = Math.sqrt(variance);
double iqr = q3 - q1;
double lowerBound = q1 - 1.5 * iqr;
double upperBound = q3 + 1.5 * iqr;
List<Map<String, Object>> outlierValues = new ArrayList<>();
for (NumericEntry entry : numericEntries) {
if (entry.value < lowerBound || entry.value > upperBound) {
Map<String, Object> outlier = new LinkedHashMap<>();
outlier.put("rowNumber", entry.rowNumber);
outlier.put("value", entry.value);
outlierValues.add(outlier);
}
}
result.put("nonMissingValueCount", n);
result.put("missingValueCount", missingCount);
result.put("invalidNumericValueCount", invalidCount);
result.put("mean", mean);
result.put("median", q2);
result.put("standardDeviation", stdDev);
result.put("variance", variance);
result.put("min", numericValues.get(0));
result.put("max", numericValues.get(n - 1));
Map<String, Object> percentiles = new LinkedHashMap<>();
percentiles.put("p25", q1);
percentiles.put("p50", q2);
percentiles.put("p75", q3);
result.put("percentiles", percentiles);
Map<String, Object> outliers = new LinkedHashMap<>();
outliers.put("lowerBound", lowerBound);
outliers.put("upperBound", upperBound);
outliers.put("iqr", iqr);
outliers.put("count", outlierValues.size());
outliers.put("values", outlierValues);
result.put("outliers", outliers);
return result;
}
private static Map<String, Object> analyzeCategoricalColumn(String name, List<ColumnEntry> entries, TypeInfo typeInfo) {
int missingCount = 0;
Map<String, Integer> frequencies = new HashMap<>();
for (ColumnEntry entry : entries) {
if (isMissing(entry.value)) {
missingCount++;
continue;
}
String value = entry.value.trim();
frequencies.put(value, frequencies.getOrDefault(value, 0) + 1);
}
List<Map.Entry<String, Integer>> sorted = new ArrayList<>(frequencies.entrySet());
sorted.sort((a, b) -> {
int byCount = Integer.compare(b.getValue(), a.getValue());
if (byCount != 0) {
return byCount;
}
return a.getKey().compareTo(b.getKey());
});
List<Map<String, Object>> topFrequencies = new ArrayList<>();
for (int i = 0; i < Math.min(10, sorted.size()); i++) {
Map.Entry<String, Integer> item = sorted.get(i);
Map<String, Object> freq = new LinkedHashMap<>();
freq.put("value", item.getKey());
freq.put("count", item.getValue());
topFrequencies.add(freq);
}
Map<String, Object> result = new LinkedHashMap<>();
result.put("column", name);
result.put("type", "categorical");
result.put("typeConfidence", round(1.0 - typeInfo.numericRatio, 4));
result.put("nonMissingValueCount", entries.size() - missingCount);
result.put("missingValueCount", missingCount);
result.put("uniqueCount", frequencies.size());
result.put("mostFrequentValue", sorted.isEmpty() ? null : sorted.get(0).getKey());
result.put("topFrequencies", topFrequencies);
return result;
}
private static boolean isMissing(String value) {
if (value == null) {
return true;
}
String normalized = value.trim().toLowerCase(Locale.ROOT);
return MISSING_TOKENS.contains(normalized);
}
private static Double parsePotentialNumber(String value) {
if (value == null) {
return null;
}
String text = value.trim();
if (text.isEmpty()) {
return null;
}
try {
if (NUMBER_PATTERN.matcher(text).matches()) {
double parsed = Double.parseDouble(text);
return Double.isFinite(parsed) ? parsed : null;
}
if (COMMA_NUMBER_PATTERN.matcher(text).matches()) {
double parsed = Double.parseDouble(text.replace(",", ""));
return Double.isFinite(parsed) ? parsed : null;
}
} catch (NumberFormatException ignored) {
return null;
}
return null;
}
private static double percentile(List<Double> sortedValues, int p) {
if (sortedValues.isEmpty()) {
return Double.NaN;
}
if (sortedValues.size() == 1) {
return sortedValues.get(0);
}
double rank = (p / 100.0) * (sortedValues.size() - 1);
int lower = (int) Math.floor(rank);
int upper = (int) Math.ceil(rank);
if (lower == upper) {
return sortedValues.get(lower);
}
double weight = rank - lower;
return sortedValues.get(lower) * (1.0 - weight) + sortedValues.get(upper) * weight;
}
private static void summarizeForConsole(Map<String, Object> report) {
@SuppressWarnings("unchecked")
Map<String, Object> metadata = (Map<String, Object>) report.get("metadata");
@SuppressWarnings("unchecked")
List<String> warnings = (List<String>) report.get("warnings");
@SuppressWarnings("unchecked")
Map<String, Object> numericColumns = (Map<String, Object>) report.get("numericColumns");
@SuppressWarnings("unchecked")
Map<String, Object> categoricalColumns = (Map<String, Object>) report.get("categoricalColumns");
System.out.println("CSV Statistical Analyzer");
System.out.println("========================");
System.out.println("Input File: " + metadata.get("inputFile"));
System.out.println("Generated : " + metadata.get("generatedAt"));
System.out.println("Columns : " + metadata.get("totalColumns"));
System.out.println("Data Rows : " + metadata.get("dataRowCount"));
System.out.println("Malformed : " + metadata.get("malformedRowCount"));
if (!warnings.isEmpty()) {
System.out.println();
System.out.println("Warnings:");
for (String warning : warnings) {
System.out.println("- " + warning);
}
}
List<List<String>> numericRows = new ArrayList<>();
for (Object value : numericColumns.values()) {
@SuppressWarnings("unchecked")
Map<String, Object> col = (Map<String, Object>) value;
@SuppressWarnings("unchecked")
Map<String, Object> percentiles = (Map<String, Object>) col.get("percentiles");
@SuppressWarnings("unchecked")
Map<String, Object> outliers = (Map<String, Object>) col.get("outliers");
List<String> row = new ArrayList<>();
row.add(String.valueOf(col.get("column")));
row.add(String.valueOf(col.get("nonMissingValueCount")));
row.add(formatNumber(col.get("mean")));
row.add(formatNumber(col.get("median")));
row.add(formatNumber(col.get("standardDeviation")));
row.add(formatNumber(col.get("variance")));
row.add(formatNumber(col.get("min")));
row.add(formatNumber(percentiles.get("p25")));
row.add(formatNumber(percentiles.get("p50")));
row.add(formatNumber(percentiles.get("p75")));
row.add(formatNumber(col.get("max")));
row.add(String.valueOf(outliers.get("count")));
numericRows.add(row);
}
printTable(
"Numeric Columns",
Arrays.asList("Column", "Count", "Mean", "Median", "StdDev", "Variance", "Min", "P25", "P50", "P75", "Max", "Outliers"),
numericRows
);
List<List<String>> categoricalRows = new ArrayList<>();
for (Object value : categoricalColumns.values()) {
@SuppressWarnings("unchecked")
Map<String, Object> col = (Map<String, Object>) value;
@SuppressWarnings("unchecked")
List<Map<String, Object>> topFrequencies = (List<Map<String, Object>>) col.get("topFrequencies");
List<String> previewItems = new ArrayList<>();
for (int i = 0; i < Math.min(3, topFrequencies.size()); i++) {
Map<String, Object> entry = topFrequencies.get(i);
previewItems.add(entry.get("value") + " (" + entry.get("count") + ")");
}
List<String> row = new ArrayList<>();
row.add(String.valueOf(col.get("column")));
row.add(String.valueOf(col.get("nonMissingValueCount")));
row.add(String.valueOf(col.get("uniqueCount")));
row.add(col.get("mostFrequentValue") == null ? "N/A" : String.valueOf(col.get("mostFrequentValue")));
row.add(previewItems.isEmpty() ? "N/A" : String.join(", ", previewItems));
categoricalRows.add(row);
}
printTable(
"Categorical Columns",
Arrays.asList("Column", "NonMissing", "Unique", "Most Frequent", "Top Values (up to 3 shown)"),
categoricalRows
);
System.out.println();
System.out.println("Outlier Details (IQR Method)");
if (numericColumns.isEmpty()) {
System.out.println(" No numeric columns were detected.");
return;
}
for (Object value : numericColumns.values()) {
@SuppressWarnings("unchecked")
Map<String, Object> col = (Map<String, Object>) value;
@SuppressWarnings("unchecked")
Map<String, Object> outliers = (Map<String, Object>) col.get("outliers");
@SuppressWarnings("unchecked")
List<Map<String, Object>> values = (List<Map<String, Object>>) outliers.get("values");
if (values.isEmpty()) {
System.out.println("- " + col.get("column") + ": none");
continue;
}
List<String> pieces = new ArrayList<>();
for (Map<String, Object> outlier : values) {
pieces.add("row " + outlier.get("rowNumber") + " -> " + formatNumber(outlier.get("value")));
}
System.out.println("- " + col.get("column") + ": " + String.join(", ", pieces));
}
}
private static void printTable(String title, List<String> columns, List<List<String>> rows) {
System.out.println();
System.out.println(title);
if (rows.isEmpty()) {
System.out.println(" (none)");
return;
}
int[] widths = new int[columns.size()];
for (int i = 0; i < columns.size(); i++) {
widths[i] = columns.get(i).length();
}
for (List<String> row : rows) {
for (int i = 0; i < row.size(); i++) {
widths[i] = Math.max(widths[i], row.get(i).length());
}
}
System.out.println(renderRow(columns, widths));
List<String> separatorParts = new ArrayList<>();
for (int width : widths) {
separatorParts.add("-".repeat(width));
}
System.out.println(String.join("-+-", separatorParts));
for (List<String> row : rows) {
System.out.println(renderRow(row, widths));
}
}
private static String renderRow(List<String> values, int[] widths) {
List<String> padded = new ArrayList<>();
for (int i = 0; i < widths.length; i++) {
String value = i < values.size() ? values.get(i) : "";
padded.add(padRight(value, widths[i]));
}
return String.join(" | ", padded);
}
private static String padRight(String text, int width) {
if (text.length() >= width) {
return text;
}
StringBuilder builder = new StringBuilder(text);
while (builder.length() < width) {
builder.append(' ');
}
return builder.toString();
}
private static String formatNumber(Object value) {
if (!(value instanceof Number)) {
return "N/A";
}
double d = ((Number) value).doubleValue();
if (!Double.isFinite(d)) {
return String.valueOf(d);
}
if (d == 0.0) {
return "0";
}
double abs = Math.abs(d);
if (abs >= 1_000_000 || abs < 0.0001) {
return String.format(Locale.ROOT, "%.4e", d);
}
String format = abs >= 1000 ? "%.2f" : "%.4f";
String result = String.format(Locale.ROOT, format, d);
result = result.replaceAll("\\.0+$", "").replaceAll("(\\.\\d*?)0+$", "$1");
return result;
}
private static void generateSampleCsv(Path filePath, int rowCount) throws IOException {
List<String> headers = Arrays.asList(
"temperature_c",
"revenue_usd",
"latency_ms",
"quality_score",
"units_sold",
"region",
"segment"
);
List<String> regions = Arrays.asList("North", "South", "East", "West", "New York, NY");
List<String> segments = Arrays.asList("Consumer", "Enterprise", "SMB", "Public Sector");
Random random = new Random(42L);
List<String> lines = new ArrayList<>();
lines.add(String.join(",", headers));
for (int i = 0; i < rowCount; i++) {
double temperature = gaussian(random, 20.0, 7.0);
double revenue = gaussian(random, 15000.0, 3500.0);
Object latency = gaussian(random, 120.0, 35.0);
Object quality = clamp(gaussian(random, 82.0, 9.0), 0.0, 100.0);
Object units = Math.max(0L, Math.round(gaussian(random, 240.0, 80.0)));
if (i % 57 == 0) {
revenue *= 3.5;
}
if (i % 41 == 0) {
latency = "timeout";
}
if (i % 23 == 0) {
quality = "";
}
if (i % 37 == 0) {
units = "1,250";
}
String region = regions.get(i % regions.size());
String segment = segments.get((i * 3) % segments.size());
List<String> values = Arrays.asList(
String.format(Locale.ROOT, "%.3f", temperature),
String.format(Locale.ROOT, "%.2f", revenue),
latency instanceof Number ? String.format(Locale.ROOT, "%.3f", ((Number) latency).doubleValue()) : String.valueOf(latency),
quality instanceof Number ? String.format(Locale.ROOT, "%.2f", ((Number) quality).doubleValue()) : String.valueOf(quality),
String.valueOf(units),
region,
segment
);
List<String> escaped = new ArrayList<>();
for (String value : values) {
escaped.add(csvEscape(value));
}
lines.add(String.join(",", escaped));
}
Files.writeString(filePath, String.join(System.lineSeparator(), lines) + System.lineSeparator(), StandardCharsets.UTF_8);
}
private static double gaussian(Random random, double mean, double stdDev) {
return mean + random.nextGaussian() * stdDev;
}
private static double clamp(double value, double min, double max) {
return Math.max(min, Math.min(max, value));
}
private static String csvEscape(String value) {
String text = value == null ? "" : value;
if (text.contains("\"") || text.contains(",") || text.contains("\n") || text.contains("\r")) {
return "\"" + text.replace("\"", "\"\"") + "\"";
}
return text;
}
private static String toJson(Object value) {
return toJson(value, 0);
}
@SuppressWarnings("unchecked")
private static String toJson(Object value, int indent) {
if (value == null) {
return "null";
}
if (value instanceof String) {
return "\"" + escapeJson((String) value) + "\"";
}
if (value instanceof Boolean || value instanceof Integer || value instanceof Long || value instanceof Short || value instanceof Byte) {
return String.valueOf(value);
}
if (value instanceof Float || value instanceof Double) {
double d = ((Number) value).doubleValue();
if (!Double.isFinite(d)) {
return "null";
}
return formatJsonDouble(d);
}
if (value instanceof Number) {
return String.valueOf(value);
}
if (value instanceof Map) {
Map<String, Object> map = (Map<String, Object>) value;
if (map.isEmpty()) {
return "{}";
}
StringBuilder builder = new StringBuilder();
builder.append("{\n");
int index = 0;
for (Map.Entry<String, Object> entry : map.entrySet()) {
builder.append(" ".repeat(indent + 1));
builder.append("\"").append(escapeJson(entry.getKey())).append("\": ");
builder.append(toJson(entry.getValue(), indent + 1));
if (index < map.size() - 1) {
builder.append(',');
}
builder.append('\n');
index++;
}
builder.append(" ".repeat(indent)).append('}');
return builder.toString();
}
if (value instanceof List) {
List<Object> list = (List<Object>) value;
if (list.isEmpty()) {
return "[]";
}
StringBuilder builder = new StringBuilder();
builder.append("[\n");
for (int i = 0; i < list.size(); i++) {
builder.append(" ".repeat(indent + 1));
builder.append(toJson(list.get(i), indent + 1));
if (i < list.size() - 1) {
builder.append(',');
}
builder.append('\n');
}
builder.append(" ".repeat(indent)).append(']');
return builder.toString();
}
return "\"" + escapeJson(String.valueOf(value)) + "\"";
}
private static String formatJsonDouble(double value) {
String text = String.format(Locale.ROOT, "%.15f", value);
text = text.replaceAll("\\.0+$", "").replaceAll("(\\.\\d*?)0+$", "$1");
if (text.equals("-0")) {
return "0";
}
return text;
}
private static String escapeJson(String text) {
StringBuilder escaped = new StringBuilder();
for (int i = 0; i < text.length(); i++) {
char ch = text.charAt(i);
switch (ch) {
case '\\':
escaped.append("\\\\");
break;
case '"':
escaped.append("\\\"");
break;
case '\b':
escaped.append("\\b");
break;
case '\f':
escaped.append("\\f");
break;
case '\n':
escaped.append("\\n");
break;
case '\r':
escaped.append("\\r");
break;
case '\t':
escaped.append("\\t");
break;
default:
if (ch < 0x20) {
escaped.append(String.format("\\u%04x", (int) ch));
} else {
escaped.append(ch);
}
}
}
return escaped.toString();
}
private static double round(double value, int decimals) {
double factor = Math.pow(10, decimals);
return Math.round(value * factor) / factor;
}
private static final class RowData {
final int rowNumber;
final List<String> values;
RowData(int rowNumber, List<String> values) {
this.rowNumber = rowNumber;
this.values = values;
}
}
private static final class ColumnEntry {
final int rowNumber;
final String value;
ColumnEntry(int rowNumber, String value) {
this.rowNumber = rowNumber;
this.value = value;
}
}
private static final class NumericEntry {
final int rowNumber;
final double value;
NumericEntry(int rowNumber, double value) {
this.rowNumber = rowNumber;
this.value = value;
}
}
private static final class TypeInfo {
final String type;
final int nonMissingCount;
final int numericCount;
final double numericRatio;
final boolean allValuesMissing;
TypeInfo(String type, int nonMissingCount, int numericCount, double numericRatio, boolean allValuesMissing) {
this.type = type;
this.nonMissingCount = nonMissingCount;
this.numericCount = numericCount;
this.numericRatio = numericRatio;
this.allValuesMissing = allValuesMissing;
}
}
private static final class NormalizedCsv {
final List<String> headers;
final List<RowData> dataRows;
final List<Map<String, Object>> malformedRows;
final int skippedBlankRows;
final int sourceRowCount;
NormalizedCsv(List<String> headers, List<RowData> dataRows, List<Map<String, Object>> malformedRows, int skippedBlankRows, int sourceRowCount) {
this.headers = headers;
this.dataRows = dataRows;
this.malformedRows = malformedRows;
this.skippedBlankRows = skippedBlankRows;
this.sourceRowCount = sourceRowCount;
}
}
private static final class AnalysisResult {
final Map<String, Object> columnClassifications;
final Map<String, Object> numericColumns;
final Map<String, Object> categoricalColumns;
AnalysisResult(Map<String, Object> columnClassifications, Map<String, Object> numericColumns, Map<String, Object> categoricalColumns) {
this.columnClassifications = columnClassifications;
this.numericColumns = numericColumns;
this.categoricalColumns = categoricalColumns;
}
}
}