Data Profiling Tool (java, written by Codex)
envgap__codex__java-t1-7
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/java-t1 #7 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Data Profiling Tool Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report. FUNCTIONAL REQUIREMENTS: - Accept a CSV or JSON data file path as a command-line argument - Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings) - For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th) - For string columns: compute min/max/average length, most common values (top 10), and unique count - For all columns: count total values, missing/null values, missing percentage, and unique value count - Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates - Compute a pairwise correlation matrix for all numeric columns - Print a formatted summary report to console showing key statistics per column - Save the full profiling report as a JSON file with --output flag (default: data_profile.json) - If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it - Handle files with inconsistent delimiters or encoding issues gracefully Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>tmlr.codex_generated.p07</groupId>
<artifactId>data-profiling-tool</artifactId>
<version>1.0.0</version>
<name>Data Profiling Tool</name>
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<maven.compiler.release>17</maven.compiler.release>
</properties>
<dependencies>
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-databind</artifactId>
<version>2.18.2</version>
</dependency>
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-core</artifactId>
<version>2.18.2</version>
</dependency>
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-annotations</artifactId>
<version>2.18.2</version>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.13.0</version>
</plugin>
<plugin>
<groupId>org.codehaus.mojo</groupId>
<artifactId>exec-maven-plugin</artifactId>
<version>3.5.0</version>
<configuration>
<mainClass>DataProfilingTool</mainClass>
</configuration>
</plugin>
</plugins>
</build>
</project>
README.md
# Data Profiling Tool (Java) Profiles CSV/JSON datasets for column typing, missingness, distributions, outliers, correlations, and quality issues. ## Requirements - Ubuntu 22.04 - JDK 17+ - Maven 3.9+ ## Dependencies - Direct: - `com.fasterxml.jackson.core:jackson-databind:2.18.2` - `com.fasterxml.jackson.core:jackson-core:2.18.2` - `com.fasterxml.jackson.core:jackson-annotations:2.18.2` - Transitive: - none (explicitly pinned above) All versions are pinned in `pom.xml`. ## Build ```bash mvn -q -DskipTests compile ``` ## Run With input: ```bash mvn -q exec:java -Dexec.args="/path/to/data.csv --output data_profile.json" ``` JSON input: ```bash mvn -q exec:java -Dexec.args="/path/to/data.json --output data_profile.json" ``` No input (generates sample dataset with intentional quality issues): ```bash mvn -q exec:java ``` ## Output - Console summary table per column - Full JSON report (`data_profile.json` by default)
src/main/java/DataProfilingTool.java
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.BufferedWriter;
import java.io.IOException;
import java.nio.ByteBuffer;
import java.nio.charset.CharacterCodingException;
import java.nio.charset.CharsetDecoder;
import java.nio.charset.CodingErrorAction;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.time.LocalDate;
import java.time.LocalDateTime;
import java.time.OffsetDateTime;
import java.time.format.DateTimeFormatter;
import java.time.format.DateTimeParseException;
import java.util.ArrayList;
import java.util.Collections;
import java.util.Comparator;
import java.util.HashMap;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
public class DataProfilingTool {
private static final ObjectMapper MAPPER = new ObjectMapper();
public static void main(String[] args) throws Exception {
ParsedArgs parsed = parseArgs(args);
String output = parsed.options.getOrDefault("output", "data_profile.json");
Path outputPath = Paths.get(output).toAbsolutePath();
List<Map<String, String>> rows;
Map<String, Object> metadata = new LinkedHashMap<>();
if (parsed.positional.isEmpty()) {
rows = generateSampleRows();
Path samplePath = Paths.get("sample_profile_data.csv").toAbsolutePath();
writeSampleCsv(rows, samplePath);
metadata.put("generated_sample", samplePath.toString());
System.out.println("No input file provided. Generated sample dataset: " + samplePath);
} else {
Path inputPath = Paths.get(parsed.positional.get(0)).toAbsolutePath();
if (!Files.exists(inputPath)) {
System.err.println("Input file not found: " + inputPath);
System.exit(1);
return;
}
String ext = extension(inputPath);
if (ext.equals("json") || ext.equals("jsonl") || ext.equals("ndjson")) {
LoadResult load = loadJson(inputPath);
rows = load.rows;
metadata.put("input_format", "json");
metadata.put("encoding", load.encoding);
} else {
LoadResult load = loadCsv(inputPath);
rows = load.rows;
metadata.put("input_format", "csv");
metadata.put("encoding", load.encoding);
metadata.put("delimiter", load.delimiter);
}
metadata.put("input_file", inputPath.toString());
}
Map<String, Object> report = profile(rows);
report.put("generated_at", Instant.now().toString());
for (Map.Entry<String, Object> e : metadata.entrySet()) {
report.put(e.getKey(), e.getValue());
}
printSummary(report);
Files.writeString(outputPath, MAPPER.writerWithDefaultPrettyPrinter().writeValueAsString(report) + "\n", StandardCharsets.UTF_8);
System.out.println("Saved JSON profile: " + outputPath);
}
private static String extension(Path path) {
String name = path.getFileName().toString();
int dot = name.lastIndexOf('.');
return dot >= 0 ? name.substring(dot + 1).toLowerCase() : "";
}
private record ParsedArgs(Map<String, String> options, List<String> positional) {}
private record LoadResult(List<Map<String, String>> rows, String encoding, String delimiter) {}
private static ParsedArgs parseArgs(String[] args) {
Map<String, String> options = new HashMap<>();
List<String> positional = new ArrayList<>();
for (int i = 0; i < args.length; i++) {
String token = args[i];
if (token.startsWith("--")) {
String key = token.substring(2);
if (i + 1 < args.length && !args[i + 1].startsWith("--")) {
options.put(key, args[++i]);
} else {
options.put(key, "true");
}
} else {
positional.add(token);
}
}
return new ParsedArgs(options, positional);
}
private static String decodeUtf8ThenLatin1(byte[] bytes) {
CharsetDecoder dec = StandardCharsets.UTF_8.newDecoder();
dec.onMalformedInput(CodingErrorAction.REPORT);
dec.onUnmappableCharacter(CodingErrorAction.REPORT);
try {
return dec.decode(ByteBuffer.wrap(bytes)).toString();
} catch (CharacterCodingException e) {
return new String(bytes, StandardCharsets.ISO_8859_1);
}
}
private static String detectEncoding(byte[] bytes) {
CharsetDecoder dec = StandardCharsets.UTF_8.newDecoder();
dec.onMalformedInput(CodingErrorAction.REPORT);
dec.onUnmappableCharacter(CodingErrorAction.REPORT);
try {
dec.decode(ByteBuffer.wrap(bytes));
return "utf-8";
} catch (CharacterCodingException e) {
return "latin1";
}
}
private static List<String> parseCsvLine(String line, char delimiter) {
List<String> out = new ArrayList<>();
StringBuilder current = new StringBuilder();
boolean inQuotes = false;
for (int i = 0; i < line.length(); i++) {
char ch = line.charAt(i);
if (ch == '"') {
if (inQuotes && i + 1 < line.length() && line.charAt(i + 1) == '"') {
current.append('"');
i++;
} else {
inQuotes = !inQuotes;
}
} else if (ch == delimiter && !inQuotes) {
out.add(current.toString());
current.setLength(0);
} else {
current.append(ch);
}
}
out.add(current.toString());
return out;
}
private static char detectDelimiter(List<String> lines) {
char[] cands = new char[]{',', ';', '\t', '|'};
char best = ',';
int bestScore = -1;
for (char c : cands) {
int score = 0;
for (String line : lines) {
for (int i = 0; i < line.length(); i++) if (line.charAt(i) == c) score++;
}
if (score > bestScore) {
bestScore = score;
best = c;
}
}
return best;
}
private static LoadResult loadCsv(Path path) throws IOException {
byte[] bytes = Files.readAllBytes(path);
String encoding = detectEncoding(bytes);
String text = decodeUtf8ThenLatin1(bytes);
List<String> lines = new ArrayList<>();
for (String line : text.replace("\r\n", "\n").replace("\r", "\n").split("\n")) {
if (!line.trim().isEmpty()) lines.add(line);
}
if (lines.isEmpty()) return new LoadResult(new ArrayList<>(), encoding, ",");
char delimiter = detectDelimiter(lines.subList(0, Math.min(5, lines.size())));
List<String> headers = parseCsvLine(lines.get(0), delimiter);
List<Map<String, String>> rows = new ArrayList<>();
for (int i = 1; i < lines.size(); i++) {
List<String> values = parseCsvLine(lines.get(i), delimiter);
Map<String, String> row = new LinkedHashMap<>();
for (int j = 0; j < headers.size(); j++) {
row.put(headers.get(j).trim(), j < values.size() ? values.get(j) : "");
}
rows.add(row);
}
return new LoadResult(rows, encoding, String.valueOf(delimiter));
}
private static LoadResult loadJson(Path path) throws IOException {
byte[] bytes = Files.readAllBytes(path);
String encoding = detectEncoding(bytes);
String text = decodeUtf8ThenLatin1(bytes);
List<Map<String, String>> rows = new ArrayList<>();
try {
JsonNode node = MAPPER.readTree(text);
if (node != null && node.isArray()) {
for (JsonNode item : node) {
if (item.isObject()) rows.add(nodeToFlatMap(item));
}
} else if (node != null && node.isObject()) {
rows.add(nodeToFlatMap(node));
}
} catch (Exception e) {
for (String line : text.split("\\R")) {
if (line.trim().isEmpty()) continue;
JsonNode item = MAPPER.readTree(line);
if (item.isObject()) rows.add(nodeToFlatMap(item));
}
}
return new LoadResult(rows, encoding, null);
}
private static Map<String, String> nodeToFlatMap(JsonNode node) {
Map<String, String> out = new LinkedHashMap<>();
node.fields().forEachRemaining(e -> out.put(e.getKey(), e.getValue().isNull() ? "" : e.getValue().asText()));
return out;
}
private static boolean isMissing(String v) {
if (v == null) return true;
String s = v.trim().toLowerCase();
return s.isEmpty() || s.equals("null") || s.equals("na") || s.equals("n/a") || s.equals("none");
}
private static Double parseNumber(String v) {
if (v == null) return null;
String s = v.trim();
if (s.isEmpty()) return null;
if (!s.matches("^[-+]?\\d+(\\.\\d+)?$")) return null;
try {
return Double.parseDouble(s);
} catch (NumberFormatException e) {
return null;
}
}
private static Boolean parseBoolean(String v) {
if (v == null) return null;
String s = v.trim().toLowerCase();
if (Set.of("true", "1", "yes", "y").contains(s)) return true;
if (Set.of("false", "0", "no", "n").contains(s)) return false;
return null;
}
private static Double parseDateEpoch(String v) {
if (v == null) return null;
String s = v.trim();
if (s.isEmpty()) return null;
try { return (double) Instant.parse(s).getEpochSecond(); } catch (DateTimeParseException ignored) {}
try { return (double) OffsetDateTime.parse(s).toEpochSecond(); } catch (DateTimeParseException ignored) {}
try { return (double) LocalDateTime.parse(s, DateTimeFormatter.ISO_LOCAL_DATE_TIME).toEpochSecond(java.time.ZoneOffset.UTC); } catch (DateTimeParseException ignored) {}
try { return (double) LocalDate.parse(s, DateTimeFormatter.ISO_LOCAL_DATE).atStartOfDay(java.time.ZoneOffset.UTC).toEpochSecond(); } catch (DateTimeParseException ignored) {}
return null;
}
private static String detectType(List<String> values) {
if (values.isEmpty()) return "string";
int boolCount = 0, numCount = 0, dateCount = 0;
for (String v : values) {
if (parseBoolean(v) != null) boolCount++;
if (parseNumber(v) != null) numCount++;
if (parseDateEpoch(v) != null) dateCount++;
}
if (boolCount == values.size()) return "boolean";
if (numCount == values.size()) {
boolean hasFloat = values.stream().anyMatch(s -> s.contains("."));
return hasFloat ? "float" : "integer";
}
if (dateCount >= Math.max(3, (int) Math.floor(values.size() * 0.9))) return "date";
Set<String> unique = new HashSet<>(values);
double ratio = values.isEmpty() ? 0.0 : (double) unique.size() / values.size();
if (unique.size() <= 20 || ratio <= 0.1) return "categorical";
return "string";
}
private static Double mean(List<Double> values) {
if (values.isEmpty()) return null;
double sum = 0.0;
for (Double v : values) sum += v;
return sum / values.size();
}
private static double stddev(List<Double> values, Double mean) {
if (values.size() <= 1) return 0.0;
double mu = mean != null ? mean : mean(values);
double acc = 0.0;
for (Double v : values) acc += (v - mu) * (v - mu);
return Math.sqrt(acc / values.size());
}
private static double skewness(List<Double> values, Double mean, Double sd) {
if (values.size() < 3) return 0.0;
double mu = mean != null ? mean : mean(values);
double sigma = sd != null ? sd : stddev(values, mu);
if (sigma == 0.0) return 0.0;
double m3 = 0.0;
for (Double v : values) m3 += Math.pow(v - mu, 3);
m3 /= values.size();
return m3 / Math.pow(sigma, 3);
}
private static Double percentile(List<Double> sorted, double p) {
if (sorted.isEmpty()) return null;
if (sorted.size() == 1) return sorted.get(0);
double pos = (sorted.size() - 1) * p;
int lo = (int) Math.floor(pos);
int hi = (int) Math.ceil(pos);
if (lo == hi) return sorted.get(lo);
double w = pos - lo;
return sorted.get(lo) + (sorted.get(hi) - sorted.get(lo)) * w;
}
private static Double correlation(List<Double> xs, List<Double> ys) {
List<Double> x = new ArrayList<>();
List<Double> y = new ArrayList<>();
int n = Math.min(xs.size(), ys.size());
for (int i = 0; i < n; i++) {
if (xs.get(i) == null || ys.get(i) == null) continue;
x.add(xs.get(i));
y.add(ys.get(i));
}
if (x.size() < 2) return null;
Double mx = mean(x), my = mean(y);
double sx = stddev(x, mx), sy = stddev(y, my);
if (sx == 0.0 || sy == 0.0) return 0.0;
double cov = 0.0;
for (int i = 0; i < x.size(); i++) cov += (x.get(i) - mx) * (y.get(i) - my);
cov /= x.size();
return cov / (sx * sy);
}
private static List<Map<String, String>> topValues(List<String> values, int n) {
Map<String, Integer> freq = new HashMap<>();
for (String v : values) freq.put(v, freq.getOrDefault(v, 0) + 1);
List<Map.Entry<String, Integer>> items = new ArrayList<>(freq.entrySet());
items.sort((a, b) -> {
int c = Integer.compare(b.getValue(), a.getValue());
return c != 0 ? c : a.getKey().compareTo(b.getKey());
});
List<Map<String, String>> out = new ArrayList<>();
for (int i = 0; i < Math.min(n, items.size()); i++) {
Map<String, String> row = new LinkedHashMap<>();
row.put("value", items.get(i).getKey());
row.put("count", String.valueOf(items.get(i).getValue()));
out.add(row);
}
return out;
}
private static Map<String, Object> profile(List<Map<String, String>> rows) {
Set<String> colSet = new HashSet<>();
for (Map<String, String> row : rows) colSet.addAll(row.keySet());
List<String> columns = new ArrayList<>(colSet);
Collections.sort(columns);
Map<String, Object> columnProfiles = new LinkedHashMap<>();
List<Map<String, Object>> issues = new ArrayList<>();
List<String> numericColumns = new ArrayList<>();
Map<String, List<Double>> numericSeries = new HashMap<>();
for (String col : columns) {
List<String> raw = new ArrayList<>();
for (Map<String, String> row : rows) raw.add(row.get(col));
int total = raw.size();
int missing = 0;
List<String> nonMissing = new ArrayList<>();
for (String v : raw) {
if (isMissing(v)) missing++;
else nonMissing.add(v);
}
int unique = new HashSet<>(nonMissing).size();
String type = detectType(nonMissing);
double missingPct = total == 0 ? 0.0 : ((double) missing / total) * 100.0;
Map<String, Object> p = new LinkedHashMap<>();
p.put("type", type);
p.put("total_values", total);
p.put("missing_values", missing);
p.put("missing_percentage", Math.round(missingPct * 10000.0) / 10000.0);
p.put("unique_values", unique);
List<String> quality = new ArrayList<>();
if (missing == total) {
quality.add("entirely_null");
issues.add(Map.of("column", col, "issue", "entirely_null"));
}
if (unique == 1 && !nonMissing.isEmpty()) {
quality.add("single_unique_value");
issues.add(Map.of("column", col, "issue", "single_unique_value"));
}
if (type.equals("integer") || type.equals("float")) {
List<Double> nums = new ArrayList<>();
for (String v : nonMissing) {
Double d = parseNumber(v);
if (d != null) nums.add(d);
}
nums.sort(Comparator.naturalOrder());
Double mu = mean(nums);
double sd = stddev(nums, mu);
List<Double> outliers = new ArrayList<>();
for (Double d : nums) {
if (sd > 0 && mu != null && Math.abs(d - mu) > 4 * sd) outliers.add(d);
}
if (!outliers.isEmpty()) {
quality.add("extreme_outliers");
issues.add(Map.of("column", col, "issue", "extreme_outliers", "count", outliers.size()));
}
Map<String, Object> ns = new LinkedHashMap<>();
ns.put("min", nums.isEmpty() ? null : nums.get(0));
ns.put("max", nums.isEmpty() ? null : nums.get(nums.size() - 1));
ns.put("mean", mu);
ns.put("median", percentile(nums, 0.5));
ns.put("stddev", sd);
ns.put("skewness", skewness(nums, mu, sd));
ns.put("percentiles", Map.of(
"p25", percentile(nums, 0.25),
"p50", percentile(nums, 0.50),
"p75", percentile(nums, 0.75),
"p95", percentile(nums, 0.95),
"p99", percentile(nums, 0.99)
));
ns.put("outliers_beyond_4std", outliers);
p.put("numeric_stats", ns);
numericColumns.add(col);
List<Double> series = new ArrayList<>();
for (String v : raw) series.add(isMissing(v) ? null : parseNumber(v));
numericSeries.put(col, series);
} else if (type.equals("string") || type.equals("categorical")) {
List<Integer> lengths = new ArrayList<>();
int numericLike = 0, dateLike = 0;
for (String v : nonMissing) {
lengths.add(v.length());
if (parseNumber(v) != null) numericLike++;
if (parseDateEpoch(v) != null) dateLike++;
}
if (!nonMissing.isEmpty() && (double) numericLike / nonMissing.size() >= 0.8) {
quality.add("string_looks_numeric");
issues.add(Map.of("column", col, "issue", "string_looks_numeric"));
}
if (!nonMissing.isEmpty() && (double) dateLike / nonMissing.size() >= 0.8) {
quality.add("string_looks_date");
issues.add(Map.of("column", col, "issue", "string_looks_date"));
}
Map<String, Object> ss = new LinkedHashMap<>();
ss.put("min_length", lengths.isEmpty() ? null : Collections.min(lengths));
ss.put("max_length", lengths.isEmpty() ? null : Collections.max(lengths));
ss.put("avg_length", lengths.isEmpty() ? null : lengths.stream().mapToInt(x -> x).average().orElse(0.0));
ss.put("unique_count", unique);
ss.put("top_values", topValues(nonMissing, 10));
p.put("string_stats", ss);
}
p.put("quality_issues", quality);
columnProfiles.put(col, p);
}
Map<String, Object> correlations = new LinkedHashMap<>();
for (String c1 : numericColumns) {
Map<String, Object> row = new LinkedHashMap<>();
for (String c2 : numericColumns) {
row.put(c2, c1.equals(c2) ? 1.0 : correlation(numericSeries.get(c1), numericSeries.get(c2)));
}
correlations.put(c1, row);
}
Map<String, Object> out = new LinkedHashMap<>();
out.put("row_count", rows.size());
out.put("column_count", columns.size());
out.put("columns", columnProfiles);
out.put("correlations", correlations);
out.put("issues", issues);
return out;
}
@SuppressWarnings("unchecked")
private static void printSummary(Map<String, Object> report) {
Map<String, Map<String, Object>> columns = (Map<String, Map<String, Object>>) report.get("columns");
List<String[]> rows = new ArrayList<>();
for (Map.Entry<String, Map<String, Object>> e : columns.entrySet()) {
String name = e.getKey();
Map<String, Object> info = e.getValue();
List<String> qi = (List<String>) info.get("quality_issues");
rows.add(new String[]{
name,
String.valueOf(info.get("type")),
String.valueOf(info.get("total_values")),
String.format("%.2f", ((Number) info.get("missing_percentage")).doubleValue()),
String.valueOf(info.get("unique_values")),
String.join(",", qi)
});
}
String[] header = {"Column", "Type", "Total", "Missing%", "Unique", "Notes"};
int[] widths = new int[header.length];
for (int i = 0; i < header.length; i++) widths[i] = header[i].length();
for (String[] row : rows) {
for (int i = 0; i < row.length; i++) widths[i] = Math.max(widths[i], row[i].length());
}
StringBuilder sep = new StringBuilder("+-");
for (int i = 0; i < widths.length; i++) {
sep.append("-".repeat(widths[i]));
sep.append(i + 1 < widths.length ? "-+-" : "-+");
}
System.out.println("Data Profiling Summary");
System.out.println("======================");
System.out.println("Rows: " + report.get("row_count"));
System.out.println("Columns: " + report.get("column_count"));
System.out.println(sep);
System.out.println(formatRow(header, widths));
System.out.println(sep);
for (String[] row : rows) System.out.println(formatRow(row, widths));
System.out.println(sep);
}
private static String formatRow(String[] row, int[] widths) {
StringBuilder sb = new StringBuilder("| ");
for (int i = 0; i < row.length; i++) {
sb.append(row[i]);
sb.append(" ".repeat(Math.max(0, widths[i] - row[i].length())));
sb.append(i + 1 < row.length ? " | " : " |");
}
return sb.toString();
}
private static List<Map<String, String>> generateSampleRows() {
List<Map<String, String>> rows = new ArrayList<>();
String[] countries = {"US", "CA", "GB", "DE", "IN"};
for (int i = 0; i < 1000; i++) {
int age = i % 200 == 0 ? 140 : 20 + (i % 45);
double income = i % 150 == 0 ? 250000 : 30000 + (i % 120) * 800 + (i % 7) * 13.5;
Map<String, String> row = new LinkedHashMap<>();
row.put("id", String.valueOf(i + 1));
row.put("age", String.valueOf(age));
row.put("income", String.format("%.2f", income));
row.put("is_active", i % 2 == 0 ? "true" : "false");
row.put("signup_date", String.format("2025-%02d-%02dT12:00:00Z", (i % 12) + 1, (i % 28) + 1));
row.put("country", countries[i % countries.length]);
row.put("status_code_str", String.valueOf(1000 + (i % 4)));
row.put("comment", i % 25 == 0 ? "" : "note_" + (i % 17));
row.put("constant_col", "CONST");
row.put("all_null_col", "");
rows.add(row);
}
return rows;
}
private static void writeSampleCsv(List<Map<String, String>> rows, Path file) throws IOException {
List<String> headers = new ArrayList<>(rows.get(0).keySet());
try (BufferedWriter writer = Files.newBufferedWriter(file, StandardCharsets.UTF_8)) {
writer.write(String.join(",", headers));
writer.newLine();
for (Map<String, String> row : rows) {
List<String> vals = new ArrayList<>();
for (String h : headers) {
String value = row.getOrDefault(h, "");
if (value.contains(",") || value.contains("\"") || value.contains("\n")) {
vals.add("\"" + value.replace("\"", "\"\"") + "\"");
} else {
vals.add(value);
}
}
writer.write(String.join(",", vals));
writer.newLine();
}
}
}
}