← All tasks
javacodex/java-t1 #7Not a task: already works

Data Profiling Tool (java, written by Codex)

envgap__codex__java-t1-7

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
pom.xml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/java-t1 #7 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Data Profiling Tool

Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV or JSON data file path as a command-line argument
- Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings)
- For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th)
- For string columns: compute min/max/average length, most common values (top 10), and unique count
- For all columns: count total values, missing/null values, missing percentage, and unique value count
- Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates
- Compute a pairwise correlation matrix for all numeric columns
- Print a formatted summary report to console showing key statistics per column
- Save the full profiling report as a JSON file with --output flag (default: data_profile.json)
- If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it
- Handle files with inconsistent delimiters or encoding issues gracefully

Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include:
- Source code
- pom.xml with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
  <modelVersion>4.0.0</modelVersion>

  <groupId>tmlr.codex_generated.p07</groupId>
  <artifactId>data-profiling-tool</artifactId>
  <version>1.0.0</version>
  <name>Data Profiling Tool</name>

  <properties>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
    <maven.compiler.release>17</maven.compiler.release>
  </properties>

  <dependencies>
    <dependency>
      <groupId>com.fasterxml.jackson.core</groupId>
      <artifactId>jackson-databind</artifactId>
      <version>2.18.2</version>
    </dependency>
    <dependency>
      <groupId>com.fasterxml.jackson.core</groupId>
      <artifactId>jackson-core</artifactId>
      <version>2.18.2</version>
    </dependency>
    <dependency>
      <groupId>com.fasterxml.jackson.core</groupId>
      <artifactId>jackson-annotations</artifactId>
      <version>2.18.2</version>
    </dependency>
  </dependencies>

  <build>
    <plugins>
      <plugin>
        <groupId>org.apache.maven.plugins</groupId>
        <artifactId>maven-compiler-plugin</artifactId>
        <version>3.13.0</version>
      </plugin>
      <plugin>
        <groupId>org.codehaus.mojo</groupId>
        <artifactId>exec-maven-plugin</artifactId>
        <version>3.5.0</version>
        <configuration>
          <mainClass>DataProfilingTool</mainClass>
        </configuration>
      </plugin>
    </plugins>
  </build>
</project>
README.md
# Data Profiling Tool (Java)

Profiles CSV/JSON datasets for column typing, missingness, distributions, outliers, correlations, and quality issues.

## Requirements

- Ubuntu 22.04
- JDK 17+
- Maven 3.9+

## Dependencies

- Direct:
  - `com.fasterxml.jackson.core:jackson-databind:2.18.2`
  - `com.fasterxml.jackson.core:jackson-core:2.18.2`
  - `com.fasterxml.jackson.core:jackson-annotations:2.18.2`
- Transitive:
  - none (explicitly pinned above)

All versions are pinned in `pom.xml`.

## Build

```bash
mvn -q -DskipTests compile
```

## Run

With input:

```bash
mvn -q exec:java -Dexec.args="/path/to/data.csv --output data_profile.json"
```

JSON input:

```bash
mvn -q exec:java -Dexec.args="/path/to/data.json --output data_profile.json"
```

No input (generates sample dataset with intentional quality issues):

```bash
mvn -q exec:java
```

## Output

- Console summary table per column
- Full JSON report (`data_profile.json` by default)
src/main/java/DataProfilingTool.java
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;

import java.io.BufferedWriter;
import java.io.IOException;
import java.nio.ByteBuffer;
import java.nio.charset.CharacterCodingException;
import java.nio.charset.CharsetDecoder;
import java.nio.charset.CodingErrorAction;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.time.LocalDate;
import java.time.LocalDateTime;
import java.time.OffsetDateTime;
import java.time.format.DateTimeFormatter;
import java.time.format.DateTimeParseException;
import java.util.ArrayList;
import java.util.Collections;
import java.util.Comparator;
import java.util.HashMap;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;

public class DataProfilingTool {
    private static final ObjectMapper MAPPER = new ObjectMapper();

    public static void main(String[] args) throws Exception {
        ParsedArgs parsed = parseArgs(args);
        String output = parsed.options.getOrDefault("output", "data_profile.json");
        Path outputPath = Paths.get(output).toAbsolutePath();

        List<Map<String, String>> rows;
        Map<String, Object> metadata = new LinkedHashMap<>();
        if (parsed.positional.isEmpty()) {
            rows = generateSampleRows();
            Path samplePath = Paths.get("sample_profile_data.csv").toAbsolutePath();
            writeSampleCsv(rows, samplePath);
            metadata.put("generated_sample", samplePath.toString());
            System.out.println("No input file provided. Generated sample dataset: " + samplePath);
        } else {
            Path inputPath = Paths.get(parsed.positional.get(0)).toAbsolutePath();
            if (!Files.exists(inputPath)) {
                System.err.println("Input file not found: " + inputPath);
                System.exit(1);
                return;
            }
            String ext = extension(inputPath);
            if (ext.equals("json") || ext.equals("jsonl") || ext.equals("ndjson")) {
                LoadResult load = loadJson(inputPath);
                rows = load.rows;
                metadata.put("input_format", "json");
                metadata.put("encoding", load.encoding);
            } else {
                LoadResult load = loadCsv(inputPath);
                rows = load.rows;
                metadata.put("input_format", "csv");
                metadata.put("encoding", load.encoding);
                metadata.put("delimiter", load.delimiter);
            }
            metadata.put("input_file", inputPath.toString());
        }

        Map<String, Object> report = profile(rows);
        report.put("generated_at", Instant.now().toString());
        for (Map.Entry<String, Object> e : metadata.entrySet()) {
            report.put(e.getKey(), e.getValue());
        }

        printSummary(report);
        Files.writeString(outputPath, MAPPER.writerWithDefaultPrettyPrinter().writeValueAsString(report) + "\n", StandardCharsets.UTF_8);
        System.out.println("Saved JSON profile: " + outputPath);
    }

    private static String extension(Path path) {
        String name = path.getFileName().toString();
        int dot = name.lastIndexOf('.');
        return dot >= 0 ? name.substring(dot + 1).toLowerCase() : "";
    }

    private record ParsedArgs(Map<String, String> options, List<String> positional) {}
    private record LoadResult(List<Map<String, String>> rows, String encoding, String delimiter) {}

    private static ParsedArgs parseArgs(String[] args) {
        Map<String, String> options = new HashMap<>();
        List<String> positional = new ArrayList<>();
        for (int i = 0; i < args.length; i++) {
            String token = args[i];
            if (token.startsWith("--")) {
                String key = token.substring(2);
                if (i + 1 < args.length && !args[i + 1].startsWith("--")) {
                    options.put(key, args[++i]);
                } else {
                    options.put(key, "true");
                }
            } else {
                positional.add(token);
            }
        }
        return new ParsedArgs(options, positional);
    }

    private static String decodeUtf8ThenLatin1(byte[] bytes) {
        CharsetDecoder dec = StandardCharsets.UTF_8.newDecoder();
        dec.onMalformedInput(CodingErrorAction.REPORT);
        dec.onUnmappableCharacter(CodingErrorAction.REPORT);
        try {
            return dec.decode(ByteBuffer.wrap(bytes)).toString();
        } catch (CharacterCodingException e) {
            return new String(bytes, StandardCharsets.ISO_8859_1);
        }
    }

    private static String detectEncoding(byte[] bytes) {
        CharsetDecoder dec = StandardCharsets.UTF_8.newDecoder();
        dec.onMalformedInput(CodingErrorAction.REPORT);
        dec.onUnmappableCharacter(CodingErrorAction.REPORT);
        try {
            dec.decode(ByteBuffer.wrap(bytes));
            return "utf-8";
        } catch (CharacterCodingException e) {
            return "latin1";
        }
    }

    private static List<String> parseCsvLine(String line, char delimiter) {
        List<String> out = new ArrayList<>();
        StringBuilder current = new StringBuilder();
        boolean inQuotes = false;
        for (int i = 0; i < line.length(); i++) {
            char ch = line.charAt(i);
            if (ch == '"') {
                if (inQuotes && i + 1 < line.length() && line.charAt(i + 1) == '"') {
                    current.append('"');
                    i++;
                } else {
                    inQuotes = !inQuotes;
                }
            } else if (ch == delimiter && !inQuotes) {
                out.add(current.toString());
                current.setLength(0);
            } else {
                current.append(ch);
            }
        }
        out.add(current.toString());
        return out;
    }

    private static char detectDelimiter(List<String> lines) {
        char[] cands = new char[]{',', ';', '\t', '|'};
        char best = ',';
        int bestScore = -1;
        for (char c : cands) {
            int score = 0;
            for (String line : lines) {
                for (int i = 0; i < line.length(); i++) if (line.charAt(i) == c) score++;
            }
            if (score > bestScore) {
                bestScore = score;
                best = c;
            }
        }
        return best;
    }

    private static LoadResult loadCsv(Path path) throws IOException {
        byte[] bytes = Files.readAllBytes(path);
        String encoding = detectEncoding(bytes);
        String text = decodeUtf8ThenLatin1(bytes);
        List<String> lines = new ArrayList<>();
        for (String line : text.replace("\r\n", "\n").replace("\r", "\n").split("\n")) {
            if (!line.trim().isEmpty()) lines.add(line);
        }
        if (lines.isEmpty()) return new LoadResult(new ArrayList<>(), encoding, ",");

        char delimiter = detectDelimiter(lines.subList(0, Math.min(5, lines.size())));
        List<String> headers = parseCsvLine(lines.get(0), delimiter);
        List<Map<String, String>> rows = new ArrayList<>();
        for (int i = 1; i < lines.size(); i++) {
            List<String> values = parseCsvLine(lines.get(i), delimiter);
            Map<String, String> row = new LinkedHashMap<>();
            for (int j = 0; j < headers.size(); j++) {
                row.put(headers.get(j).trim(), j < values.size() ? values.get(j) : "");
            }
            rows.add(row);
        }
        return new LoadResult(rows, encoding, String.valueOf(delimiter));
    }

    private static LoadResult loadJson(Path path) throws IOException {
        byte[] bytes = Files.readAllBytes(path);
        String encoding = detectEncoding(bytes);
        String text = decodeUtf8ThenLatin1(bytes);
        List<Map<String, String>> rows = new ArrayList<>();
        try {
            JsonNode node = MAPPER.readTree(text);
            if (node != null && node.isArray()) {
                for (JsonNode item : node) {
                    if (item.isObject()) rows.add(nodeToFlatMap(item));
                }
            } else if (node != null && node.isObject()) {
                rows.add(nodeToFlatMap(node));
            }
        } catch (Exception e) {
            for (String line : text.split("\\R")) {
                if (line.trim().isEmpty()) continue;
                JsonNode item = MAPPER.readTree(line);
                if (item.isObject()) rows.add(nodeToFlatMap(item));
            }
        }
        return new LoadResult(rows, encoding, null);
    }

    private static Map<String, String> nodeToFlatMap(JsonNode node) {
        Map<String, String> out = new LinkedHashMap<>();
        node.fields().forEachRemaining(e -> out.put(e.getKey(), e.getValue().isNull() ? "" : e.getValue().asText()));
        return out;
    }

    private static boolean isMissing(String v) {
        if (v == null) return true;
        String s = v.trim().toLowerCase();
        return s.isEmpty() || s.equals("null") || s.equals("na") || s.equals("n/a") || s.equals("none");
    }

    private static Double parseNumber(String v) {
        if (v == null) return null;
        String s = v.trim();
        if (s.isEmpty()) return null;
        if (!s.matches("^[-+]?\\d+(\\.\\d+)?$")) return null;
        try {
            return Double.parseDouble(s);
        } catch (NumberFormatException e) {
            return null;
        }
    }

    private static Boolean parseBoolean(String v) {
        if (v == null) return null;
        String s = v.trim().toLowerCase();
        if (Set.of("true", "1", "yes", "y").contains(s)) return true;
        if (Set.of("false", "0", "no", "n").contains(s)) return false;
        return null;
    }

    private static Double parseDateEpoch(String v) {
        if (v == null) return null;
        String s = v.trim();
        if (s.isEmpty()) return null;
        try { return (double) Instant.parse(s).getEpochSecond(); } catch (DateTimeParseException ignored) {}
        try { return (double) OffsetDateTime.parse(s).toEpochSecond(); } catch (DateTimeParseException ignored) {}
        try { return (double) LocalDateTime.parse(s, DateTimeFormatter.ISO_LOCAL_DATE_TIME).toEpochSecond(java.time.ZoneOffset.UTC); } catch (DateTimeParseException ignored) {}
        try { return (double) LocalDate.parse(s, DateTimeFormatter.ISO_LOCAL_DATE).atStartOfDay(java.time.ZoneOffset.UTC).toEpochSecond(); } catch (DateTimeParseException ignored) {}
        return null;
    }

    private static String detectType(List<String> values) {
        if (values.isEmpty()) return "string";
        int boolCount = 0, numCount = 0, dateCount = 0;
        for (String v : values) {
            if (parseBoolean(v) != null) boolCount++;
            if (parseNumber(v) != null) numCount++;
            if (parseDateEpoch(v) != null) dateCount++;
        }
        if (boolCount == values.size()) return "boolean";
        if (numCount == values.size()) {
            boolean hasFloat = values.stream().anyMatch(s -> s.contains("."));
            return hasFloat ? "float" : "integer";
        }
        if (dateCount >= Math.max(3, (int) Math.floor(values.size() * 0.9))) return "date";
        Set<String> unique = new HashSet<>(values);
        double ratio = values.isEmpty() ? 0.0 : (double) unique.size() / values.size();
        if (unique.size() <= 20 || ratio <= 0.1) return "categorical";
        return "string";
    }

    private static Double mean(List<Double> values) {
        if (values.isEmpty()) return null;
        double sum = 0.0;
        for (Double v : values) sum += v;
        return sum / values.size();
    }

    private static double stddev(List<Double> values, Double mean) {
        if (values.size() <= 1) return 0.0;
        double mu = mean != null ? mean : mean(values);
        double acc = 0.0;
        for (Double v : values) acc += (v - mu) * (v - mu);
        return Math.sqrt(acc / values.size());
    }

    private static double skewness(List<Double> values, Double mean, Double sd) {
        if (values.size() < 3) return 0.0;
        double mu = mean != null ? mean : mean(values);
        double sigma = sd != null ? sd : stddev(values, mu);
        if (sigma == 0.0) return 0.0;
        double m3 = 0.0;
        for (Double v : values) m3 += Math.pow(v - mu, 3);
        m3 /= values.size();
        return m3 / Math.pow(sigma, 3);
    }

    private static Double percentile(List<Double> sorted, double p) {
        if (sorted.isEmpty()) return null;
        if (sorted.size() == 1) return sorted.get(0);
        double pos = (sorted.size() - 1) * p;
        int lo = (int) Math.floor(pos);
        int hi = (int) Math.ceil(pos);
        if (lo == hi) return sorted.get(lo);
        double w = pos - lo;
        return sorted.get(lo) + (sorted.get(hi) - sorted.get(lo)) * w;
    }

    private static Double correlation(List<Double> xs, List<Double> ys) {
        List<Double> x = new ArrayList<>();
        List<Double> y = new ArrayList<>();
        int n = Math.min(xs.size(), ys.size());
        for (int i = 0; i < n; i++) {
            if (xs.get(i) == null || ys.get(i) == null) continue;
            x.add(xs.get(i));
            y.add(ys.get(i));
        }
        if (x.size() < 2) return null;
        Double mx = mean(x), my = mean(y);
        double sx = stddev(x, mx), sy = stddev(y, my);
        if (sx == 0.0 || sy == 0.0) return 0.0;
        double cov = 0.0;
        for (int i = 0; i < x.size(); i++) cov += (x.get(i) - mx) * (y.get(i) - my);
        cov /= x.size();
        return cov / (sx * sy);
    }

    private static List<Map<String, String>> topValues(List<String> values, int n) {
        Map<String, Integer> freq = new HashMap<>();
        for (String v : values) freq.put(v, freq.getOrDefault(v, 0) + 1);
        List<Map.Entry<String, Integer>> items = new ArrayList<>(freq.entrySet());
        items.sort((a, b) -> {
            int c = Integer.compare(b.getValue(), a.getValue());
            return c != 0 ? c : a.getKey().compareTo(b.getKey());
        });
        List<Map<String, String>> out = new ArrayList<>();
        for (int i = 0; i < Math.min(n, items.size()); i++) {
            Map<String, String> row = new LinkedHashMap<>();
            row.put("value", items.get(i).getKey());
            row.put("count", String.valueOf(items.get(i).getValue()));
            out.add(row);
        }
        return out;
    }

    private static Map<String, Object> profile(List<Map<String, String>> rows) {
        Set<String> colSet = new HashSet<>();
        for (Map<String, String> row : rows) colSet.addAll(row.keySet());
        List<String> columns = new ArrayList<>(colSet);
        Collections.sort(columns);

        Map<String, Object> columnProfiles = new LinkedHashMap<>();
        List<Map<String, Object>> issues = new ArrayList<>();
        List<String> numericColumns = new ArrayList<>();
        Map<String, List<Double>> numericSeries = new HashMap<>();

        for (String col : columns) {
            List<String> raw = new ArrayList<>();
            for (Map<String, String> row : rows) raw.add(row.get(col));
            int total = raw.size();
            int missing = 0;
            List<String> nonMissing = new ArrayList<>();
            for (String v : raw) {
                if (isMissing(v)) missing++;
                else nonMissing.add(v);
            }
            int unique = new HashSet<>(nonMissing).size();
            String type = detectType(nonMissing);
            double missingPct = total == 0 ? 0.0 : ((double) missing / total) * 100.0;

            Map<String, Object> p = new LinkedHashMap<>();
            p.put("type", type);
            p.put("total_values", total);
            p.put("missing_values", missing);
            p.put("missing_percentage", Math.round(missingPct * 10000.0) / 10000.0);
            p.put("unique_values", unique);
            List<String> quality = new ArrayList<>();

            if (missing == total) {
                quality.add("entirely_null");
                issues.add(Map.of("column", col, "issue", "entirely_null"));
            }
            if (unique == 1 && !nonMissing.isEmpty()) {
                quality.add("single_unique_value");
                issues.add(Map.of("column", col, "issue", "single_unique_value"));
            }

            if (type.equals("integer") || type.equals("float")) {
                List<Double> nums = new ArrayList<>();
                for (String v : nonMissing) {
                    Double d = parseNumber(v);
                    if (d != null) nums.add(d);
                }
                nums.sort(Comparator.naturalOrder());
                Double mu = mean(nums);
                double sd = stddev(nums, mu);
                List<Double> outliers = new ArrayList<>();
                for (Double d : nums) {
                    if (sd > 0 && mu != null && Math.abs(d - mu) > 4 * sd) outliers.add(d);
                }
                if (!outliers.isEmpty()) {
                    quality.add("extreme_outliers");
                    issues.add(Map.of("column", col, "issue", "extreme_outliers", "count", outliers.size()));
                }
                Map<String, Object> ns = new LinkedHashMap<>();
                ns.put("min", nums.isEmpty() ? null : nums.get(0));
                ns.put("max", nums.isEmpty() ? null : nums.get(nums.size() - 1));
                ns.put("mean", mu);
                ns.put("median", percentile(nums, 0.5));
                ns.put("stddev", sd);
                ns.put("skewness", skewness(nums, mu, sd));
                ns.put("percentiles", Map.of(
                        "p25", percentile(nums, 0.25),
                        "p50", percentile(nums, 0.50),
                        "p75", percentile(nums, 0.75),
                        "p95", percentile(nums, 0.95),
                        "p99", percentile(nums, 0.99)
                ));
                ns.put("outliers_beyond_4std", outliers);
                p.put("numeric_stats", ns);

                numericColumns.add(col);
                List<Double> series = new ArrayList<>();
                for (String v : raw) series.add(isMissing(v) ? null : parseNumber(v));
                numericSeries.put(col, series);
            } else if (type.equals("string") || type.equals("categorical")) {
                List<Integer> lengths = new ArrayList<>();
                int numericLike = 0, dateLike = 0;
                for (String v : nonMissing) {
                    lengths.add(v.length());
                    if (parseNumber(v) != null) numericLike++;
                    if (parseDateEpoch(v) != null) dateLike++;
                }
                if (!nonMissing.isEmpty() && (double) numericLike / nonMissing.size() >= 0.8) {
                    quality.add("string_looks_numeric");
                    issues.add(Map.of("column", col, "issue", "string_looks_numeric"));
                }
                if (!nonMissing.isEmpty() && (double) dateLike / nonMissing.size() >= 0.8) {
                    quality.add("string_looks_date");
                    issues.add(Map.of("column", col, "issue", "string_looks_date"));
                }
                Map<String, Object> ss = new LinkedHashMap<>();
                ss.put("min_length", lengths.isEmpty() ? null : Collections.min(lengths));
                ss.put("max_length", lengths.isEmpty() ? null : Collections.max(lengths));
                ss.put("avg_length", lengths.isEmpty() ? null : lengths.stream().mapToInt(x -> x).average().orElse(0.0));
                ss.put("unique_count", unique);
                ss.put("top_values", topValues(nonMissing, 10));
                p.put("string_stats", ss);
            }

            p.put("quality_issues", quality);
            columnProfiles.put(col, p);
        }

        Map<String, Object> correlations = new LinkedHashMap<>();
        for (String c1 : numericColumns) {
            Map<String, Object> row = new LinkedHashMap<>();
            for (String c2 : numericColumns) {
                row.put(c2, c1.equals(c2) ? 1.0 : correlation(numericSeries.get(c1), numericSeries.get(c2)));
            }
            correlations.put(c1, row);
        }

        Map<String, Object> out = new LinkedHashMap<>();
        out.put("row_count", rows.size());
        out.put("column_count", columns.size());
        out.put("columns", columnProfiles);
        out.put("correlations", correlations);
        out.put("issues", issues);
        return out;
    }

    @SuppressWarnings("unchecked")
    private static void printSummary(Map<String, Object> report) {
        Map<String, Map<String, Object>> columns = (Map<String, Map<String, Object>>) report.get("columns");
        List<String[]> rows = new ArrayList<>();
        for (Map.Entry<String, Map<String, Object>> e : columns.entrySet()) {
            String name = e.getKey();
            Map<String, Object> info = e.getValue();
            List<String> qi = (List<String>) info.get("quality_issues");
            rows.add(new String[]{
                    name,
                    String.valueOf(info.get("type")),
                    String.valueOf(info.get("total_values")),
                    String.format("%.2f", ((Number) info.get("missing_percentage")).doubleValue()),
                    String.valueOf(info.get("unique_values")),
                    String.join(",", qi)
            });
        }
        String[] header = {"Column", "Type", "Total", "Missing%", "Unique", "Notes"};
        int[] widths = new int[header.length];
        for (int i = 0; i < header.length; i++) widths[i] = header[i].length();
        for (String[] row : rows) {
            for (int i = 0; i < row.length; i++) widths[i] = Math.max(widths[i], row[i].length());
        }

        StringBuilder sep = new StringBuilder("+-");
        for (int i = 0; i < widths.length; i++) {
            sep.append("-".repeat(widths[i]));
            sep.append(i + 1 < widths.length ? "-+-" : "-+");
        }

        System.out.println("Data Profiling Summary");
        System.out.println("======================");
        System.out.println("Rows: " + report.get("row_count"));
        System.out.println("Columns: " + report.get("column_count"));
        System.out.println(sep);
        System.out.println(formatRow(header, widths));
        System.out.println(sep);
        for (String[] row : rows) System.out.println(formatRow(row, widths));
        System.out.println(sep);
    }

    private static String formatRow(String[] row, int[] widths) {
        StringBuilder sb = new StringBuilder("| ");
        for (int i = 0; i < row.length; i++) {
            sb.append(row[i]);
            sb.append(" ".repeat(Math.max(0, widths[i] - row[i].length())));
            sb.append(i + 1 < row.length ? " | " : " |");
        }
        return sb.toString();
    }

    private static List<Map<String, String>> generateSampleRows() {
        List<Map<String, String>> rows = new ArrayList<>();
        String[] countries = {"US", "CA", "GB", "DE", "IN"};
        for (int i = 0; i < 1000; i++) {
            int age = i % 200 == 0 ? 140 : 20 + (i % 45);
            double income = i % 150 == 0 ? 250000 : 30000 + (i % 120) * 800 + (i % 7) * 13.5;
            Map<String, String> row = new LinkedHashMap<>();
            row.put("id", String.valueOf(i + 1));
            row.put("age", String.valueOf(age));
            row.put("income", String.format("%.2f", income));
            row.put("is_active", i % 2 == 0 ? "true" : "false");
            row.put("signup_date", String.format("2025-%02d-%02dT12:00:00Z", (i % 12) + 1, (i % 28) + 1));
            row.put("country", countries[i % countries.length]);
            row.put("status_code_str", String.valueOf(1000 + (i % 4)));
            row.put("comment", i % 25 == 0 ? "" : "note_" + (i % 17));
            row.put("constant_col", "CONST");
            row.put("all_null_col", "");
            rows.add(row);
        }
        return rows;
    }

    private static void writeSampleCsv(List<Map<String, String>> rows, Path file) throws IOException {
        List<String> headers = new ArrayList<>(rows.get(0).keySet());
        try (BufferedWriter writer = Files.newBufferedWriter(file, StandardCharsets.UTF_8)) {
            writer.write(String.join(",", headers));
            writer.newLine();
            for (Map<String, String> row : rows) {
                List<String> vals = new ArrayList<>();
                for (String h : headers) {
                    String value = row.getOrDefault(h, "");
                    if (value.contains(",") || value.contains("\"") || value.contains("\n")) {
                        vals.add("\"" + value.replace("\"", "\"\"") + "\"");
                    } else {
                        vals.add(value);
                    }
                }
                writer.write(String.join(",", vals));
                writer.newLine();
            }
        }
    }
}