← All tasks
javaclaude-code/java-t1 #7Not a task: repair changed code

Data Profiling Tool (java, written by Claude Code)

envgap__claude-code__java-t1-7

Written by a coding agent; not on GitHubWritten 2026-02-27

01 / FAILURE SIGNATURE

As the study recorded it

NullPointerException: Cannot invoke java.lang.Integer.intValue() - null unboxing in ternary expression
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
pom.xml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/java-t1 #7 · read the task the agent was given
Claude Code wrote this java project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Data Profiling Tool

Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV or JSON data file path as a command-line argument
- Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings)
- For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th)
- For string columns: compute min/max/average length, most common values (top 10), and unique count
- For all columns: count total values, missing/null values, missing percentage, and unique value count
- Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates
- Compute a pairwise correlation matrix for all numeric columns
- Print a formatted summary report to console showing key statistics per column
- Save the full profiling report as a JSON file with --output flag (default: data_profile.json)
- If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it
- Handle files with inconsistent delimiters or encoding issues gracefully

Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include:
- Source code
- pom.xml with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

2 files, exactly as written, before any repair.

pom.xml
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
    <modelVersion>4.0.0</modelVersion>

    <groupId>com.profiler</groupId>
    <artifactId>data-profiler-trial1</artifactId>
    <version>1.0.0</version>
    <packaging>jar</packaging>

    <properties>
        <maven.compiler.source>11</maven.compiler.source>
        <maven.compiler.target>11</maven.compiler.target>
        <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
    </properties>

    <dependencies>
        <dependency>
            <groupId>org.apache.commons</groupId>
            <artifactId>commons-csv</artifactId>
            <version>1.11.0</version>
        </dependency>
        <dependency>
            <groupId>org.apache.commons</groupId>
            <artifactId>commons-math3</artifactId>
            <version>3.6.1</version>
        </dependency>
        <dependency>
            <groupId>com.google.code.gson</groupId>
            <artifactId>gson</artifactId>
            <version>2.11.0</version>
        </dependency>
    </dependencies>

    <build>
        <plugins>
            <plugin>
                <groupId>org.apache.maven.plugins</groupId>
                <artifactId>maven-jar-plugin</artifactId>
                <version>3.4.1</version>
                <configuration>
                    <archive>
                        <manifest>
                            <mainClass>profiler.DataProfiler</mainClass>
                        </manifest>
                    </archive>
                </configuration>
            </plugin>
            <plugin>
                <groupId>org.apache.maven.plugins</groupId>
                <artifactId>maven-shade-plugin</artifactId>
                <version>3.5.3</version>
                <executions>
                    <execution>
                        <phase>package</phase>
                        <goals><goal>shade</goal></goals>
                    </execution>
                </executions>
            </plugin>
        </plugins>
    </build>
</project>
src/main/java/profiler/DataProfiler.java
package profiler;

import org.apache.commons.csv.CSVFormat;
import org.apache.commons.csv.CSVParser;
import org.apache.commons.csv.CSVRecord;
import org.apache.commons.math3.stat.descriptive.DescriptiveStatistics;
import org.apache.commons.math3.stat.correlation.PearsonsCorrelation;
import com.google.gson.Gson;
import com.google.gson.GsonBuilder;

import java.io.*;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.time.LocalDateTime;
import java.time.format.DateTimeFormatter;
import java.util.*;
import java.util.stream.Collectors;

/**
 * Data Profiling Tool - Trial 1
 * Dependencies: Apache Commons CSV, Commons Math3, Gson
 */
public class DataProfiler {

    private static final Set<String> NULL_VALUES = new HashSet<>(
            Arrays.asList("", "null", "NULL", "None", "NA", "N/A", "nan", "NaN", "undefined"));

    public static void main(String[] args) {
        String filePath;
        if (args.length > 0) {
            filePath = args[0];
        } else {
            System.out.println("[INFO] No input file provided. Generating sample dataset...");
            filePath = "sample_data.csv";
            generateSampleDataset(filePath);
        }

        File file = new File(filePath);
        if (!file.exists()) {
            System.err.println("[ERROR] File not found: " + filePath);
            System.exit(1);
        }

        try {
            profileDataset(filePath);
        } catch (Exception e) {
            System.err.println("[ERROR] " + e.getMessage());
            e.printStackTrace();
            System.exit(1);
        }
    }

    static void generateSampleDataset(String filePath) {
        Random rng = new Random(42);
        String[] names = {"Alice", "Bob", "Charlie", "Diana", "Eve", "", "Frank", "Grace", null, "Heidi"};
        String[] departments = {"Engineering", "Sales", "HR", "Marketing", null, "Engineering", "Sales", ""};
        String[] dates = {"2020-01-15", "2021-06-30", "2019-12-01", null, "invalid-date", "2022-03-10"};
        String[] actives = {"true", "false", null, "yes", "no", "1", "0"};

        try (PrintWriter writer = new PrintWriter(new FileWriter(filePath))) {
            writer.println("id,name,age,salary,department,score,join_date,is_active");
            for (int i = 0; i < 200; i++) {
                String name = names[rng.nextInt(names.length)];
                Integer age = rng.nextBoolean() ? rng.nextInt(62) + 18 : (rng.nextBoolean() ? null : (rng.nextBoolean() ? -1 : 999));
                Double salary = rng.nextBoolean() ? rng.nextDouble() * 125000 + 25000 : (rng.nextBoolean() ? null : (rng.nextBoolean() ? 0.0 : -500.0));
                String dept = departments[rng.nextInt(departments.length)];
                Double score = rng.nextBoolean() ? rng.nextDouble() * 100 : null;
                String date = dates[rng.nextInt(dates.length)];
                String active = actives[rng.nextInt(actives.length)];

                writer.printf("%d,%s,%s,%s,%s,%s,%s,%s%n",
                        i + 1,
                        name != null ? name : "",
                        age != null ? age.toString() : "",
                        salary != null ? String.format("%.2f", salary) : "",
                        dept != null ? dept : "",
                        score != null ? String.format("%.1f", score) : "",
                        date != null ? date : "",
                        active != null ? active : "");
            }
            System.out.println("[INFO] Generated sample dataset: " + filePath + " (200 rows)");
        } catch (IOException e) {
            System.err.println("[ERROR] Failed to generate sample: " + e.getMessage());
        }
    }

    static void profileDataset(String filePath) throws Exception {
        System.out.println("\n" + "=".repeat(60));
        System.out.println("  DATA PROFILING TOOL - Trial 1 (Commons CSV + Math3 + Gson)");
        System.out.println("=".repeat(60) + "\n");

        // Load data
        List<Map<String, String>> data = loadCsv(filePath);
        List<String> columns = new ArrayList<>(data.get(0).keySet());

        System.out.println("[INFO] Loaded file: " + filePath);
        System.out.println("[INFO] Shape: " + data.size() + " rows x " + columns.size() + " columns\n");

        // Detect types
        Map<String, String> colTypes = new LinkedHashMap<>();
        for (String col : columns) {
            List<String> values = data.stream().map(r -> r.get(col)).collect(Collectors.toList());
            colTypes.put(col, detectColumnType(values));
        }

        // Profile columns
        Map<String, Map<String, Object>> columnProfiles = new LinkedHashMap<>();
        for (String col : columns) {
            List<String> values = data.stream().map(r -> r.get(col)).collect(Collectors.toList());
            columnProfiles.put(col, profileColumn(col, values, colTypes.get(col)));
        }

        // Correlation matrix
        Map<String, Map<String, Double>> corrMatrix = computeCorrelationMatrix(data, columns, colTypes);

        // Data quality score
        double qualityScore = computeDataQualityScore(data, columns);

        // Total missing
        int totalMissing = 0;
        for (String col : columns) {
            for (Map<String, String> row : data) {
                if (isNull(row.get(col))) totalMissing++;
            }
        }
        int totalCells = data.size() * columns.size();

        // Build report
        Map<String, Object> report = new LinkedHashMap<>();
        report.put("file", new File(filePath).getName());
        report.put("generated_at", LocalDateTime.now().format(DateTimeFormatter.ISO_LOCAL_DATE_TIME));

        Map<String, Object> overview = new LinkedHashMap<>();
        overview.put("rows", data.size());
        overview.put("columns", columns.size());
        overview.put("total_cells", totalCells);
        overview.put("total_missing", totalMissing);
        overview.put("total_missing_pct", round((double) totalMissing / totalCells * 100, 2));
        report.put("dataset_overview", overview);
        report.put("data_quality_score", qualityScore);
        report.put("column_profiles", columnProfiles);
        report.put("correlation_matrix", corrMatrix);

        // Console output
        printConsoleReport(report, columns);

        // Save JSON
        Gson gson = new GsonBuilder().setPrettyPrinting().serializeNulls().create();
        String outputDir = new File(filePath).getAbsoluteFile().getParent();
        String reportPath = outputDir + File.separator + "profile_report.json";
        try (Writer writer = new FileWriter(reportPath)) {
            gson.toJson(report, writer);
        }
        System.out.println("\n[INFO] Profile report saved: " + reportPath);
    }

    static List<Map<String, String>> loadCsv(String filePath) throws Exception {
        String ext = filePath.substring(filePath.lastIndexOf('.')).toLowerCase();
        List<Map<String, String>> data = new ArrayList<>();

        if (ext.equals(".csv")) {
            try (Reader reader = new FileReader(filePath, StandardCharsets.UTF_8);
                 CSVParser parser = CSVFormat.DEFAULT.builder()
                         .setHeader()
                         .setSkipHeaderRecord(true)
                         .setTrim(true)
                         .setIgnoreEmptyLines(true)
                         .build()
                         .parse(reader)) {

                for (CSVRecord record : parser) {
                    Map<String, String> row = new LinkedHashMap<>();
                    for (String header : parser.getHeaderNames()) {
                        row.put(header, record.get(header));
                    }
                    data.add(row);
                }
            }
        } else if (ext.equals(".json")) {
            String content = new String(Files.readAllBytes(Paths.get(filePath)), StandardCharsets.UTF_8);
            Gson gson = new Gson();
            List<Map<String, Object>> jsonData = gson.fromJson(content,
                    new com.google.gson.reflect.TypeToken<List<Map<String, Object>>>() {}.getType());
            for (Map<String, Object> item : jsonData) {
                Map<String, String> row = new LinkedHashMap<>();
                for (Map.Entry<String, Object> entry : item.entrySet()) {
                    row.put(entry.getKey(), entry.getValue() != null ? entry.getValue().toString() : "");
                }
                data.add(row);
            }
        } else {
            throw new IllegalArgumentException("Unsupported format: " + ext);
        }

        return data;
    }

    static boolean isNull(String val) {
        return val == null || NULL_VALUES.contains(val.trim());
    }

    static List<String> cleanValues(List<String> values) {
        return values.stream().filter(v -> !isNull(v)).collect(Collectors.toList());
    }

    static String detectColumnType(List<String> values) {
        List<String> clean = cleanValues(values);
        if (clean.isEmpty()) return "empty";

        int numericCount = 0;
        boolean allInts = true;
        for (String v : clean) {
            try {
                double d = Double.parseDouble(v.trim());
                numericCount++;
                if (d != Math.floor(d) || v.contains(".")) allInts = false;
            } catch (NumberFormatException e) {
                allInts = false;
            }
        }

        if ((double) numericCount / clean.size() > 0.8) {
            if (numericCount == clean.size() && allInts) return "integer";
            if (numericCount == clean.size()) return "float";
            return "numeric_mixed";
        }

        Set<String> boolVals = new HashSet<>(Arrays.asList("true", "false", "yes", "no", "1", "0", "y", "n"));
        boolean allBool = clean.stream().allMatch(v -> boolVals.contains(v.toLowerCase().trim()));
        if (allBool) return "boolean";

        Set<String> unique = new HashSet<>(clean);
        if ((double) unique.size() / clean.size() < 0.5 || unique.size() <= 20) return "categorical";
        return "text";
    }

    static Map<String, Object> profileColumn(String colName, List<String> values, String colType) {
        Map<String, Object> profile = new LinkedHashMap<>();
        profile.put("name", colName);
        profile.put("detected_type", colType);
        profile.put("total_count", values.size());

        int missingCount = (int) values.stream().filter(DataProfiler::isNull).count();
        double missingPct = values.size() > 0 ? round((double) missingCount / values.size() * 100, 2) : 0;
        profile.put("missing_count", missingCount);
        profile.put("missing_percentage", missingPct);

        Set<String> unique = new HashSet<>(cleanValues(values));
        profile.put("unique_count", unique.size());

        if (colType.equals("integer") || colType.equals("float") || colType.equals("numeric_mixed")) {
            profile.put("numeric_stats", computeNumericStats(values));
        } else if (colType.equals("categorical") || colType.equals("boolean") || colType.equals("text")) {
            profile.put("categorical_stats", computeCategoricalStats(values));
        }

        profile.put("distribution", computeDistribution(values, colType));
        return profile;
    }

    static Map<String, Object> computeNumericStats(List<String> values) {
        DescriptiveStatistics stats = new DescriptiveStatistics();
        for (String v : values) {
            if (!isNull(v)) {
                try {
                    stats.addValue(Double.parseDouble(v.trim()));
                } catch (NumberFormatException ignored) {}
            }
        }

        if (stats.getN() == 0) return new LinkedHashMap<>();

        Map<String, Object> result = new LinkedHashMap<>();
        result.put("count", (int) stats.getN());
        result.put("mean", round(stats.getMean(), 4));
        result.put("median", round(stats.getPercentile(50), 4));
        result.put("std", round(stats.getStandardDeviation(), 4));
        result.put("min", round(stats.getMin(), 4));
        result.put("max", round(stats.getMax(), 4));
        result.put("q1", round(stats.getPercentile(25), 4));
        result.put("q3", round(stats.getPercentile(75), 4));
        result.put("skewness", round(stats.getSkewness(), 4));
        result.put("kurtosis", round(stats.getKurtosis(), 4));

        long zeros = 0, negatives = 0;
        for (double v : stats.getValues()) {
            if (v == 0) zeros++;
            if (v < 0) negatives++;
        }
        result.put("zeros", (int) zeros);
        result.put("negatives", (int) negatives);
        return result;
    }

    static Map<String, Object> computeCategoricalStats(List<String> values) {
        List<String> clean = cleanValues(values).stream()
                .map(String::trim)
                .filter(v -> !v.isEmpty())
                .collect(Collectors.toList());

        if (clean.isEmpty()) return new LinkedHashMap<>();

        Map<String, Integer> freq = new LinkedHashMap<>();
        for (String v : clean) {
            freq.merge(v, 1, Integer::sum);
        }

        List<Map.Entry<String, Integer>> sorted = freq.entrySet().stream()
                .sorted(Map.Entry.<String, Integer>comparingByValue().reversed())
                .collect(Collectors.toList());

        Map<String, Integer> topValues = new LinkedHashMap<>();
        for (int i = 0; i < Math.min(10, sorted.size()); i++) {
            topValues.put(sorted.get(i).getKey(), sorted.get(i).getValue());
        }

        Map<String, Object> result = new LinkedHashMap<>();
        result.put("unique_count", new HashSet<>(clean).size());
        result.put("top_values", topValues);
        result.put("mode", sorted.isEmpty() ? null : sorted.get(0).getKey());
        result.put("mode_frequency", sorted.isEmpty() ? 0 : sorted.get(0).getValue());
        return result;
    }

    static Map<String, Object> computeDistribution(List<String> values, String colType) {
        Map<String, Object> result = new LinkedHashMap<>();

        if (colType.equals("integer") || colType.equals("float") || colType.equals("numeric_mixed")) {
            List<Double> nums = new ArrayList<>();
            for (String v : values) {
                if (!isNull(v)) {
                    try { nums.add(Double.parseDouble(v.trim())); } catch (NumberFormatException ignored) {}
                }
            }
            if (nums.isEmpty()) return result;

            double min = nums.stream().mapToDouble(Double::doubleValue).min().orElse(0);
            double max = nums.stream().mapToDouble(Double::doubleValue).max().orElse(0);
            int binCount = 10;
            double binWidth = (max - min) / binCount;
            if (binWidth == 0) binWidth = 1;

            List<Double> bins = new ArrayList<>();
            int[] counts = new int[binCount];
            for (int i = 0; i <= binCount; i++) {
                bins.add(round(min + i * binWidth, 4));
            }
            for (double n : nums) {
                int idx = (int) Math.floor((n - min) / binWidth);
                if (idx >= binCount) idx = binCount - 1;
                if (idx < 0) idx = 0;
                counts[idx]++;
            }

            result.put("type", "histogram");
            result.put("bins", bins);
            List<Integer> countsList = new ArrayList<>();
            for (int c : counts) countsList.add(c);
            result.put("counts", countsList);

        } else if (colType.equals("categorical") || colType.equals("boolean") || colType.equals("text")) {
            Map<String, Integer> freq = new LinkedHashMap<>();
            for (String v : values) {
                if (!isNull(v)) {
                    freq.merge(v, 1, Integer::sum);
                }
            }
            Map<String, Integer> topFreq = freq.entrySet().stream()
                    .sorted(Map.Entry.<String, Integer>comparingByValue().reversed())
                    .limit(20)
                    .collect(Collectors.toMap(Map.Entry::getKey, Map.Entry::getValue, (a, b) -> a, LinkedHashMap::new));

            result.put("type", "frequency");
            result.put("values", topFreq);
        }

        return result;
    }

    static Map<String, Map<String, Double>> computeCorrelationMatrix(
            List<Map<String, String>> data, List<String> columns, Map<String, String> colTypes) {

        List<String> numericCols = columns.stream()
                .filter(c -> {
                    String t = colTypes.get(c);
                    return "integer".equals(t) || "float".equals(t) || "numeric_mixed".equals(t);
                })
                .collect(Collectors.toList());

        Map<String, Map<String, Double>> result = new LinkedHashMap<>();
        if (numericCols.size() < 2) return result;

        for (String colA : numericCols) {
            result.put(colA, new LinkedHashMap<>());
            for (String colB : numericCols) {
                List<double[]> pairs = new ArrayList<>();
                for (Map<String, String> row : data) {
                    String va = row.get(colA);
                    String vb = row.get(colB);
                    if (!isNull(va) && !isNull(vb)) {
                        try {
                            double a = Double.parseDouble(va.trim());
                            double b = Double.parseDouble(vb.trim());
                            pairs.add(new double[]{a, b});
                        } catch (NumberFormatException ignored) {}
                    }
                }

                if (pairs.size() < 2) {
                    result.get(colA).put(colB, null);
                    continue;
                }

                double[] xs = pairs.stream().mapToDouble(p -> p[0]).toArray();
                double[] ys = pairs.stream().mapToDouble(p -> p[1]).toArray();

                try {
                    PearsonsCorrelation pc = new PearsonsCorrelation();
                    double corr = pc.correlation(xs, ys);
                    result.get(colA).put(colB, Double.isNaN(corr) ? null : round(corr, 4));
                } catch (Exception e) {
                    result.get(colA).put(colB, null);
                }
            }
        }

        return result;
    }

    static double computeDataQualityScore(List<Map<String, String>> data, List<String> columns) {
        int totalCells = data.size() * columns.size();
        if (totalCells == 0) return 0;

        int nullCount = 0;
        for (Map<String, String> row : data) {
            for (String col : columns) {
                if (isNull(row.get(col))) nullCount++;
            }
        }
        double completeness = 1.0 - (double) nullCount / totalCells;

        List<Double> uniquenessScores = new ArrayList<>();
        for (String col : columns) {
            List<String> clean = cleanValues(data.stream().map(r -> r.get(col)).collect(Collectors.toList()));
            if (!clean.isEmpty()) {
                uniquenessScores.add((double) new HashSet<>(clean).size() / clean.size());
            }
        }
        double avgUniqueness = uniquenessScores.isEmpty() ? 0 :
                uniquenessScores.stream().mapToDouble(Double::doubleValue).average().orElse(0);

        List<Double> consistencyScores = new ArrayList<>();
        for (String col : columns) {
            List<String> clean = cleanValues(data.stream().map(r -> r.get(col)).collect(Collectors.toList()));
            if (!clean.isEmpty()) {
                long numCount = clean.stream().filter(v -> {
                    try { Double.parseDouble(v.trim()); return true; }
                    catch (NumberFormatException e) { return false; }
                }).count();
                double ratio = (double) numCount / clean.size();
                consistencyScores.add(Math.max(ratio, 1 - ratio));
            }
        }
        double avgConsistency = consistencyScores.isEmpty() ? 0 :
                consistencyScores.stream().mapToDouble(Double::doubleValue).average().orElse(0);

        return round((completeness * 0.5 + avgConsistency * 0.3 + Math.min(avgUniqueness, 1.0) * 0.2) * 100, 2);
    }

    @SuppressWarnings("unchecked")
    static void printConsoleReport(Map<String, Object> report, List<String> columns) {
        Map<String, Object> overview = (Map<String, Object>) report.get("dataset_overview");
        System.out.println("--- Dataset Overview ---");
        System.out.printf("  Rows:            %s%n", overview.get("rows"));
        System.out.printf("  Columns:         %s%n", overview.get("columns"));
        System.out.printf("  Total cells:     %s%n", overview.get("total_cells"));
        System.out.printf("  Total missing:   %s (%s%%)%n", overview.get("total_missing"), overview.get("total_missing_pct"));
        System.out.printf("%n  Data Quality Score: %s / 100%n%n", report.get("data_quality_score"));

        System.out.println("--- Column Profiles ---");
        Map<String, Map<String, Object>> profiles = (Map<String, Map<String, Object>>) report.get("column_profiles");

        System.out.printf("  %-15s %-15s %-10s %-10s %-10s %-12s %-12s %-12s %-12s%n",
                "Column", "Type", "Missing", "Miss%", "Unique", "Min", "Max", "Mean", "Median");
        System.out.println("  " + "-".repeat(108));

        for (Map.Entry<String, Map<String, Object>> entry : profiles.entrySet()) {
            Map<String, Object> p = entry.getValue();
            String min = "", max = "", mean = "", median = "";

            Map<String, Object> ns = (Map<String, Object>) p.get("numeric_stats");
            if (ns != null && !ns.isEmpty()) {
                min = String.valueOf(ns.get("min"));
                max = String.valueOf(ns.get("max"));
                mean = String.valueOf(ns.get("mean"));
                median = String.valueOf(ns.get("median"));
            }

            System.out.printf("  %-15s %-15s %-10s %-10s %-10s %-12s %-12s %-12s %-12s%n",
                    truncate(entry.getKey(), 14),
                    p.get("detected_type"),
                    p.get("missing_count"),
                    p.get("missing_percentage") + "%",
                    p.get("unique_count"),
                    min, max, mean, median);
        }

        // Detailed output
        for (Map.Entry<String, Map<String, Object>> entry : profiles.entrySet()) {
            Map<String, Object> p = entry.getValue();
            System.out.printf("%n  [%s]%n", entry.getKey());
            System.out.printf("    Type:          %s%n", p.get("detected_type"));
            System.out.printf("    Missing:       %s (%s%%)%n", p.get("missing_count"), p.get("missing_percentage"));
            System.out.printf("    Unique:        %s%n", p.get("unique_count"));

            Map<String, Object> ns = (Map<String, Object>) p.get("numeric_stats");
            if (ns != null && !ns.isEmpty()) {
                for (Map.Entry<String, Object> stat : ns.entrySet()) {
                    System.out.printf("    %-15s %s%n", stat.getKey() + ":", stat.getValue());
                }
            }

            Map<String, Object> cs = (Map<String, Object>) p.get("categorical_stats");
            if (cs != null && !cs.isEmpty()) {
                System.out.printf("    Unique values: %s%n", cs.get("unique_count"));
                System.out.printf("    Mode:          %s (freq: %s)%n", cs.get("mode"), cs.get("mode_frequency"));
                Map<String, Integer> topVals = (Map<String, Integer>) cs.get("top_values");
                if (topVals != null) {
                    System.out.println("    Top values:");
                    int count = 0;
                    for (Map.Entry<String, Integer> tv : topVals.entrySet()) {
                        if (count++ >= 5) break;
                        System.out.printf("      %s: %s%n", tv.getKey(), tv.getValue());
                    }
                }
            }
        }

        // Correlation matrix
        Map<String, Map<String, Double>> corrMatrix =
                (Map<String, Map<String, Double>>) report.get("correlation_matrix");
        if (corrMatrix != null && !corrMatrix.isEmpty()) {
            System.out.println("\n--- Correlation Matrix ---");
            List<String> corrCols = new ArrayList<>(corrMatrix.keySet());
            System.out.printf("  %-15s", "");
            for (String col : corrCols) System.out.printf("%-14s", truncate(col, 12));
            System.out.println();

            for (String row : corrCols) {
                System.out.printf("  %-15s", truncate(row, 14));
                for (String col : corrCols) {
                    Double val = corrMatrix.get(row).get(col);
                    System.out.printf("%-14s", val != null ? val.toString() : "N/A");
                }
                System.out.println();
            }
        }
    }

    static String truncate(String s, int maxLen) {
        return s.length() > maxLen ? s.substring(0, maxLen) : s;
    }

    static double round(double val, int places) {
        double factor = Math.pow(10, places);
        return Math.round(val * factor) / factor;
    }
}