← All tasks
javacodex/java-t1 #10Not a task: already works

Duplicate Record Finder (java, written by Codex)

envgap__codex__java-t1-10

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
pom.xml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/java-t1 #10 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Duplicate Record Finder

Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Support exact duplicate detection: find rows where all specified columns match exactly
- Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric
- Accept a --columns flag to specify which columns to compare (default: all columns)
- Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85)
- Group duplicates into clusters and assign each cluster an ID
- For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates
- Compute similarity scores for each pair within a cluster
- Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range
- Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns
- Export a deduplicated CSV (keeping only primary records) via --deduplicate flag
- If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it
- Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair

Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include:
- Source code
- pom.xml with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
  <modelVersion>4.0.0</modelVersion>

  <groupId>tmlr.codex_generated.p10</groupId>
  <artifactId>duplicate-record-finder</artifactId>
  <version>1.0.0</version>
  <name>Duplicate Record Finder</name>

  <properties>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
    <maven.compiler.release>17</maven.compiler.release>
  </properties>

  <dependencies>
  </dependencies>

  <build>
    <plugins>
      <plugin>
        <groupId>org.apache.maven.plugins</groupId>
        <artifactId>maven-compiler-plugin</artifactId>
        <version>3.13.0</version>
      </plugin>
      <plugin>
        <groupId>org.codehaus.mojo</groupId>
        <artifactId>exec-maven-plugin</artifactId>
        <version>3.5.0</version>
        <configuration>
          <mainClass>DuplicateRecordFinder</mainClass>
        </configuration>
      </plugin>
    </plugins>
  </build>
</project>
README.md
# Duplicate Record Finder (Java)

Identifies exact and fuzzy duplicate clusters in CSV data with blocking/indexing for scalable matching.

## Requirements

- Ubuntu 22.04
- JDK 17+
- Maven 3.9+

## Dependencies

- Direct: none (standard library only)
- Transitive: none

Build plugins are pinned in `pom.xml` for reproducibility.

## Build

```bash
mvn -q -DskipTests compile
```

## Run

With input:

```bash
mvn -q exec:java -Dexec.args="/path/to/data.csv --columns name,email,city --threshold 0.85 --output duplicates_report.json --deduplicate deduplicated.csv"
```

No input (generates 500-record sample):

```bash
mvn -q exec:java
```
src/main/java/DuplicateRecordFinder.java
import java.io.BufferedWriter;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.util.ArrayList;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;

public class DuplicateRecordFinder {
    private record ParsedArgs(Map<String, String> options, List<String> positional) {}
    private record LoadResult(List<String> headers, List<Map<String, String>> rows) {}
    private record FindResult(List<List<Integer>> clusters, List<Map<String, Object>> pairScores) {}

    public static void main(String[] args) {
        try {
            ParsedArgs parsed = parseArgs(args);
            double threshold = Double.parseDouble(parsed.options.getOrDefault("threshold", "0.85"));
            threshold = Math.max(0.0, Math.min(1.0, threshold));
            Path outputPath = Paths.get(parsed.options.getOrDefault("output", "duplicates_report.json")).toAbsolutePath();

            Path inputPath;
            if (parsed.positional.isEmpty()) {
                inputPath = Paths.get("sample_duplicates.csv").toAbsolutePath();
                generateSample(inputPath);
                System.out.println("No input provided. Generated sample dataset: " + inputPath);
            } else {
                inputPath = Paths.get(parsed.positional.get(0)).toAbsolutePath();
                if (!Files.exists(inputPath)) {
                    System.err.println("Input file not found: " + inputPath);
                    System.exit(1);
                    return;
                }
            }

            LoadResult loaded = loadCsv(inputPath);
            if (loaded.rows.isEmpty()) {
                System.err.println("No rows found.");
                System.exit(1);
                return;
            }

            List<String> columns;
            if (parsed.options.containsKey("columns")) {
                columns = new ArrayList<>();
                for (String c : parsed.options.get("columns").split(",")) {
                    String t = c.trim();
                    if (!t.isEmpty() && loaded.headers.contains(t)) columns.add(t);
                }
            } else {
                columns = new ArrayList<>(loaded.headers);
            }
            if (columns.isEmpty()) {
                System.err.println("No valid comparison columns found.");
                System.exit(1);
                return;
            }

            FindResult found = findDuplicates(loaded.rows, columns, threshold);
            Map<String, Object> report = buildReport(loaded.rows, loaded.headers, columns, found, threshold);
            Files.writeString(outputPath, toJson(report, 0) + "\n", StandardCharsets.UTF_8);

            Path dedupPath = null;
            if (parsed.options.containsKey("deduplicate")) {
                String value = parsed.options.get("deduplicate");
                dedupPath = Paths.get(value.equals("true") ? "deduplicated.csv" : value).toAbsolutePath();
                writeDeduplicated(dedupPath, loaded.headers, loaded.rows, found.clusters);
            }

            printSummary(report, outputPath, dedupPath);
        } catch (Exception e) {
            System.err.println("Failed: " + e.getMessage());
            System.exit(1);
        }
    }

    private static ParsedArgs parseArgs(String[] args) {
        Map<String, String> options = new HashMap<>();
        List<String> positional = new ArrayList<>();
        for (int i = 0; i < args.length; i++) {
            String token = args[i];
            if (token.startsWith("--")) {
                String key = token.substring(2);
                if (i + 1 < args.length && !args[i + 1].startsWith("--")) options.put(key, args[++i]);
                else options.put(key, "true");
            } else {
                positional.add(token);
            }
        }
        return new ParsedArgs(options, positional);
    }

    private static List<String> parseCsvLine(String line) {
        List<String> out = new ArrayList<>();
        StringBuilder current = new StringBuilder();
        boolean inQuotes = false;
        for (int i = 0; i < line.length(); i++) {
            char ch = line.charAt(i);
            if (ch == '"') {
                if (inQuotes && i + 1 < line.length() && line.charAt(i + 1) == '"') {
                    current.append('"');
                    i++;
                } else {
                    inQuotes = !inQuotes;
                }
            } else if (ch == ',' && !inQuotes) {
                out.add(current.toString());
                current.setLength(0);
            } else {
                current.append(ch);
            }
        }
        out.add(current.toString());
        return out;
    }

    private static String csvEscape(String value) {
        String s = value == null ? "" : value;
        if (s.contains(",") || s.contains("\"") || s.contains("\n")) return "\"" + s.replace("\"", "\"\"") + "\"";
        return s;
    }

    private static LoadResult loadCsv(Path path) throws IOException {
        List<String> lines = Files.readAllLines(path, StandardCharsets.UTF_8);
        if (lines.isEmpty()) return new LoadResult(new ArrayList<>(), new ArrayList<>());
        if (!lines.get(0).isEmpty() && lines.get(0).charAt(0) == '\uFEFF') lines.set(0, lines.get(0).substring(1));
        List<String> headers = parseCsvLine(lines.get(0));
        List<Map<String, String>> rows = new ArrayList<>();
        for (int i = 1; i < lines.size(); i++) {
            if (lines.get(i).trim().isEmpty()) continue;
            List<String> parts = parseCsvLine(lines.get(i));
            Map<String, String> row = new LinkedHashMap<>();
            row.put("__index", String.valueOf(i - 1));
            for (int j = 0; j < headers.size(); j++) row.put(headers.get(j), j < parts.size() ? parts.get(j) : "");
            rows.add(row);
        }
        return new LoadResult(headers, rows);
    }

    private static String normalizeText(String value) {
        return value == null ? "" : value.toLowerCase().replace(".", "").replaceAll("\\s+", " ").trim();
    }

    private static int levenshtein(String a, String b) {
        if (a.equals(b)) return 0;
        if (a.isEmpty()) return b.length();
        if (b.isEmpty()) return a.length();
        int[][] dp = new int[a.length() + 1][b.length() + 1];
        for (int i = 0; i <= a.length(); i++) dp[i][0] = i;
        for (int j = 0; j <= b.length(); j++) dp[0][j] = j;
        for (int i = 1; i <= a.length(); i++) {
            for (int j = 1; j <= b.length(); j++) {
                int cost = a.charAt(i - 1) == b.charAt(j - 1) ? 0 : 1;
                dp[i][j] = Math.min(Math.min(dp[i - 1][j] + 1, dp[i][j - 1] + 1), dp[i - 1][j - 1] + cost);
            }
        }
        return dp[a.length()][b.length()];
    }

    private static double similarity(String a, String b) {
        String x = normalizeText(a);
        String y = normalizeText(b);
        int m = Math.max(x.length(), y.length());
        if (m == 0) return 1.0;
        return 1.0 - ((double) levenshtein(x, y) / m);
    }

    private static double recordSimilarity(Map<String, String> r1, Map<String, String> r2, List<String> columns) {
        if (columns.isEmpty()) return 1.0;
        double sum = 0.0;
        for (String c : columns) sum += similarity(r1.get(c), r2.get(c));
        return sum / columns.size();
    }

    private static String blockingKey(Map<String, String> row, List<String> columns) {
        List<String> tokens = new ArrayList<>();
        for (String c : columns) {
            String t = normalizeText(row.get(c)).replaceAll("[^a-z0-9]", "");
            tokens.add(t.length() <= 4 ? t : t.substring(0, 4));
        }
        return String.join("|", tokens);
    }

    private static class DSU {
        private final int[] parent;
        private final int[] rank;

        DSU(int n) {
            parent = new int[n];
            rank = new int[n];
            for (int i = 0; i < n; i++) parent[i] = i;
        }

        int find(int x) {
            if (parent[x] != x) parent[x] = find(parent[x]);
            return parent[x];
        }

        void union(int a, int b) {
            int ra = find(a), rb = find(b);
            if (ra == rb) return;
            if (rank[ra] < rank[rb]) {
                int tmp = ra;
                ra = rb;
                rb = tmp;
            }
            parent[rb] = ra;
            if (rank[ra] == rank[rb]) rank[ra]++;
        }
    }

    private static FindResult findDuplicates(List<Map<String, String>> rows, List<String> columns, double threshold) {
        DSU dsu = new DSU(rows.size());
        List<Map<String, Object>> pairScores = new ArrayList<>();

        Map<String, List<Integer>> exactGroups = new HashMap<>();
        for (int i = 0; i < rows.size(); i++) {
            StringBuilder key = new StringBuilder();
            for (String c : columns) key.append(rows.get(i).getOrDefault(c, "")).append('\u0001');
            exactGroups.computeIfAbsent(key.toString(), k -> new ArrayList<>()).add(i);
        }
        for (List<Integer> members : exactGroups.values()) {
            if (members.size() <= 1) continue;
            for (int i = 1; i < members.size(); i++) dsu.union(members.get(0), members.get(i));
            for (int i = 0; i < members.size(); i++) {
                for (int j = i + 1; j < members.size(); j++) {
                    pairScores.add(new LinkedHashMap<>(Map.of(
                            "i", members.get(i),
                            "j", members.get(j),
                            "score", 1.0,
                            "type", "exact"
                    )));
                }
            }
        }

        Map<String, List<Integer>> blocks = new HashMap<>();
        for (int i = 0; i < rows.size(); i++) {
            blocks.computeIfAbsent(blockingKey(rows.get(i), columns), k -> new ArrayList<>()).add(i);
        }
        for (List<Integer> bucket : blocks.values()) {
            if (bucket.size() < 2) continue;
            for (int a = 0; a < bucket.size(); a++) {
                for (int b = a + 1; b < bucket.size(); b++) {
                    int i = bucket.get(a), j = bucket.get(b);
                    double score = recordSimilarity(rows.get(i), rows.get(j), columns);
                    if (score >= threshold) {
                        dsu.union(i, j);
                        Map<String, Object> p = new LinkedHashMap<>();
                        p.put("i", i);
                        p.put("j", j);
                        p.put("score", score);
                        p.put("type", "fuzzy");
                        pairScores.add(p);
                    }
                }
            }
        }

        Map<Integer, List<Integer>> groups = new HashMap<>();
        for (int i = 0; i < rows.size(); i++) groups.computeIfAbsent(dsu.find(i), k -> new ArrayList<>()).add(i);
        List<List<Integer>> clusters = new ArrayList<>();
        for (List<Integer> members : groups.values()) {
            if (members.size() > 1) {
                Collections.sort(members);
                clusters.add(members);
            }
        }
        clusters.sort((a, b) -> Integer.compare(a.get(0), b.get(0)));
        return new FindResult(clusters, pairScores);
    }

    private static Map<String, Object> buildReport(List<Map<String, String>> rows, List<String> headers, List<String> columns,
                                                   FindResult found, double threshold) {
        List<Map<String, Object>> clustersOut = new ArrayList<>();
        for (int i = 0; i < found.clusters.size(); i++) {
            List<Integer> members = found.clusters.get(i);
            List<Map<String, Object>> records = new ArrayList<>();
            for (int j = 0; j < members.size(); j++) {
                int idx = members.get(j);
                Map<String, Object> rec = new LinkedHashMap<>();
                rec.put("row_index", idx);
                rec.put("role", j == 0 ? "primary" : "duplicate");
                Map<String, String> data = new LinkedHashMap<>();
                for (String h : headers) data.put(h, rows.get(idx).get(h));
                rec.put("data", data);
                records.add(rec);
            }

            List<Map<String, Object>> pairs = new ArrayList<>();
            for (int a = 0; a < members.size(); a++) {
                for (int b = a + 1; b < members.size(); b++) {
                    int ia = members.get(a), ib = members.get(b);
                    double score = recordSimilarity(rows.get(ia), rows.get(ib), columns);
                    Map<String, Object> pair = new LinkedHashMap<>();
                    pair.put("row_a", ia);
                    pair.put("row_b", ib);
                    pair.put("similarity", score);
                    pairs.add(pair);
                }
            }

            Map<String, Object> cluster = new LinkedHashMap<>();
            cluster.put("cluster_id", String.format("C%04d", i + 1));
            cluster.put("primary_row_index", members.get(0));
            cluster.put("size", members.size());
            cluster.put("matching_columns", columns);
            cluster.put("records", records);
            cluster.put("pairwise_similarity", pairs);
            clustersOut.add(cluster);
        }

        List<Double> allScores = new ArrayList<>();
        for (Map<String, Object> c : clustersOut) {
            @SuppressWarnings("unchecked")
            List<Map<String, Object>> pairs = (List<Map<String, Object>>) c.get("pairwise_similarity");
            for (Map<String, Object> p : pairs) allScores.add((Double) p.get("similarity"));
        }
        Map<String, Integer> breakdown = new LinkedHashMap<>();
        breakdown.put("0.85-0.90", (int) allScores.stream().filter(s -> s >= 0.85 && s < 0.90).count());
        breakdown.put("0.90-0.95", (int) allScores.stream().filter(s -> s >= 0.90 && s < 0.95).count());
        breakdown.put("0.95-1.00", (int) allScores.stream().filter(s -> s >= 0.95 && s <= 1.00).count());

        Map<String, Object> metadata = new LinkedHashMap<>();
        metadata.put("total_records", rows.size());
        metadata.put("threshold", threshold);
        metadata.put("compared_columns", columns);
        metadata.put("generated_at", Instant.now().toString());

        Map<String, Object> summary = new LinkedHashMap<>();
        summary.put("duplicate_clusters", clustersOut.size());
        summary.put("total_duplicate_records", clustersOut.stream().mapToInt(c -> (int) c.get("size") - 1).sum());
        summary.put("similarity_breakdown", breakdown);

        Map<String, Object> report = new LinkedHashMap<>();
        report.put("metadata", metadata);
        report.put("summary", summary);
        report.put("clusters", clustersOut);
        return report;
    }

    private static void writeDeduplicated(Path path, List<String> headers, List<Map<String, String>> rows, List<List<Integer>> clusters) throws IOException {
        Set<Integer> duplicates = new LinkedHashSet<>();
        for (List<Integer> c : clusters) {
            for (int i = 1; i < c.size(); i++) duplicates.add(c.get(i));
        }
        try (BufferedWriter writer = Files.newBufferedWriter(path, StandardCharsets.UTF_8)) {
            writer.write(String.join(",", headers));
            writer.write("\n");
            for (int i = 0; i < rows.size(); i++) {
                if (duplicates.contains(i)) continue;
                List<String> line = new ArrayList<>();
                for (String h : headers) line.add(csvEscape(rows.get(i).get(h)));
                writer.write(String.join(",", line));
                writer.write("\n");
            }
        }
    }

    private static void generateSample(Path path) throws IOException {
        String[] firstNames = {"Alice", "Bob", "Carol", "David", "Eva", "Frank", "Grace", "Helen"};
        String[] lastNames = {"Smith", "Johnson", "Brown", "Wilson", "Taylor", "Miller", "Davis", "Moore"};
        String[] cities = {"Austin", "Boston", "Chicago", "Denver", "Seattle"};
        List<Map<String, String>> rows = new ArrayList<>();
        for (int i = 0; i < 450; i++) {
            String fn = firstNames[i % firstNames.length];
            String ln = lastNames[(i * 3) % lastNames.length];
            String city = cities[i % cities.length];
            rows.add(new LinkedHashMap<>(Map.of(
                    "id", String.format("R%04d", i + 1),
                    "name", fn + " " + ln,
                    "email", fn.toLowerCase() + "." + ln.toLowerCase() + i + "@example.com",
                    "city", city,
                    "phone", "555-" + (1000 + i)
            )));
        }
        for (int i = 0; i < 25; i++) {
            Map<String, String> dup = new LinkedHashMap<>(rows.get(i));
            dup.put("id", "DUPX" + i);
            rows.add(dup);
        }
        for (int i = 0; i < 25; i++) {
            Map<String, String> base = new LinkedHashMap<>(rows.get(100 + i));
            base.put("id", "DUPF" + i);
            base.put("name", base.get("name").replace("Smith", "Smiht").replace("David", "Davd"));
            base.put("city", base.get("city").toLowerCase());
            base.put("email", base.get("email").replace("@example.com", "@example.co"));
            rows.add(base);
        }
        List<String> headers = List.of("id", "name", "email", "city", "phone");
        try (BufferedWriter writer = Files.newBufferedWriter(path, StandardCharsets.UTF_8)) {
            writer.write(String.join(",", headers));
            writer.write("\n");
            for (Map<String, String> row : rows) {
                List<String> line = new ArrayList<>();
                for (String h : headers) line.add(csvEscape(row.get(h)));
                writer.write(String.join(",", line));
                writer.write("\n");
            }
        }
    }

    @SuppressWarnings("unchecked")
    private static void printSummary(Map<String, Object> report, Path outputPath, Path dedupPath) {
        Map<String, Object> metadata = (Map<String, Object>) report.get("metadata");
        Map<String, Object> summary = (Map<String, Object>) report.get("summary");
        Map<String, Object> breakdown = (Map<String, Object>) summary.get("similarity_breakdown");
        System.out.println("Duplicate Record Finder");
        System.out.println("=======================");
        System.out.println("Total records         : " + metadata.get("total_records"));
        System.out.println("Duplicate clusters    : " + summary.get("duplicate_clusters"));
        System.out.println("Total duplicate rows  : " + summary.get("total_duplicate_records"));
        System.out.println("Similarity breakdown  :");
        for (Map.Entry<String, Object> e : breakdown.entrySet()) {
            System.out.println("  " + e.getKey() + ": " + e.getValue());
        }
        System.out.println("JSON report saved     : " + outputPath);
        if (dedupPath != null) System.out.println("Deduplicated CSV saved: " + dedupPath);
    }

    private static String toJson(Object obj, int indent) {
        String pad = "  ".repeat(indent);
        if (obj == null) return "null";
        if (obj instanceof String s) return "\"" + s.replace("\\", "\\\\").replace("\"", "\\\"") + "\"";
        if (obj instanceof Number || obj instanceof Boolean) return String.valueOf(obj);
        if (obj instanceof Map<?, ?> map) {
            StringBuilder out = new StringBuilder();
            out.append("{\n");
            int i = 0;
            for (Map.Entry<?, ?> e : map.entrySet()) {
                out.append(pad).append("  ").append(toJson(String.valueOf(e.getKey()), 0)).append(": ")
                        .append(toJson(e.getValue(), indent + 1));
                if (++i < map.size()) out.append(",");
                out.append("\n");
            }
            out.append(pad).append("}");
            return out.toString();
        }
        if (obj instanceof List<?> list) {
            StringBuilder out = new StringBuilder();
            out.append("[\n");
            for (int i = 0; i < list.size(); i++) {
                out.append(pad).append("  ").append(toJson(list.get(i), indent + 1));
                if (i + 1 < list.size()) out.append(",");
                out.append("\n");
            }
            out.append(pad).append("]");
            return out.toString();
        }
        return toJson(String.valueOf(obj), indent);
    }
}