Duplicate Record Finder (java, written by Codex)
envgap__codex__java-t1-10
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/java-t1 #10 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Duplicate Record Finder Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Support exact duplicate detection: find rows where all specified columns match exactly - Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric - Accept a --columns flag to specify which columns to compare (default: all columns) - Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85) - Group duplicates into clusters and assign each cluster an ID - For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates - Compute similarity scores for each pair within a cluster - Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range - Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns - Export a deduplicated CSV (keeping only primary records) via --deduplicate flag - If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it - Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>tmlr.codex_generated.p10</groupId>
<artifactId>duplicate-record-finder</artifactId>
<version>1.0.0</version>
<name>Duplicate Record Finder</name>
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<maven.compiler.release>17</maven.compiler.release>
</properties>
<dependencies>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.13.0</version>
</plugin>
<plugin>
<groupId>org.codehaus.mojo</groupId>
<artifactId>exec-maven-plugin</artifactId>
<version>3.5.0</version>
<configuration>
<mainClass>DuplicateRecordFinder</mainClass>
</configuration>
</plugin>
</plugins>
</build>
</project>
README.md
# Duplicate Record Finder (Java) Identifies exact and fuzzy duplicate clusters in CSV data with blocking/indexing for scalable matching. ## Requirements - Ubuntu 22.04 - JDK 17+ - Maven 3.9+ ## Dependencies - Direct: none (standard library only) - Transitive: none Build plugins are pinned in `pom.xml` for reproducibility. ## Build ```bash mvn -q -DskipTests compile ``` ## Run With input: ```bash mvn -q exec:java -Dexec.args="/path/to/data.csv --columns name,email,city --threshold 0.85 --output duplicates_report.json --deduplicate deduplicated.csv" ``` No input (generates 500-record sample): ```bash mvn -q exec:java ```
src/main/java/DuplicateRecordFinder.java
import java.io.BufferedWriter;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.util.ArrayList;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
public class DuplicateRecordFinder {
private record ParsedArgs(Map<String, String> options, List<String> positional) {}
private record LoadResult(List<String> headers, List<Map<String, String>> rows) {}
private record FindResult(List<List<Integer>> clusters, List<Map<String, Object>> pairScores) {}
public static void main(String[] args) {
try {
ParsedArgs parsed = parseArgs(args);
double threshold = Double.parseDouble(parsed.options.getOrDefault("threshold", "0.85"));
threshold = Math.max(0.0, Math.min(1.0, threshold));
Path outputPath = Paths.get(parsed.options.getOrDefault("output", "duplicates_report.json")).toAbsolutePath();
Path inputPath;
if (parsed.positional.isEmpty()) {
inputPath = Paths.get("sample_duplicates.csv").toAbsolutePath();
generateSample(inputPath);
System.out.println("No input provided. Generated sample dataset: " + inputPath);
} else {
inputPath = Paths.get(parsed.positional.get(0)).toAbsolutePath();
if (!Files.exists(inputPath)) {
System.err.println("Input file not found: " + inputPath);
System.exit(1);
return;
}
}
LoadResult loaded = loadCsv(inputPath);
if (loaded.rows.isEmpty()) {
System.err.println("No rows found.");
System.exit(1);
return;
}
List<String> columns;
if (parsed.options.containsKey("columns")) {
columns = new ArrayList<>();
for (String c : parsed.options.get("columns").split(",")) {
String t = c.trim();
if (!t.isEmpty() && loaded.headers.contains(t)) columns.add(t);
}
} else {
columns = new ArrayList<>(loaded.headers);
}
if (columns.isEmpty()) {
System.err.println("No valid comparison columns found.");
System.exit(1);
return;
}
FindResult found = findDuplicates(loaded.rows, columns, threshold);
Map<String, Object> report = buildReport(loaded.rows, loaded.headers, columns, found, threshold);
Files.writeString(outputPath, toJson(report, 0) + "\n", StandardCharsets.UTF_8);
Path dedupPath = null;
if (parsed.options.containsKey("deduplicate")) {
String value = parsed.options.get("deduplicate");
dedupPath = Paths.get(value.equals("true") ? "deduplicated.csv" : value).toAbsolutePath();
writeDeduplicated(dedupPath, loaded.headers, loaded.rows, found.clusters);
}
printSummary(report, outputPath, dedupPath);
} catch (Exception e) {
System.err.println("Failed: " + e.getMessage());
System.exit(1);
}
}
private static ParsedArgs parseArgs(String[] args) {
Map<String, String> options = new HashMap<>();
List<String> positional = new ArrayList<>();
for (int i = 0; i < args.length; i++) {
String token = args[i];
if (token.startsWith("--")) {
String key = token.substring(2);
if (i + 1 < args.length && !args[i + 1].startsWith("--")) options.put(key, args[++i]);
else options.put(key, "true");
} else {
positional.add(token);
}
}
return new ParsedArgs(options, positional);
}
private static List<String> parseCsvLine(String line) {
List<String> out = new ArrayList<>();
StringBuilder current = new StringBuilder();
boolean inQuotes = false;
for (int i = 0; i < line.length(); i++) {
char ch = line.charAt(i);
if (ch == '"') {
if (inQuotes && i + 1 < line.length() && line.charAt(i + 1) == '"') {
current.append('"');
i++;
} else {
inQuotes = !inQuotes;
}
} else if (ch == ',' && !inQuotes) {
out.add(current.toString());
current.setLength(0);
} else {
current.append(ch);
}
}
out.add(current.toString());
return out;
}
private static String csvEscape(String value) {
String s = value == null ? "" : value;
if (s.contains(",") || s.contains("\"") || s.contains("\n")) return "\"" + s.replace("\"", "\"\"") + "\"";
return s;
}
private static LoadResult loadCsv(Path path) throws IOException {
List<String> lines = Files.readAllLines(path, StandardCharsets.UTF_8);
if (lines.isEmpty()) return new LoadResult(new ArrayList<>(), new ArrayList<>());
if (!lines.get(0).isEmpty() && lines.get(0).charAt(0) == '\uFEFF') lines.set(0, lines.get(0).substring(1));
List<String> headers = parseCsvLine(lines.get(0));
List<Map<String, String>> rows = new ArrayList<>();
for (int i = 1; i < lines.size(); i++) {
if (lines.get(i).trim().isEmpty()) continue;
List<String> parts = parseCsvLine(lines.get(i));
Map<String, String> row = new LinkedHashMap<>();
row.put("__index", String.valueOf(i - 1));
for (int j = 0; j < headers.size(); j++) row.put(headers.get(j), j < parts.size() ? parts.get(j) : "");
rows.add(row);
}
return new LoadResult(headers, rows);
}
private static String normalizeText(String value) {
return value == null ? "" : value.toLowerCase().replace(".", "").replaceAll("\\s+", " ").trim();
}
private static int levenshtein(String a, String b) {
if (a.equals(b)) return 0;
if (a.isEmpty()) return b.length();
if (b.isEmpty()) return a.length();
int[][] dp = new int[a.length() + 1][b.length() + 1];
for (int i = 0; i <= a.length(); i++) dp[i][0] = i;
for (int j = 0; j <= b.length(); j++) dp[0][j] = j;
for (int i = 1; i <= a.length(); i++) {
for (int j = 1; j <= b.length(); j++) {
int cost = a.charAt(i - 1) == b.charAt(j - 1) ? 0 : 1;
dp[i][j] = Math.min(Math.min(dp[i - 1][j] + 1, dp[i][j - 1] + 1), dp[i - 1][j - 1] + cost);
}
}
return dp[a.length()][b.length()];
}
private static double similarity(String a, String b) {
String x = normalizeText(a);
String y = normalizeText(b);
int m = Math.max(x.length(), y.length());
if (m == 0) return 1.0;
return 1.0 - ((double) levenshtein(x, y) / m);
}
private static double recordSimilarity(Map<String, String> r1, Map<String, String> r2, List<String> columns) {
if (columns.isEmpty()) return 1.0;
double sum = 0.0;
for (String c : columns) sum += similarity(r1.get(c), r2.get(c));
return sum / columns.size();
}
private static String blockingKey(Map<String, String> row, List<String> columns) {
List<String> tokens = new ArrayList<>();
for (String c : columns) {
String t = normalizeText(row.get(c)).replaceAll("[^a-z0-9]", "");
tokens.add(t.length() <= 4 ? t : t.substring(0, 4));
}
return String.join("|", tokens);
}
private static class DSU {
private final int[] parent;
private final int[] rank;
DSU(int n) {
parent = new int[n];
rank = new int[n];
for (int i = 0; i < n; i++) parent[i] = i;
}
int find(int x) {
if (parent[x] != x) parent[x] = find(parent[x]);
return parent[x];
}
void union(int a, int b) {
int ra = find(a), rb = find(b);
if (ra == rb) return;
if (rank[ra] < rank[rb]) {
int tmp = ra;
ra = rb;
rb = tmp;
}
parent[rb] = ra;
if (rank[ra] == rank[rb]) rank[ra]++;
}
}
private static FindResult findDuplicates(List<Map<String, String>> rows, List<String> columns, double threshold) {
DSU dsu = new DSU(rows.size());
List<Map<String, Object>> pairScores = new ArrayList<>();
Map<String, List<Integer>> exactGroups = new HashMap<>();
for (int i = 0; i < rows.size(); i++) {
StringBuilder key = new StringBuilder();
for (String c : columns) key.append(rows.get(i).getOrDefault(c, "")).append('\u0001');
exactGroups.computeIfAbsent(key.toString(), k -> new ArrayList<>()).add(i);
}
for (List<Integer> members : exactGroups.values()) {
if (members.size() <= 1) continue;
for (int i = 1; i < members.size(); i++) dsu.union(members.get(0), members.get(i));
for (int i = 0; i < members.size(); i++) {
for (int j = i + 1; j < members.size(); j++) {
pairScores.add(new LinkedHashMap<>(Map.of(
"i", members.get(i),
"j", members.get(j),
"score", 1.0,
"type", "exact"
)));
}
}
}
Map<String, List<Integer>> blocks = new HashMap<>();
for (int i = 0; i < rows.size(); i++) {
blocks.computeIfAbsent(blockingKey(rows.get(i), columns), k -> new ArrayList<>()).add(i);
}
for (List<Integer> bucket : blocks.values()) {
if (bucket.size() < 2) continue;
for (int a = 0; a < bucket.size(); a++) {
for (int b = a + 1; b < bucket.size(); b++) {
int i = bucket.get(a), j = bucket.get(b);
double score = recordSimilarity(rows.get(i), rows.get(j), columns);
if (score >= threshold) {
dsu.union(i, j);
Map<String, Object> p = new LinkedHashMap<>();
p.put("i", i);
p.put("j", j);
p.put("score", score);
p.put("type", "fuzzy");
pairScores.add(p);
}
}
}
}
Map<Integer, List<Integer>> groups = new HashMap<>();
for (int i = 0; i < rows.size(); i++) groups.computeIfAbsent(dsu.find(i), k -> new ArrayList<>()).add(i);
List<List<Integer>> clusters = new ArrayList<>();
for (List<Integer> members : groups.values()) {
if (members.size() > 1) {
Collections.sort(members);
clusters.add(members);
}
}
clusters.sort((a, b) -> Integer.compare(a.get(0), b.get(0)));
return new FindResult(clusters, pairScores);
}
private static Map<String, Object> buildReport(List<Map<String, String>> rows, List<String> headers, List<String> columns,
FindResult found, double threshold) {
List<Map<String, Object>> clustersOut = new ArrayList<>();
for (int i = 0; i < found.clusters.size(); i++) {
List<Integer> members = found.clusters.get(i);
List<Map<String, Object>> records = new ArrayList<>();
for (int j = 0; j < members.size(); j++) {
int idx = members.get(j);
Map<String, Object> rec = new LinkedHashMap<>();
rec.put("row_index", idx);
rec.put("role", j == 0 ? "primary" : "duplicate");
Map<String, String> data = new LinkedHashMap<>();
for (String h : headers) data.put(h, rows.get(idx).get(h));
rec.put("data", data);
records.add(rec);
}
List<Map<String, Object>> pairs = new ArrayList<>();
for (int a = 0; a < members.size(); a++) {
for (int b = a + 1; b < members.size(); b++) {
int ia = members.get(a), ib = members.get(b);
double score = recordSimilarity(rows.get(ia), rows.get(ib), columns);
Map<String, Object> pair = new LinkedHashMap<>();
pair.put("row_a", ia);
pair.put("row_b", ib);
pair.put("similarity", score);
pairs.add(pair);
}
}
Map<String, Object> cluster = new LinkedHashMap<>();
cluster.put("cluster_id", String.format("C%04d", i + 1));
cluster.put("primary_row_index", members.get(0));
cluster.put("size", members.size());
cluster.put("matching_columns", columns);
cluster.put("records", records);
cluster.put("pairwise_similarity", pairs);
clustersOut.add(cluster);
}
List<Double> allScores = new ArrayList<>();
for (Map<String, Object> c : clustersOut) {
@SuppressWarnings("unchecked")
List<Map<String, Object>> pairs = (List<Map<String, Object>>) c.get("pairwise_similarity");
for (Map<String, Object> p : pairs) allScores.add((Double) p.get("similarity"));
}
Map<String, Integer> breakdown = new LinkedHashMap<>();
breakdown.put("0.85-0.90", (int) allScores.stream().filter(s -> s >= 0.85 && s < 0.90).count());
breakdown.put("0.90-0.95", (int) allScores.stream().filter(s -> s >= 0.90 && s < 0.95).count());
breakdown.put("0.95-1.00", (int) allScores.stream().filter(s -> s >= 0.95 && s <= 1.00).count());
Map<String, Object> metadata = new LinkedHashMap<>();
metadata.put("total_records", rows.size());
metadata.put("threshold", threshold);
metadata.put("compared_columns", columns);
metadata.put("generated_at", Instant.now().toString());
Map<String, Object> summary = new LinkedHashMap<>();
summary.put("duplicate_clusters", clustersOut.size());
summary.put("total_duplicate_records", clustersOut.stream().mapToInt(c -> (int) c.get("size") - 1).sum());
summary.put("similarity_breakdown", breakdown);
Map<String, Object> report = new LinkedHashMap<>();
report.put("metadata", metadata);
report.put("summary", summary);
report.put("clusters", clustersOut);
return report;
}
private static void writeDeduplicated(Path path, List<String> headers, List<Map<String, String>> rows, List<List<Integer>> clusters) throws IOException {
Set<Integer> duplicates = new LinkedHashSet<>();
for (List<Integer> c : clusters) {
for (int i = 1; i < c.size(); i++) duplicates.add(c.get(i));
}
try (BufferedWriter writer = Files.newBufferedWriter(path, StandardCharsets.UTF_8)) {
writer.write(String.join(",", headers));
writer.write("\n");
for (int i = 0; i < rows.size(); i++) {
if (duplicates.contains(i)) continue;
List<String> line = new ArrayList<>();
for (String h : headers) line.add(csvEscape(rows.get(i).get(h)));
writer.write(String.join(",", line));
writer.write("\n");
}
}
}
private static void generateSample(Path path) throws IOException {
String[] firstNames = {"Alice", "Bob", "Carol", "David", "Eva", "Frank", "Grace", "Helen"};
String[] lastNames = {"Smith", "Johnson", "Brown", "Wilson", "Taylor", "Miller", "Davis", "Moore"};
String[] cities = {"Austin", "Boston", "Chicago", "Denver", "Seattle"};
List<Map<String, String>> rows = new ArrayList<>();
for (int i = 0; i < 450; i++) {
String fn = firstNames[i % firstNames.length];
String ln = lastNames[(i * 3) % lastNames.length];
String city = cities[i % cities.length];
rows.add(new LinkedHashMap<>(Map.of(
"id", String.format("R%04d", i + 1),
"name", fn + " " + ln,
"email", fn.toLowerCase() + "." + ln.toLowerCase() + i + "@example.com",
"city", city,
"phone", "555-" + (1000 + i)
)));
}
for (int i = 0; i < 25; i++) {
Map<String, String> dup = new LinkedHashMap<>(rows.get(i));
dup.put("id", "DUPX" + i);
rows.add(dup);
}
for (int i = 0; i < 25; i++) {
Map<String, String> base = new LinkedHashMap<>(rows.get(100 + i));
base.put("id", "DUPF" + i);
base.put("name", base.get("name").replace("Smith", "Smiht").replace("David", "Davd"));
base.put("city", base.get("city").toLowerCase());
base.put("email", base.get("email").replace("@example.com", "@example.co"));
rows.add(base);
}
List<String> headers = List.of("id", "name", "email", "city", "phone");
try (BufferedWriter writer = Files.newBufferedWriter(path, StandardCharsets.UTF_8)) {
writer.write(String.join(",", headers));
writer.write("\n");
for (Map<String, String> row : rows) {
List<String> line = new ArrayList<>();
for (String h : headers) line.add(csvEscape(row.get(h)));
writer.write(String.join(",", line));
writer.write("\n");
}
}
}
@SuppressWarnings("unchecked")
private static void printSummary(Map<String, Object> report, Path outputPath, Path dedupPath) {
Map<String, Object> metadata = (Map<String, Object>) report.get("metadata");
Map<String, Object> summary = (Map<String, Object>) report.get("summary");
Map<String, Object> breakdown = (Map<String, Object>) summary.get("similarity_breakdown");
System.out.println("Duplicate Record Finder");
System.out.println("=======================");
System.out.println("Total records : " + metadata.get("total_records"));
System.out.println("Duplicate clusters : " + summary.get("duplicate_clusters"));
System.out.println("Total duplicate rows : " + summary.get("total_duplicate_records"));
System.out.println("Similarity breakdown :");
for (Map.Entry<String, Object> e : breakdown.entrySet()) {
System.out.println(" " + e.getKey() + ": " + e.getValue());
}
System.out.println("JSON report saved : " + outputPath);
if (dedupPath != null) System.out.println("Deduplicated CSV saved: " + dedupPath);
}
private static String toJson(Object obj, int indent) {
String pad = " ".repeat(indent);
if (obj == null) return "null";
if (obj instanceof String s) return "\"" + s.replace("\\", "\\\\").replace("\"", "\\\"") + "\"";
if (obj instanceof Number || obj instanceof Boolean) return String.valueOf(obj);
if (obj instanceof Map<?, ?> map) {
StringBuilder out = new StringBuilder();
out.append("{\n");
int i = 0;
for (Map.Entry<?, ?> e : map.entrySet()) {
out.append(pad).append(" ").append(toJson(String.valueOf(e.getKey()), 0)).append(": ")
.append(toJson(e.getValue(), indent + 1));
if (++i < map.size()) out.append(",");
out.append("\n");
}
out.append(pad).append("}");
return out.toString();
}
if (obj instanceof List<?> list) {
StringBuilder out = new StringBuilder();
out.append("[\n");
for (int i = 0; i < list.size(); i++) {
out.append(pad).append(" ").append(toJson(list.get(i), indent + 1));
if (i + 1 < list.size()) out.append(",");
out.append("\n");
}
out.append(pad).append("]");
return out.toString();
}
return toJson(String.valueOf(obj), indent);
}
}