Duplicate Record Finder (java, written by Claude Code)
envgap__claude-code__java-t1-10
Written by a coding agent; not on GitHubWritten 2026-02-27
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
claude-code/java-t1 #10 · read the task the agent was given
Claude Code wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Duplicate Record Finder Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Support exact duplicate detection: find rows where all specified columns match exactly - Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric - Accept a --columns flag to specify which columns to compare (default: all columns) - Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85) - Group duplicates into clusters and assign each cluster an ID - For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates - Compute similarity scores for each pair within a cluster - Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range - Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns - Export a deduplicated CSV (keeping only primary records) via --deduplicate flag - If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it - Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>com.dupfinder</groupId>
<artifactId>duplicate-record-finder-t1</artifactId>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>Duplicate Record Finder - Trial 1</name>
<description>Duplicate detection using Commons Text, Commons CSV, and Gson</description>
<properties>
<maven.compiler.source>17</maven.compiler.source>
<maven.compiler.target>17</maven.compiler.target>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
</properties>
<dependencies>
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-text</artifactId>
<version>1.11.0</version>
</dependency>
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-csv</artifactId>
<version>1.10.0</version>
</dependency>
<dependency>
<groupId>com.google.code.gson</groupId>
<artifactId>gson</artifactId>
<version>2.10.1</version>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-jar-plugin</artifactId>
<version>3.3.0</version>
<configuration>
<archive>
<manifest>
<mainClass>dupfinder.DuplicateRecordFinder</mainClass>
</manifest>
</archive>
</configuration>
</plugin>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-shade-plugin</artifactId>
<version>3.5.1</version>
<executions>
<execution>
<phase>package</phase>
<goals><goal>shade</goal></goals>
</execution>
</executions>
</plugin>
</plugins>
</build>
</project>
README.md
# Duplicate Record Finder - Java Trial 1 Identifies duplicate and near-duplicate records in CSV files using exact matching and fuzzy string matching with configurable similarity thresholds. ## Dependencies - Apache Commons Text 1.11.0 - Apache Commons CSV 1.10.0 - Gson 2.10.1 ## Build ```bash mvn clean package ``` ## Usage ```bash # With sample data (auto-generated) java -jar target/duplicate-record-finder-t1-1.0.0.jar # With custom CSV java -jar target/duplicate-record-finder-t1-1.0.0.jar --input data.csv # With custom threshold and strategy java -jar target/duplicate-record-finder-t1-1.0.0.jar --input data.csv --threshold 0.85 --strategy jaro-winkler # With exact match columns java -jar target/duplicate-record-finder-t1-1.0.0.jar --input data.csv --exact-cols email phone ``` ## Strategies - `levenshtein` - Levenshtein distance-based similarity - `jaro-winkler` - Jaro-Winkler similarity - `cosine` - Cosine similarity using word tokens ## Output - Console report with duplicate groups and similarity scores - `duplicates_report.json` with detailed findings
src/main/java/dupfinder/DuplicateRecordFinder.java
package dupfinder;
import com.google.gson.Gson;
import com.google.gson.GsonBuilder;
import org.apache.commons.csv.CSVFormat;
import org.apache.commons.csv.CSVParser;
import org.apache.commons.csv.CSVPrinter;
import org.apache.commons.csv.CSVRecord;
import org.apache.commons.text.similarity.CosineDistance;
import org.apache.commons.text.similarity.JaroWinklerSimilarity;
import org.apache.commons.text.similarity.LevenshteinDistance;
import java.io.*;
import java.nio.charset.StandardCharsets;
import java.util.*;
import java.util.stream.Collectors;
/**
* Duplicate Record Finder - Trial 1
* Uses Apache Commons Text + Commons CSV + Gson.
* Supports exact matching, fuzzy string matching with configurable thresholds,
* and multiple comparison strategies (Levenshtein, Jaro-Winkler, Cosine).
*/
public class DuplicateRecordFinder {
private static final LevenshteinDistance LEVENSHTEIN = new LevenshteinDistance();
private static final JaroWinklerSimilarity JARO_WINKLER = new JaroWinklerSimilarity();
private static final CosineDistance COSINE_DISTANCE = new CosineDistance();
private final double threshold;
private final String defaultStrategy;
private final Set<String> exactCols;
public DuplicateRecordFinder(double threshold, String defaultStrategy, Set<String> exactCols) {
this.threshold = threshold;
this.defaultStrategy = defaultStrategy;
this.exactCols = exactCols != null ? exactCols : new HashSet<>();
}
// --- Similarity strategies ---
public static double levenshteinSimilarity(String s1, String s2) {
if (s1 == null) s1 = "";
if (s2 == null) s2 = "";
if (s1.isEmpty() && s2.isEmpty()) return 1.0;
int maxLen = Math.max(s1.length(), s2.length());
if (maxLen == 0) return 1.0;
int dist = LEVENSHTEIN.apply(s1, s2);
return 1.0 - ((double) dist / maxLen);
}
public static double jaroWinklerSimilarity(String s1, String s2) {
if (s1 == null) s1 = "";
if (s2 == null) s2 = "";
return JARO_WINKLER.apply(s1, s2);
}
public static double cosineSimilarity(String s1, String s2) {
if (s1 == null) s1 = "";
if (s2 == null) s2 = "";
if (s1.isEmpty() || s2.isEmpty()) return 0.0;
try {
double dist = COSINE_DISTANCE.apply(s1, s2);
return 1.0 - dist;
} catch (Exception e) {
return s1.equals(s2) ? 1.0 : 0.0;
}
}
public double computeSimilarity(String s1, String s2, String strategy) {
return switch (strategy) {
case "jaro-winkler" -> jaroWinklerSimilarity(s1, s2);
case "cosine" -> cosineSimilarity(s1, s2);
default -> levenshteinSimilarity(s1, s2);
};
}
// --- Data structures ---
static class ColumnScore {
double score;
String strategy;
ColumnScore(double score, String strategy) {
this.score = Math.round(score * 10000.0) / 10000.0;
this.strategy = strategy;
}
}
static class PairScore {
double overall;
Map<String, ColumnScore> columns;
PairScore(double overall, Map<String, ColumnScore> columns) {
this.overall = Math.round(overall * 10000.0) / 10000.0;
this.columns = columns;
}
}
static class DuplicateGroup {
List<Integer> indices;
Map<String, PairScore> pairScores;
DuplicateGroup(List<Integer> indices, Map<String, PairScore> pairScores) {
this.indices = indices;
this.pairScores = pairScores;
}
}
static class RecordData {
int rowIndex;
Map<String, String> data;
RecordData(int rowIndex, Map<String, String> data) {
this.rowIndex = rowIndex;
this.data = data;
}
}
// --- Core logic ---
public PairScore computeRecordSimilarity(Map<String, String> rec1, Map<String, String> rec2, List<String> columns) {
Map<String, ColumnScore> colScores = new LinkedHashMap<>();
double weightedScore = 0.0;
double weightsSum = 0.0;
for (String col : columns) {
if (col.equalsIgnoreCase("id") || col.equalsIgnoreCase("index")) continue;
String strategy = defaultStrategy;
double weight = 1.0;
boolean exact = exactCols.contains(col);
String val1 = rec1.getOrDefault(col, "").trim();
String val2 = rec2.getOrDefault(col, "").trim();
double score;
if (exact) {
score = val1.equalsIgnoreCase(val2) ? 1.0 : 0.0;
} else {
score = computeSimilarity(val1, val2, strategy);
}
colScores.put(col, new ColumnScore(score, strategy));
weightedScore += score * weight;
weightsSum += weight;
}
double overall = weightsSum > 0 ? weightedScore / weightsSum : 0.0;
return new PairScore(overall, colScores);
}
public List<List<Integer>> findExactDuplicates(List<Map<String, String>> records) {
Map<String, List<Integer>> groups = new LinkedHashMap<>();
for (int i = 0; i < records.size(); i++) {
String key = records.get(i).entrySet().stream()
.sorted(Map.Entry.comparingByKey())
.map(e -> e.getKey() + "=" + e.getValue())
.collect(Collectors.joining("|"));
groups.computeIfAbsent(key, k -> new ArrayList<>()).add(i);
}
return groups.values().stream()
.filter(g -> g.size() > 1)
.collect(Collectors.toList());
}
public List<DuplicateGroup> findFuzzyDuplicates(List<Map<String, String>> records, List<String> columns) {
int n = records.size();
int[] parent = new int[n];
for (int i = 0; i < n; i++) parent[i] = i;
Map<String, PairScore> pairScores = new LinkedHashMap<>();
for (int i = 0; i < n; i++) {
for (int j = i + 1; j < n; j++) {
PairScore ps = computeRecordSimilarity(records.get(i), records.get(j), columns);
if (ps.overall >= threshold) {
union(parent, i, j);
pairScores.put(i + "-" + j, ps);
}
}
}
Map<Integer, List<Integer>> groupMap = new LinkedHashMap<>();
for (int i = 0; i < n; i++) {
int root = find(parent, i);
groupMap.computeIfAbsent(root, k -> new ArrayList<>()).add(i);
}
List<DuplicateGroup> result = new ArrayList<>();
for (List<Integer> members : groupMap.values()) {
if (members.size() > 1) {
Map<String, PairScore> groupPairs = new LinkedHashMap<>();
for (int a = 0; a < members.size(); a++) {
for (int b = a + 1; b < members.size(); b++) {
int idxA = members.get(a), idxB = members.get(b);
String key = Math.min(idxA, idxB) + "-" + Math.max(idxA, idxB);
if (pairScores.containsKey(key)) {
groupPairs.put(key, pairScores.get(key));
}
}
}
result.add(new DuplicateGroup(new ArrayList<>(members), groupPairs));
}
}
return result;
}
private int find(int[] parent, int x) {
while (parent[x] != x) {
parent[x] = parent[parent[x]];
x = parent[x];
}
return x;
}
private void union(int[] parent, int a, int b) {
int ra = find(parent, a), rb = find(parent, b);
if (ra != rb) parent[ra] = rb;
}
// --- CSV I/O ---
public static List<Map<String, String>> loadCsv(String path) throws IOException {
List<Map<String, String>> records = new ArrayList<>();
try (Reader reader = new FileReader(path, StandardCharsets.UTF_8);
CSVParser parser = CSVFormat.DEFAULT.builder().setHeader().setSkipHeaderRecord(true).build().parse(reader)) {
for (CSVRecord csvRecord : parser) {
Map<String, String> row = new LinkedHashMap<>(csvRecord.toMap());
records.add(row);
}
}
return records;
}
public static String generateSampleDataset(String outputPath) throws IOException {
String[][] data = {
{"1", "John", "Smith", "john.smith@email.com", "555-0101", "New York"},
{"2", "John", "Smith", "john.smith@email.com", "555-0101", "New York"},
{"3", "Jon", "Smyth", "jon.smyth@email.com", "555-0101", "New York"},
{"4", "Jane", "Doe", "jane.doe@email.com", "555-0202", "Los Angeles"},
{"5", "Jane", "Doe", "jane.doe@email.com", "555-0202", "Los Angeles"},
{"6", "Jayne", "Doe", "jayne.doe@email.com", "555-0203", "Los Angeles"},
{"7", "Robert", "Johnson", "r.johnson@email.com", "555-0303", "Chicago"},
{"8", "Bob", "Johnson", "bob.johnson@email.com", "555-0304", "Chicago"},
{"9", "Alice", "Williams", "alice.w@email.com", "555-0404", "Houston"},
{"10", "Alice", "Willams", "alice.w@email.com", "555-0404", "Houston"},
{"11", "Michael", "Brown", "m.brown@email.com", "555-0505", "Phoenix"},
{"12", "Emily", "Davis", "emily.d@email.com", "555-0606", "Philadelphia"},
{"13", "Emilie", "Davis", "emilie.davis@email.com", "555-0607", "Philadelphia"},
{"14", "David", "Garcia", "d.garcia@email.com", "555-0707", "San Antonio"},
{"15", "David", "Garcia", "d.garcia@email.com", "555-0707", "San Antonio"},
};
try (Writer writer = new FileWriter(outputPath, StandardCharsets.UTF_8);
CSVPrinter printer = CSVFormat.DEFAULT.builder()
.setHeader("id", "first_name", "last_name", "email", "phone", "city")
.build().print(writer)) {
for (String[] row : data) {
printer.printRecord((Object[]) row);
}
}
System.out.println("Sample dataset generated: " + outputPath + " (" + data.length + " records)");
return outputPath;
}
// --- Reporting ---
public void printConsoleReport(List<Map<String, String>> records, List<List<Integer>> exactGroups, List<DuplicateGroup> fuzzyGroups) {
System.out.println("\n" + "=".repeat(70));
System.out.println(" DUPLICATE RECORD FINDER - REPORT (Commons Text + Commons CSV + Gson)");
System.out.println("=".repeat(70));
System.out.println("\nTotal records analyzed: " + records.size());
System.out.println("\n--- Exact Duplicates: " + exactGroups.size() + " group(s) ---");
int g = 1;
for (List<Integer> group : exactGroups) {
System.out.println("\n Group " + g++ + " (" + group.size() + " records):");
for (int idx : group) {
System.out.println(" Row " + idx + ": " + records.get(idx));
}
}
System.out.println("\n--- Fuzzy Duplicate Groups: " + fuzzyGroups.size() + " group(s) ---");
g = 1;
for (DuplicateGroup dg : fuzzyGroups) {
System.out.println("\n Group " + g++ + " (" + dg.indices.size() + " records):");
for (int idx : dg.indices) {
System.out.println(" Row " + idx + ": " + records.get(idx));
}
for (Map.Entry<String, PairScore> entry : dg.pairScores.entrySet()) {
PairScore ps = entry.getValue();
System.out.println(" Pair " + entry.getKey() + ": overall=" + ps.overall);
for (Map.Entry<String, ColumnScore> colEntry : ps.columns.entrySet()) {
ColumnScore cs = colEntry.getValue();
System.out.println(" " + colEntry.getKey() + ": " + cs.score + " (" + cs.strategy + ")");
}
}
}
System.out.println("\n" + "=".repeat(70));
}
public void generateJsonReport(List<Map<String, String>> records, List<List<Integer>> exactGroups,
List<DuplicateGroup> fuzzyGroups, String outputPath) throws IOException {
Map<String, Object> report = new LinkedHashMap<>();
Map<String, Object> summary = new LinkedHashMap<>();
summary.put("total_records", records.size());
summary.put("exact_duplicate_groups", exactGroups.size());
summary.put("fuzzy_duplicate_groups", fuzzyGroups.size());
report.put("summary", summary);
List<Object> exactList = new ArrayList<>();
for (List<Integer> group : exactGroups) {
List<Map<String, Object>> groupRecords = new ArrayList<>();
for (int idx : group) {
Map<String, Object> entry = new LinkedHashMap<>();
entry.put("row_index", idx);
entry.put("data", records.get(idx));
groupRecords.add(entry);
}
Map<String, Object> gMap = new LinkedHashMap<>();
gMap.put("records", groupRecords);
exactList.add(gMap);
}
report.put("exact_duplicates", exactList);
List<Object> fuzzyList = new ArrayList<>();
for (DuplicateGroup dg : fuzzyGroups) {
List<Map<String, Object>> groupRecords = new ArrayList<>();
for (int idx : dg.indices) {
Map<String, Object> entry = new LinkedHashMap<>();
entry.put("row_index", idx);
entry.put("data", records.get(idx));
groupRecords.add(entry);
}
Map<String, Object> gMap = new LinkedHashMap<>();
gMap.put("records", groupRecords);
Map<String, Object> pairMap = new LinkedHashMap<>();
for (Map.Entry<String, PairScore> pe : dg.pairScores.entrySet()) {
Map<String, Object> psMap = new LinkedHashMap<>();
psMap.put("overall", pe.getValue().overall);
Map<String, Object> colMap = new LinkedHashMap<>();
for (Map.Entry<String, ColumnScore> ce : pe.getValue().columns.entrySet()) {
Map<String, Object> csMap = new LinkedHashMap<>();
csMap.put("score", ce.getValue().score);
csMap.put("strategy", ce.getValue().strategy);
colMap.put(ce.getKey(), csMap);
}
psMap.put("columns", colMap);
pairMap.put(pe.getKey(), psMap);
}
gMap.put("pair_scores", pairMap);
fuzzyList.add(gMap);
}
report.put("fuzzy_duplicates", fuzzyList);
Gson gson = new GsonBuilder().setPrettyPrinting().create();
try (Writer writer = new FileWriter(outputPath, StandardCharsets.UTF_8)) {
gson.toJson(report, writer);
}
System.out.println("\nJSON report saved to: " + outputPath);
}
// --- Main ---
public static void main(String[] args) {
String inputPath = null;
double threshold = 0.8;
String strategy = "levenshtein";
Set<String> exactCols = new HashSet<>();
String outputPath = "duplicates_report.json";
for (int i = 0; i < args.length; i++) {
switch (args[i]) {
case "--input", "-i" -> inputPath = args[++i];
case "--threshold", "-t" -> threshold = Double.parseDouble(args[++i]);
case "--strategy", "-s" -> strategy = args[++i];
case "--exact-cols" -> {
while (i + 1 < args.length && !args[i + 1].startsWith("-")) {
exactCols.add(args[++i]);
}
}
case "--output", "-o" -> outputPath = args[++i];
}
}
try {
String csvPath;
if (inputPath != null) {
if (!new File(inputPath).exists()) {
System.err.println("Error: File '" + inputPath + "' not found.");
System.exit(1);
}
csvPath = inputPath;
} else {
System.out.println("No input file specified. Generating sample dataset...");
csvPath = generateSampleDataset("sample_data.csv");
}
System.out.println("Loading data from: " + csvPath);
List<Map<String, String>> records = loadCsv(csvPath);
List<String> columns = new ArrayList<>(records.isEmpty() ? Collections.emptyList() : records.get(0).keySet());
System.out.println("Loaded " + records.size() + " records with columns: " + columns);
DuplicateRecordFinder finder = new DuplicateRecordFinder(threshold, strategy, exactCols);
System.out.println("Using strategy: " + strategy + ", threshold: " + threshold);
System.out.println("\nFinding exact duplicates...");
List<List<Integer>> exactGroups = finder.findExactDuplicates(records);
System.out.println("Finding fuzzy duplicates...");
List<DuplicateGroup> fuzzyGroups = finder.findFuzzyDuplicates(records, columns);
finder.printConsoleReport(records, exactGroups, fuzzyGroups);
finder.generateJsonReport(records, exactGroups, fuzzyGroups, outputPath);
} catch (Exception e) {
System.err.println("Error: " + e.getMessage());
e.printStackTrace();
System.exit(1);
}
}
}