Duplicate Record Finder (cpp, written by Claude Code)
envgap__claude-code__cpp-t1-10
Written by a coding agent; not on GitHubWritten 2026-02-27
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
claude-code/cpp-t1 #10 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Duplicate Record Finder Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Support exact duplicate detection: find rows where all specified columns match exactly - Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric - Accept a --columns flag to specify which columns to compare (default: all columns) - Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85) - Group duplicates into clusters and assign each cluster an ID - For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates - Compute similarity scores for each pair within a cluster - Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range - Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns - Export a deduplicated CSV (keeping only primary records) via --deduplicate flag - If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it - Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.20)
project(DuplicateRecordFinder VERSION 1.0.0 LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
include(FetchContent)
# rapidfuzz-cpp (latest)
FetchContent_Declare(
rapidfuzz
GIT_REPOSITORY https://github.com/rapidfuzz/rapidfuzz-cpp.git
GIT_TAG main
)
FetchContent_MakeAvailable(rapidfuzz)
# nlohmann/json 3.11.3
FetchContent_Declare(
nlohmann_json
GIT_REPOSITORY https://github.com/nlohmann/json.git
GIT_TAG v3.11.3
)
FetchContent_MakeAvailable(nlohmann_json)
# fast-cpp-csv-parser (master)
FetchContent_Declare(
fast_csv
GIT_REPOSITORY https://github.com/ben-strasser/fast-cpp-csv-parser.git
GIT_TAG master
)
FetchContent_MakeAvailable(fast_csv)
add_executable(finder main.cpp)
target_include_directories(finder PRIVATE ${fast_csv_SOURCE_DIR})
target_link_libraries(finder
PRIVATE
rapidfuzz::rapidfuzz
nlohmann_json::nlohmann_json
)
main.cpp
/**
* Duplicate Record Finder - C++ Trial 1
* Uses rapidfuzz-cpp + nlohmann/json + fast-cpp-csv-parser.
* Supports exact matching, fuzzy string matching with configurable thresholds,
* and multiple comparison strategies (Levenshtein, Jaro-Winkler, Cosine).
*/
#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
#include <vector>
#include <map>
#include <set>
#include <cmath>
#include <algorithm>
#include <numeric>
#include <functional>
#include <iomanip>
#include <nlohmann/json.hpp>
#include <csv.h>
#include <rapidfuzz/fuzz.hpp>
#include <rapidfuzz/distance.hpp>
using json = nlohmann::json;
// --- Data Structures ---
struct Record {
std::map<std::string, std::string> fields;
};
struct ColumnRule {
std::string strategy;
double weight;
bool exact;
};
struct ColumnScore {
double score;
std::string strategy;
};
struct PairResult {
double overall;
std::map<std::string, ColumnScore> columns;
};
struct DuplicateGroup {
std::vector<int> indices;
std::map<std::string, PairResult> pairScores;
};
// --- Similarity Strategies ---
double levenshteinSimilarity(const std::string& s1, const std::string& s2) {
if (s1.empty() && s2.empty()) return 1.0;
double dist = static_cast<double>(rapidfuzz::levenshtein_distance(s1, s2));
double maxLen = static_cast<double>(std::max(s1.size(), s2.size()));
if (maxLen == 0.0) return 1.0;
return 1.0 - dist / maxLen;
}
double jaroSimilarity(const std::string& s1, const std::string& s2) {
if (s1 == s2) return 1.0;
if (s1.empty() || s2.empty()) return 0.0;
int maxDist = std::max(0, (int)(std::max(s1.size(), s2.size()) / 2) - 1);
std::vector<bool> s1Matches(s1.size(), false);
std::vector<bool> s2Matches(s2.size(), false);
int matches = 0, transpositions = 0;
for (int i = 0; i < (int)s1.size(); i++) {
int start = std::max(0, i - maxDist);
int end = std::min(i + maxDist + 1, (int)s2.size());
for (int j = start; j < end; j++) {
if (s2Matches[j] || s1[i] != s2[j]) continue;
s1Matches[i] = true;
s2Matches[j] = true;
matches++;
break;
}
}
if (matches == 0) return 0.0;
int k = 0;
for (int i = 0; i < (int)s1.size(); i++) {
if (!s1Matches[i]) continue;
while (!s2Matches[k]) k++;
if (s1[i] != s2[k]) transpositions++;
k++;
}
return (static_cast<double>(matches) / s1.size()
+ static_cast<double>(matches) / s2.size()
+ static_cast<double>(matches - transpositions / 2.0) / matches) / 3.0;
}
double jaroWinklerSimilarity(const std::string& s1, const std::string& s2) {
double jaro = jaroSimilarity(s1, s2);
int prefix = 0;
int limit = std::min(4, (int)std::min(s1.size(), s2.size()));
for (int i = 0; i < limit; i++) {
if (s1[i] == s2[i]) prefix++;
else break;
}
return jaro + prefix * 0.1 * (1.0 - jaro);
}
double cosineSimilarity(const std::string& s1, const std::string& s2) {
auto toLower = [](std::string s) {
std::transform(s.begin(), s.end(), s.begin(), ::tolower);
return s;
};
std::string a = toLower(s1);
std::string b = toLower(s2);
if (a.empty() || b.empty()) return 0.0;
auto bigrams = [](const std::string& s) {
std::map<std::string, int> bg;
for (size_t i = 0; i + 1 < s.size(); i++) {
bg[s.substr(i, 2)]++;
}
return bg;
};
auto bg1 = bigrams(a);
auto bg2 = bigrams(b);
if (bg1.empty() || bg2.empty()) return (a == b) ? 1.0 : 0.0;
std::set<std::string> allKeys;
for (auto& p : bg1) allKeys.insert(p.first);
for (auto& p : bg2) allKeys.insert(p.first);
double dot = 0, mag1 = 0, mag2 = 0;
for (auto& k : allKeys) {
double v1 = bg1.count(k) ? bg1[k] : 0;
double v2 = bg2.count(k) ? bg2[k] : 0;
dot += v1 * v2;
mag1 += v1 * v1;
mag2 += v2 * v2;
}
mag1 = std::sqrt(mag1);
mag2 = std::sqrt(mag2);
if (mag1 == 0 || mag2 == 0) return 0.0;
return dot / (mag1 * mag2);
}
using SimilarityFn = std::function<double(const std::string&, const std::string&)>;
std::map<std::string, SimilarityFn> STRATEGIES = {
{"levenshtein", levenshteinSimilarity},
{"jaro-winkler", jaroWinklerSimilarity},
{"cosine", cosineSimilarity},
};
// --- Core Logic ---
std::string trim(const std::string& s) {
size_t start = s.find_first_not_of(" \t\r\n");
size_t end = s.find_last_not_of(" \t\r\n");
return (start == std::string::npos) ? "" : s.substr(start, end - start + 1);
}
std::string toLowerStr(const std::string& s) {
std::string result = s;
std::transform(result.begin(), result.end(), result.begin(), ::tolower);
return result;
}
double roundTo4(double v) {
return std::round(v * 10000.0) / 10000.0;
}
PairResult computeRecordSimilarity(const Record& rec1, const Record& rec2,
const std::map<std::string, ColumnRule>& columnRules,
const std::string& defaultStrategy) {
PairResult result;
double weightedScore = 0.0, weightsSum = 0.0;
for (auto& [col, rule] : columnRules) {
std::string strategyName = rule.strategy.empty() ? defaultStrategy : rule.strategy;
double weight = rule.weight;
std::string val1 = trim(rec1.fields.count(col) ? rec1.fields.at(col) : "");
std::string val2 = trim(rec2.fields.count(col) ? rec2.fields.at(col) : "");
double score;
if (rule.exact) {
score = (toLowerStr(val1) == toLowerStr(val2)) ? 1.0 : 0.0;
} else {
auto it = STRATEGIES.find(strategyName);
SimilarityFn fn = (it != STRATEGIES.end()) ? it->second : levenshteinSimilarity;
score = fn(val1, val2);
}
result.columns[col] = {roundTo4(score), strategyName};
weightedScore += score * weight;
weightsSum += weight;
}
result.overall = weightsSum > 0 ? roundTo4(weightedScore / weightsSum) : 0.0;
return result;
}
std::vector<std::vector<int>> findExactDuplicates(const std::vector<Record>& records) {
std::map<std::string, std::vector<int>> groups;
for (int i = 0; i < (int)records.size(); i++) {
std::string key;
std::map<std::string, std::string> sorted(records[i].fields.begin(), records[i].fields.end());
for (auto& [k, v] : sorted) {
key += k + "=" + v + "|";
}
groups[key].push_back(i);
}
std::vector<std::vector<int>> result;
for (auto& [key, indices] : groups) {
if (indices.size() > 1) {
result.push_back(indices);
}
}
return result;
}
// Union-Find
class UnionFind {
public:
std::vector<int> parent;
UnionFind(int n) : parent(n) { std::iota(parent.begin(), parent.end(), 0); }
int find(int x) {
while (parent[x] != x) { parent[x] = parent[parent[x]]; x = parent[x]; }
return x;
}
void unite(int a, int b) {
int ra = find(a), rb = find(b);
if (ra != rb) parent[ra] = rb;
}
};
std::vector<DuplicateGroup> findFuzzyDuplicates(const std::vector<Record>& records,
const std::map<std::string, ColumnRule>& columnRules,
double threshold,
const std::string& defaultStrategy) {
int n = (int)records.size();
UnionFind uf(n);
std::map<std::string, PairResult> pairScores;
for (int i = 0; i < n; i++) {
for (int j = i + 1; j < n; j++) {
PairResult result = computeRecordSimilarity(records[i], records[j], columnRules, defaultStrategy);
if (result.overall >= threshold) {
uf.unite(i, j);
pairScores[std::to_string(i) + "-" + std::to_string(j)] = result;
}
}
}
std::map<int, std::vector<int>> groupMap;
for (int i = 0; i < n; i++) {
groupMap[uf.find(i)].push_back(i);
}
std::vector<DuplicateGroup> duplicateGroups;
for (auto& [root, members] : groupMap) {
if (members.size() > 1) {
DuplicateGroup group;
group.indices = members;
std::sort(group.indices.begin(), group.indices.end());
for (size_t a = 0; a < members.size(); a++) {
for (size_t b = a + 1; b < members.size(); b++) {
int lo = std::min(members[a], members[b]);
int hi = std::max(members[a], members[b]);
std::string key = std::to_string(lo) + "-" + std::to_string(hi);
if (pairScores.count(key)) {
group.pairScores[key] = pairScores[key];
}
}
}
duplicateGroups.push_back(group);
}
}
return duplicateGroups;
}
std::map<std::string, ColumnRule> buildColumnRules(const std::vector<std::string>& columns,
const std::string& strategy,
const std::vector<std::string>& exactCols) {
std::map<std::string, ColumnRule> rules;
std::set<std::string> exactSet(exactCols.begin(), exactCols.end());
for (auto& col : columns) {
std::string lower = toLowerStr(col);
if (lower == "id" || lower == "index") continue;
rules[col] = {strategy, 1.0, exactSet.count(col) > 0};
}
return rules;
}
// --- Sample Data ---
std::vector<Record> getSampleRecords() {
return {
{{{ "id","1"},{"first_name","John"},{"last_name","Smith"},{"email","john.smith@email.com"},{"phone","555-0101"},{"city","New York"}}},
{{{ "id","2"},{"first_name","John"},{"last_name","Smith"},{"email","john.smith@email.com"},{"phone","555-0101"},{"city","New York"}}},
{{{ "id","3"},{"first_name","Jon"},{"last_name","Smyth"},{"email","jon.smyth@email.com"},{"phone","555-0101"},{"city","New York"}}},
{{{ "id","4"},{"first_name","Jane"},{"last_name","Doe"},{"email","jane.doe@email.com"},{"phone","555-0202"},{"city","Los Angeles"}}},
{{{ "id","5"},{"first_name","Jane"},{"last_name","Doe"},{"email","jane.doe@email.com"},{"phone","555-0202"},{"city","Los Angeles"}}},
{{{ "id","6"},{"first_name","Jayne"},{"last_name","Doe"},{"email","jayne.doe@email.com"},{"phone","555-0203"},{"city","Los Angeles"}}},
{{{ "id","7"},{"first_name","Robert"},{"last_name","Johnson"},{"email","r.johnson@email.com"},{"phone","555-0303"},{"city","Chicago"}}},
{{{ "id","8"},{"first_name","Bob"},{"last_name","Johnson"},{"email","bob.johnson@email.com"},{"phone","555-0304"},{"city","Chicago"}}},
{{{ "id","9"},{"first_name","Alice"},{"last_name","Williams"},{"email","alice.w@email.com"},{"phone","555-0404"},{"city","Houston"}}},
{{{ "id","10"},{"first_name","Alice"},{"last_name","Willams"},{"email","alice.w@email.com"},{"phone","555-0404"},{"city","Houston"}}},
{{{ "id","11"},{"first_name","Michael"},{"last_name","Brown"},{"email","m.brown@email.com"},{"phone","555-0505"},{"city","Phoenix"}}},
{{{ "id","12"},{"first_name","Emily"},{"last_name","Davis"},{"email","emily.d@email.com"},{"phone","555-0606"},{"city","Philadelphia"}}},
{{{ "id","13"},{"first_name","Emilie"},{"last_name","Davis"},{"email","emilie.davis@email.com"},{"phone","555-0607"},{"city","Philadelphia"}}},
{{{ "id","14"},{"first_name","David"},{"last_name","Garcia"},{"email","d.garcia@email.com"},{"phone","555-0707"},{"city","San Antonio"}}},
{{{ "id","15"},{"first_name","David"},{"last_name","Garcia"},{"email","d.garcia@email.com"},{"phone","555-0707"},{"city","San Antonio"}}},
};
}
void generateSampleCsv(const std::string& outputPath) {
auto records = getSampleRecords();
std::ofstream out(outputPath);
if (!out.is_open()) {
std::cerr << "Error: Cannot create " << outputPath << std::endl;
return;
}
out << "id,first_name,last_name,email,phone,city\n";
for (auto& rec : records) {
out << rec.fields.at("id") << ","
<< rec.fields.at("first_name") << ","
<< rec.fields.at("last_name") << ","
<< rec.fields.at("email") << ","
<< rec.fields.at("phone") << ","
<< rec.fields.at("city") << "\n";
}
out.close();
std::cout << "Sample dataset generated: " << outputPath << " (" << records.size() << " records)" << std::endl;
}
// --- CSV Loading ---
std::vector<Record> loadCsv(const std::string& filePath, std::vector<std::string>& columns) {
std::vector<Record> records;
// Read header first
std::ifstream headerFile(filePath);
std::string headerLine;
std::getline(headerFile, headerLine);
headerFile.close();
std::stringstream ss(headerLine);
std::string col;
while (std::getline(ss, col, ',')) {
col = trim(col);
columns.push_back(col);
}
// Use fast-cpp-csv-parser with 6 columns
io::CSVReader<6> reader(filePath);
reader.read_header(io::ignore_extra_column, "id", "first_name", "last_name", "email", "phone", "city");
std::string id, firstName, lastName, email, phone, city;
while (reader.read_row(id, firstName, lastName, email, phone, city)) {
Record rec;
rec.fields["id"] = id;
rec.fields["first_name"] = firstName;
rec.fields["last_name"] = lastName;
rec.fields["email"] = email;
rec.fields["phone"] = phone;
rec.fields["city"] = city;
records.push_back(rec);
}
return records;
}
// --- Reporting ---
void printConsoleReport(const std::vector<Record>& records,
const std::vector<std::vector<int>>& exactGroups,
const std::vector<DuplicateGroup>& fuzzyGroups) {
std::string sep(70, '=');
std::cout << "\n" << sep << std::endl;
std::cout << " DUPLICATE RECORD FINDER - REPORT (rapidfuzz-cpp + nlohmann/json + fast-csv)" << std::endl;
std::cout << sep << std::endl;
std::cout << "\nTotal records analyzed: " << records.size() << std::endl;
std::cout << "\n--- Exact Duplicates: " << exactGroups.size() << " group(s) ---" << std::endl;
for (size_t i = 0; i < exactGroups.size(); i++) {
std::cout << "\n Group " << (i + 1) << " (" << exactGroups[i].size() << " records):" << std::endl;
for (int idx : exactGroups[i]) {
std::cout << " Row " << idx << ": ";
for (auto& [k, v] : records[idx].fields) {
std::cout << k << "=" << v << " ";
}
std::cout << std::endl;
}
}
std::cout << "\n--- Fuzzy Duplicate Groups: " << fuzzyGroups.size() << " group(s) ---" << std::endl;
for (size_t i = 0; i < fuzzyGroups.size(); i++) {
auto& group = fuzzyGroups[i];
std::cout << "\n Group " << (i + 1) << " (" << group.indices.size() << " records):" << std::endl;
for (int idx : group.indices) {
std::cout << " Row " << idx << ": ";
for (auto& [k, v] : records[idx].fields) {
std::cout << k << "=" << v << " ";
}
std::cout << std::endl;
}
for (auto& [pairKey, scores] : group.pairScores) {
std::cout << " Pair " << pairKey << ": overall=" << scores.overall << std::endl;
for (auto& [col, info] : scores.columns) {
std::cout << " " << col << ": " << info.score << " (" << info.strategy << ")" << std::endl;
}
}
}
std::cout << "\n" << sep << std::endl;
}
void generateJsonReport(const std::vector<Record>& records,
const std::vector<std::vector<int>>& exactGroups,
const std::vector<DuplicateGroup>& fuzzyGroups,
const std::string& outputPath) {
json report;
report["summary"]["total_records"] = records.size();
report["summary"]["exact_duplicate_groups"] = exactGroups.size();
report["summary"]["fuzzy_duplicate_groups"] = fuzzyGroups.size();
json exactArr = json::array();
for (auto& group : exactGroups) {
json g;
json recs = json::array();
for (int idx : group) {
json r;
r["row_index"] = idx;
r["data"] = records[idx].fields;
recs.push_back(r);
}
g["records"] = recs;
exactArr.push_back(g);
}
report["exact_duplicates"] = exactArr;
json fuzzyArr = json::array();
for (auto& group : fuzzyGroups) {
json g;
json recs = json::array();
for (int idx : group.indices) {
json r;
r["row_index"] = idx;
r["data"] = records[idx].fields;
recs.push_back(r);
}
g["records"] = recs;
json ps;
for (auto& [key, pr] : group.pairScores) {
json pj;
pj["overall"] = pr.overall;
json cols;
for (auto& [col, cs] : pr.columns) {
cols[col] = {{"score", cs.score}, {"strategy", cs.strategy}};
}
pj["columns"] = cols;
ps[key] = pj;
}
g["pair_scores"] = ps;
fuzzyArr.push_back(g);
}
report["fuzzy_duplicates"] = fuzzyArr;
std::ofstream out(outputPath);
out << report.dump(2) << std::endl;
out.close();
std::cout << "\nJSON report saved to: " << outputPath << std::endl;
}
// --- Argument Parsing ---
struct Config {
std::string input;
double threshold = 0.8;
std::string strategy = "levenshtein";
std::vector<std::string> exactCols;
std::string output = "duplicates_report.json";
};
Config parseArgs(int argc, char* argv[]) {
Config cfg;
for (int i = 1; i < argc; i++) {
std::string arg = argv[i];
if ((arg == "--input" || arg == "-i") && i + 1 < argc) {
cfg.input = argv[++i];
} else if ((arg == "--threshold" || arg == "-t") && i + 1 < argc) {
cfg.threshold = std::stod(argv[++i]);
} else if ((arg == "--strategy" || arg == "-s") && i + 1 < argc) {
cfg.strategy = argv[++i];
} else if (arg == "--exact-cols") {
while (i + 1 < argc && argv[i + 1][0] != '-') {
cfg.exactCols.push_back(argv[++i]);
}
} else if ((arg == "--output" || arg == "-o") && i + 1 < argc) {
cfg.output = argv[++i];
}
}
return cfg;
}
// --- Main ---
int main(int argc, char* argv[]) {
Config cfg = parseArgs(argc, argv);
std::vector<Record> records;
std::vector<std::string> columns;
if (!cfg.input.empty()) {
std::ifstream test(cfg.input);
if (!test.good()) {
std::cerr << "Error: File '" << cfg.input << "' not found." << std::endl;
return 1;
}
test.close();
std::cout << "Loading data from: " << cfg.input << std::endl;
records = loadCsv(cfg.input, columns);
} else {
std::cout << "No input file specified. Generating sample dataset..." << std::endl;
std::string samplePath = "sample_data.csv";
generateSampleCsv(samplePath);
records = loadCsv(samplePath, columns);
}
std::cout << "Loaded " << records.size() << " records with " << columns.size() << " columns" << std::endl;
auto columnRules = buildColumnRules(columns, cfg.strategy, cfg.exactCols);
std::cout << "Using strategy: " << cfg.strategy << ", threshold: " << cfg.threshold << std::endl;
std::cout << "\nFinding exact duplicates..." << std::endl;
auto exactGroups = findExactDuplicates(records);
std::cout << "Finding fuzzy duplicates..." << std::endl;
auto fuzzyGroups = findFuzzyDuplicates(records, columnRules, cfg.threshold, cfg.strategy);
printConsoleReport(records, exactGroups, fuzzyGroups);
generateJsonReport(records, exactGroups, fuzzyGroups, cfg.output);
return 0;
}
README.md
# Duplicate Record Finder - C++ Trial 1 Identifies duplicate and near-duplicate records in CSV files using exact matching and fuzzy string matching with configurable similarity thresholds. ## Dependencies - rapidfuzz-cpp (latest via FetchContent) - nlohmann/json 3.11.3 - fast-cpp-csv-parser (master) ## Build & Run ```bash mkdir build && cd build cmake .. cmake --build . ./finder # With input file ./finder --input data.csv # With custom threshold and strategy ./finder --input data.csv --threshold 0.85 --strategy jaro-winkler # With exact match columns ./finder --input data.csv --exact-cols email phone # Custom output path ./finder --input data.csv --output results.json ``` ## Strategies - `levenshtein` - Levenshtein distance-based similarity (via rapidfuzz) - `jaro-winkler` - Jaro-Winkler similarity - `cosine` - Cosine similarity using character bigrams ## Output - Console report with duplicate groups and similarity scores - `duplicates_report.json` with detailed findings