Duplicate Record Finder (cpp, written by Codex)
envgap__codex__cpp-t1-10
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/cpp-t1 #10 · read the task the agent was given
Codex wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Duplicate Record Finder Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Support exact duplicate detection: find rows where all specified columns match exactly - Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric - Accept a --columns flag to specify which columns to compare (default: all columns) - Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85) - Group duplicates into clusters and assign each cluster an ID - For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates - Compute similarity scores for each pair within a cluster - Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range - Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns - Export a deduplicated CSV (keeping only primary records) via --deduplicate flag - If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it - Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.22) project(duplicate_record_finder_cpp VERSION 1.0.0 LANGUAGES CXX) set(CMAKE_CXX_STANDARD 20) set(CMAKE_CXX_STANDARD_REQUIRED ON) set(CMAKE_CXX_EXTENSIONS OFF) include(FetchContent) # Pinned dependency FetchContent_Declare( nlohmann_json URL https://github.com/nlohmann/json/releases/download/v3.11.3/json.tar.xz ) FetchContent_MakeAvailable(nlohmann_json) add_executable(duplicate_finder src/main.cpp) target_link_libraries(duplicate_finder PRIVATE nlohmann_json::nlohmann_json)
README.md
# Duplicate Record Finder (C++) Finds exact and near-duplicate records in CSV data using: - exact matching on selected columns - fuzzy matching with normalized Levenshtein similarity - blocking/indexing to avoid full all-pairs comparisons - cluster construction with primary record selection ## Requirements - Ubuntu 22.04 - G++ 12+ - CMake 3.22+ - Network access during CMake configure ## Dependencies - Direct (pinned in `CMakeLists.txt`): - `nlohmann/json v3.11.3` - Transitive: - none ## Build ```bash cmake -S . -B build cmake --build build ``` ## Run With input: ```bash ./build/duplicate_finder /path/to/data.csv --columns name,email,city --threshold 0.85 --output duplicates_report.json --deduplicate deduplicated.csv ``` No input (generates 500-record sample): ```bash ./build/duplicate_finder ```
src/main.cpp
#include <algorithm>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <map>
#include <numeric>
#include <optional>
#include <regex>
#include <set>
#include <sstream>
#include <string>
#include <unordered_map>
#include <vector>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
namespace {
struct ParsedArgs {
std::map<std::string, std::string> options;
std::vector<std::string> positional;
};
struct LoadResult {
std::vector<std::string> headers;
std::vector<std::map<std::string, std::string>> rows;
};
struct FindResult {
std::vector<std::vector<int>> clusters;
std::vector<json> pairScores;
};
std::string pad4(int n) {
std::ostringstream out;
out << std::setw(4) << std::setfill('0') << n;
return out.str();
}
ParsedArgs parseArgs(int argc, char** argv) {
ParsedArgs out;
for (int i = 1; i < argc; i++) {
std::string token = argv[i];
if (token.rfind("--", 0) == 0) {
std::string key = token.substr(2);
if (i + 1 < argc && std::string(argv[i + 1]).rfind("--", 0) != 0) out.options[key] = argv[++i];
else out.options[key] = "true";
} else {
out.positional.push_back(token);
}
}
return out;
}
std::vector<std::string> parseCsvLine(const std::string& line) {
std::vector<std::string> out;
std::string current;
bool inQuotes = false;
for (std::size_t i = 0; i < line.size(); i++) {
char ch = line[i];
if (ch == '"') {
if (inQuotes && i + 1 < line.size() && line[i + 1] == '"') {
current += '"';
i++;
} else {
inQuotes = !inQuotes;
}
} else if (ch == ',' && !inQuotes) {
out.push_back(current);
current.clear();
} else {
current += ch;
}
}
out.push_back(current);
return out;
}
std::string csvEscape(const std::string& value) {
if (value.find(',') != std::string::npos || value.find('"') != std::string::npos || value.find('\n') != std::string::npos) {
std::string out = "\"";
for (char c : value) {
if (c == '"') out += "\"\"";
else out += c;
}
out += "\"";
return out;
}
return value;
}
LoadResult loadCsv(const std::filesystem::path& path) {
std::ifstream in(path);
if (!in.is_open()) throw std::runtime_error("Failed to open CSV: " + path.string());
std::string line;
if (!std::getline(in, line)) return {};
if (!line.empty() && static_cast<unsigned char>(line[0]) == 0xEF && line.size() >= 3 &&
static_cast<unsigned char>(line[1]) == 0xBB && static_cast<unsigned char>(line[2]) == 0xBF) {
line = line.substr(3);
}
auto headers = parseCsvLine(line);
std::vector<std::map<std::string, std::string>> rows;
int idx = 0;
while (std::getline(in, line)) {
if (!line.empty() && line.back() == '\r') line.pop_back();
if (line.empty()) continue;
auto parts = parseCsvLine(line);
std::map<std::string, std::string> row;
row["__index"] = std::to_string(idx++);
for (std::size_t i = 0; i < headers.size(); i++) row[headers[i]] = i < parts.size() ? parts[i] : "";
rows.push_back(row);
}
return {headers, rows};
}
std::string normalizeText(const std::string& value) {
std::string out = value;
std::transform(out.begin(), out.end(), out.begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
out = std::regex_replace(out, std::regex(R"(\.)"), "");
out = std::regex_replace(out, std::regex(R"(\s+)"), " ");
out = std::regex_replace(out, std::regex(R"(^\s+|\s+$)"), "");
return out;
}
int levenshtein(const std::string& a, const std::string& b) {
if (a == b) return 0;
if (a.empty()) return static_cast<int>(b.size());
if (b.empty()) return static_cast<int>(a.size());
std::vector<std::vector<int>> dp(a.size() + 1, std::vector<int>(b.size() + 1, 0));
for (std::size_t i = 0; i <= a.size(); i++) dp[i][0] = static_cast<int>(i);
for (std::size_t j = 0; j <= b.size(); j++) dp[0][j] = static_cast<int>(j);
for (std::size_t i = 1; i <= a.size(); i++) {
for (std::size_t j = 1; j <= b.size(); j++) {
int cost = a[i - 1] == b[j - 1] ? 0 : 1;
dp[i][j] = std::min({dp[i - 1][j] + 1, dp[i][j - 1] + 1, dp[i - 1][j - 1] + cost});
}
}
return dp[a.size()][b.size()];
}
double similarity(const std::string& a, const std::string& b) {
std::string x = normalizeText(a);
std::string y = normalizeText(b);
std::size_t m = std::max(x.size(), y.size());
if (m == 0) return 1.0;
return 1.0 - static_cast<double>(levenshtein(x, y)) / static_cast<double>(m);
}
double recordSimilarity(const std::map<std::string, std::string>& a,
const std::map<std::string, std::string>& b,
const std::vector<std::string>& columns) {
if (columns.empty()) return 1.0;
double sum = 0.0;
for (const auto& c : columns) sum += similarity(a.at(c), b.at(c));
return sum / columns.size();
}
std::string blockingKey(const std::map<std::string, std::string>& row, const std::vector<std::string>& columns) {
std::ostringstream out;
for (std::size_t i = 0; i < columns.size(); i++) {
std::string t = normalizeText(row.at(columns[i]));
t = std::regex_replace(t, std::regex(R"([^a-z0-9])"), "");
if (t.size() > 4) t = t.substr(0, 4);
if (i) out << "|";
out << t;
}
return out.str();
}
struct DSU {
std::vector<int> parent;
std::vector<int> rank;
explicit DSU(int n) : parent(n), rank(n, 0) {
std::iota(parent.begin(), parent.end(), 0);
}
int find(int x) {
if (parent[x] != x) parent[x] = find(parent[x]);
return parent[x];
}
void unite(int a, int b) {
int ra = find(a), rb = find(b);
if (ra == rb) return;
if (rank[ra] < rank[rb]) std::swap(ra, rb);
parent[rb] = ra;
if (rank[ra] == rank[rb]) rank[ra]++;
}
};
FindResult findDuplicates(const std::vector<std::map<std::string, std::string>>& rows,
const std::vector<std::string>& columns,
double threshold) {
DSU dsu(static_cast<int>(rows.size()));
std::vector<json> pairScores;
std::unordered_map<std::string, std::vector<int>> exactGroups;
for (int i = 0; i < static_cast<int>(rows.size()); i++) {
std::ostringstream key;
for (const auto& c : columns) key << rows[i].at(c) << '\x01';
exactGroups[key.str()].push_back(i);
}
for (const auto& [_, members] : exactGroups) {
if (members.size() <= 1) continue;
for (std::size_t i = 1; i < members.size(); i++) dsu.unite(members[0], members[i]);
for (std::size_t i = 0; i < members.size(); i++) {
for (std::size_t j = i + 1; j < members.size(); j++) {
pairScores.push_back({{"i", members[i]}, {"j", members[j]}, {"score", 1.0}, {"type", "exact"}});
}
}
}
std::unordered_map<std::string, std::vector<int>> blocks;
for (int i = 0; i < static_cast<int>(rows.size()); i++) blocks[blockingKey(rows[i], columns)].push_back(i);
for (const auto& [_, bucket] : blocks) {
if (bucket.size() < 2) continue;
for (std::size_t a = 0; a < bucket.size(); a++) {
for (std::size_t b = a + 1; b < bucket.size(); b++) {
int i = bucket[a], j = bucket[b];
double score = recordSimilarity(rows[i], rows[j], columns);
if (score >= threshold) {
dsu.unite(i, j);
pairScores.push_back({{"i", i}, {"j", j}, {"score", score}, {"type", "fuzzy"}});
}
}
}
}
std::unordered_map<int, std::vector<int>> groups;
for (int i = 0; i < static_cast<int>(rows.size()); i++) groups[dsu.find(i)].push_back(i);
std::vector<std::vector<int>> clusters;
for (auto& [_, members] : groups) {
if (members.size() > 1) {
std::sort(members.begin(), members.end());
clusters.push_back(members);
}
}
std::sort(clusters.begin(), clusters.end(), [](const auto& a, const auto& b) { return a.front() < b.front(); });
return {clusters, pairScores};
}
json buildReport(const std::vector<std::map<std::string, std::string>>& rows,
const std::vector<std::string>& headers,
const std::vector<std::string>& columns,
const FindResult& found,
double threshold) {
json clusters = json::array();
for (std::size_t i = 0; i < found.clusters.size(); i++) {
const auto& members = found.clusters[i];
json records = json::array();
for (std::size_t j = 0; j < members.size(); j++) {
int idx = members[j];
json data = json::object();
for (const auto& h : headers) data[h] = rows[idx].at(h);
records.push_back({
{"row_index", idx},
{"role", j == 0 ? "primary" : "duplicate"},
{"data", data}
});
}
json pairs = json::array();
for (std::size_t a = 0; a < members.size(); a++) {
for (std::size_t b = a + 1; b < members.size(); b++) {
int ia = members[a], ib = members[b];
double score = recordSimilarity(rows[ia], rows[ib], columns);
pairs.push_back({{"row_a", ia}, {"row_b", ib}, {"similarity", score}});
}
}
clusters.push_back({
{"cluster_id", "C" + pad4(static_cast<int>(i + 1))},
{"primary_row_index", members[0]},
{"size", members.size()},
{"matching_columns", columns},
{"records", records},
{"pairwise_similarity", pairs}
});
}
std::vector<double> allScores;
for (const auto& c : clusters) for (const auto& p : c["pairwise_similarity"]) allScores.push_back(p["similarity"].get<double>());
json breakdown = {
{"0.85-0.90", std::count_if(allScores.begin(), allScores.end(), [](double s) { return s >= 0.85 && s < 0.90; })},
{"0.90-0.95", std::count_if(allScores.begin(), allScores.end(), [](double s) { return s >= 0.90 && s < 0.95; })},
{"0.95-1.00", std::count_if(allScores.begin(), allScores.end(), [](double s) { return s >= 0.95 && s <= 1.00; })}
};
return {
{"metadata", {
{"total_records", rows.size()},
{"threshold", threshold},
{"compared_columns", columns}
}},
{"summary", {
{"duplicate_clusters", clusters.size()},
{"total_duplicate_records", [&]() {
int x = 0;
for (const auto& c : clusters) x += static_cast<int>(c["size"].get<int>()) - 1;
return x;
}()},
{"similarity_breakdown", breakdown}
}},
{"clusters", clusters}
};
}
void writeDeduplicated(const std::filesystem::path& path,
const std::vector<std::string>& headers,
const std::vector<std::map<std::string, std::string>>& rows,
const std::vector<std::vector<int>>& clusters) {
std::set<int> duplicates;
for (const auto& c : clusters) for (std::size_t i = 1; i < c.size(); i++) duplicates.insert(c[i]);
std::ofstream out(path);
for (std::size_t i = 0; i < headers.size(); i++) {
if (i) out << ",";
out << headers[i];
}
out << "\n";
for (int i = 0; i < static_cast<int>(rows.size()); i++) {
if (duplicates.count(i)) continue;
for (std::size_t j = 0; j < headers.size(); j++) {
if (j) out << ",";
out << csvEscape(rows[i].at(headers[j]));
}
out << "\n";
}
}
void generateSample(const std::filesystem::path& path) {
std::vector<std::string> firstNames = {"Alice", "Bob", "Carol", "David", "Eva", "Frank", "Grace", "Helen"};
std::vector<std::string> lastNames = {"Smith", "Johnson", "Brown", "Wilson", "Taylor", "Miller", "Davis", "Moore"};
std::vector<std::string> cities = {"Austin", "Boston", "Chicago", "Denver", "Seattle"};
std::vector<std::map<std::string, std::string>> rows;
for (int i = 0; i < 450; i++) {
std::string fn = firstNames[i % firstNames.size()];
std::string ln = lastNames[(i * 3) % lastNames.size()];
std::string city = cities[i % cities.size()];
rows.push_back({
{"id", "R" + pad4(i + 1)},
{"name", fn + " " + ln},
{"email", fn + "." + ln + std::to_string(i) + "@example.com"},
{"city", city},
{"phone", "555-" + std::to_string(1000 + i)}
});
rows.back()["email"] = normalizeText(rows.back()["email"]);
}
for (int i = 0; i < 25; i++) {
auto dup = rows[i];
dup["id"] = "DUPX" + std::to_string(i);
rows.push_back(dup);
}
for (int i = 0; i < 25; i++) {
auto base = rows[100 + i];
base["id"] = "DUPF" + std::to_string(i);
base["name"] = std::regex_replace(base["name"], std::regex("Smith"), "Smiht");
base["name"] = std::regex_replace(base["name"], std::regex("David"), "Davd");
std::transform(base["city"].begin(), base["city"].end(), base["city"].begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
base["email"] = std::regex_replace(base["email"], std::regex("@example\\.com"), "@example.co");
rows.push_back(base);
}
std::ofstream out(path);
std::vector<std::string> headers = {"id", "name", "email", "city", "phone"};
for (std::size_t i = 0; i < headers.size(); i++) {
if (i) out << ",";
out << headers[i];
}
out << "\n";
for (const auto& row : rows) {
for (std::size_t i = 0; i < headers.size(); i++) {
if (i) out << ",";
out << csvEscape(row.at(headers[i]));
}
out << "\n";
}
}
void printSummary(const json& report, const std::filesystem::path& outputPath,
const std::optional<std::filesystem::path>& dedupPath) {
std::cout << "Duplicate Record Finder\n";
std::cout << "=======================\n";
std::cout << "Total records : " << report["metadata"]["total_records"] << "\n";
std::cout << "Duplicate clusters : " << report["summary"]["duplicate_clusters"] << "\n";
std::cout << "Total duplicate rows : " << report["summary"]["total_duplicate_records"] << "\n";
std::cout << "Similarity breakdown :\n";
for (auto it = report["summary"]["similarity_breakdown"].begin(); it != report["summary"]["similarity_breakdown"].end(); ++it) {
std::cout << " " << it.key() << ": " << it.value() << "\n";
}
std::cout << "JSON report saved : " << outputPath.string() << "\n";
if (dedupPath.has_value()) std::cout << "Deduplicated CSV saved: " << dedupPath->string() << "\n";
}
} // namespace
int main(int argc, char** argv) {
ParsedArgs args = parseArgs(argc, argv);
double threshold = 0.85;
if (args.options.count("threshold")) threshold = std::clamp(std::stod(args.options["threshold"]), 0.0, 1.0);
std::filesystem::path outputPath = std::filesystem::absolute(args.options.count("output") ? args.options["output"] : "duplicates_report.json");
try {
std::filesystem::path inputPath;
if (args.positional.empty()) {
inputPath = std::filesystem::absolute("sample_duplicates.csv");
generateSample(inputPath);
std::cout << "No input provided. Generated sample dataset: " << inputPath.string() << "\n";
} else {
inputPath = std::filesystem::absolute(args.positional[0]);
if (!std::filesystem::exists(inputPath)) {
std::cerr << "Input file not found: " << inputPath.string() << "\n";
return 1;
}
}
LoadResult loaded = loadCsv(inputPath);
if (loaded.rows.empty()) {
std::cerr << "No rows found.\n";
return 1;
}
std::vector<std::string> columns;
if (args.options.count("columns")) {
std::stringstream ss(args.options["columns"]);
std::string col;
while (std::getline(ss, col, ',')) {
col = std::regex_replace(col, std::regex(R"(^\s+|\s+$)"), "");
if (!col.empty() && std::find(loaded.headers.begin(), loaded.headers.end(), col) != loaded.headers.end()) columns.push_back(col);
}
} else {
columns = loaded.headers;
}
if (columns.empty()) {
std::cerr << "No valid comparison columns found.\n";
return 1;
}
FindResult found = findDuplicates(loaded.rows, columns, threshold);
json report = buildReport(loaded.rows, loaded.headers, columns, found, threshold);
std::ofstream out(outputPath);
out << report.dump(2) << "\n";
out.close();
std::optional<std::filesystem::path> dedupPath;
if (args.options.count("deduplicate")) {
std::string val = args.options["deduplicate"];
dedupPath = std::filesystem::absolute(val == "true" ? "deduplicated.csv" : val);
writeDeduplicated(*dedupPath, loaded.headers, loaded.rows, found.clusters);
}
printSummary(report, outputPath, dedupPath);
return 0;
} catch (const std::exception& e) {
std::cerr << "Failed: " << e.what() << "\n";
return 1;
}
}