← All tasks
cppcodex/cpp-t1 #10Not a task: already works

Duplicate Record Finder (cpp, written by Codex)

envgap__codex__cpp-t1-10

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/cpp-t1 #10 · read the task the agent was given
Codex wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Duplicate Record Finder

Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Support exact duplicate detection: find rows where all specified columns match exactly
- Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric
- Accept a --columns flag to specify which columns to compare (default: all columns)
- Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85)
- Group duplicates into clusters and assign each cluster an ID
- For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates
- Compute similarity scores for each pair within a cluster
- Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range
- Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns
- Export a deduplicated CSV (keeping only primary records) via --deduplicate flag
- If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it
- Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.22)
project(duplicate_record_finder_cpp VERSION 1.0.0 LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 20)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)

include(FetchContent)

# Pinned dependency
FetchContent_Declare(
  nlohmann_json
  URL https://github.com/nlohmann/json/releases/download/v3.11.3/json.tar.xz
)
FetchContent_MakeAvailable(nlohmann_json)

add_executable(duplicate_finder src/main.cpp)
target_link_libraries(duplicate_finder PRIVATE nlohmann_json::nlohmann_json)
README.md
# Duplicate Record Finder (C++)

Finds exact and near-duplicate records in CSV data using:

- exact matching on selected columns
- fuzzy matching with normalized Levenshtein similarity
- blocking/indexing to avoid full all-pairs comparisons
- cluster construction with primary record selection

## Requirements

- Ubuntu 22.04
- G++ 12+
- CMake 3.22+
- Network access during CMake configure

## Dependencies

- Direct (pinned in `CMakeLists.txt`):
  - `nlohmann/json v3.11.3`
- Transitive:
  - none

## Build

```bash
cmake -S . -B build
cmake --build build
```

## Run

With input:

```bash
./build/duplicate_finder /path/to/data.csv --columns name,email,city --threshold 0.85 --output duplicates_report.json --deduplicate deduplicated.csv
```

No input (generates 500-record sample):

```bash
./build/duplicate_finder
```
src/main.cpp
#include <algorithm>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <map>
#include <numeric>
#include <optional>
#include <regex>
#include <set>
#include <sstream>
#include <string>
#include <unordered_map>
#include <vector>

#include <nlohmann/json.hpp>

using json = nlohmann::json;

namespace {

struct ParsedArgs {
    std::map<std::string, std::string> options;
    std::vector<std::string> positional;
};

struct LoadResult {
    std::vector<std::string> headers;
    std::vector<std::map<std::string, std::string>> rows;
};

struct FindResult {
    std::vector<std::vector<int>> clusters;
    std::vector<json> pairScores;
};

std::string pad4(int n) {
    std::ostringstream out;
    out << std::setw(4) << std::setfill('0') << n;
    return out.str();
}

ParsedArgs parseArgs(int argc, char** argv) {
    ParsedArgs out;
    for (int i = 1; i < argc; i++) {
        std::string token = argv[i];
        if (token.rfind("--", 0) == 0) {
            std::string key = token.substr(2);
            if (i + 1 < argc && std::string(argv[i + 1]).rfind("--", 0) != 0) out.options[key] = argv[++i];
            else out.options[key] = "true";
        } else {
            out.positional.push_back(token);
        }
    }
    return out;
}

std::vector<std::string> parseCsvLine(const std::string& line) {
    std::vector<std::string> out;
    std::string current;
    bool inQuotes = false;
    for (std::size_t i = 0; i < line.size(); i++) {
        char ch = line[i];
        if (ch == '"') {
            if (inQuotes && i + 1 < line.size() && line[i + 1] == '"') {
                current += '"';
                i++;
            } else {
                inQuotes = !inQuotes;
            }
        } else if (ch == ',' && !inQuotes) {
            out.push_back(current);
            current.clear();
        } else {
            current += ch;
        }
    }
    out.push_back(current);
    return out;
}

std::string csvEscape(const std::string& value) {
    if (value.find(',') != std::string::npos || value.find('"') != std::string::npos || value.find('\n') != std::string::npos) {
        std::string out = "\"";
        for (char c : value) {
            if (c == '"') out += "\"\"";
            else out += c;
        }
        out += "\"";
        return out;
    }
    return value;
}

LoadResult loadCsv(const std::filesystem::path& path) {
    std::ifstream in(path);
    if (!in.is_open()) throw std::runtime_error("Failed to open CSV: " + path.string());
    std::string line;
    if (!std::getline(in, line)) return {};
    if (!line.empty() && static_cast<unsigned char>(line[0]) == 0xEF && line.size() >= 3 &&
        static_cast<unsigned char>(line[1]) == 0xBB && static_cast<unsigned char>(line[2]) == 0xBF) {
        line = line.substr(3);
    }
    auto headers = parseCsvLine(line);
    std::vector<std::map<std::string, std::string>> rows;
    int idx = 0;
    while (std::getline(in, line)) {
        if (!line.empty() && line.back() == '\r') line.pop_back();
        if (line.empty()) continue;
        auto parts = parseCsvLine(line);
        std::map<std::string, std::string> row;
        row["__index"] = std::to_string(idx++);
        for (std::size_t i = 0; i < headers.size(); i++) row[headers[i]] = i < parts.size() ? parts[i] : "";
        rows.push_back(row);
    }
    return {headers, rows};
}

std::string normalizeText(const std::string& value) {
    std::string out = value;
    std::transform(out.begin(), out.end(), out.begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
    out = std::regex_replace(out, std::regex(R"(\.)"), "");
    out = std::regex_replace(out, std::regex(R"(\s+)"), " ");
    out = std::regex_replace(out, std::regex(R"(^\s+|\s+$)"), "");
    return out;
}

int levenshtein(const std::string& a, const std::string& b) {
    if (a == b) return 0;
    if (a.empty()) return static_cast<int>(b.size());
    if (b.empty()) return static_cast<int>(a.size());
    std::vector<std::vector<int>> dp(a.size() + 1, std::vector<int>(b.size() + 1, 0));
    for (std::size_t i = 0; i <= a.size(); i++) dp[i][0] = static_cast<int>(i);
    for (std::size_t j = 0; j <= b.size(); j++) dp[0][j] = static_cast<int>(j);
    for (std::size_t i = 1; i <= a.size(); i++) {
        for (std::size_t j = 1; j <= b.size(); j++) {
            int cost = a[i - 1] == b[j - 1] ? 0 : 1;
            dp[i][j] = std::min({dp[i - 1][j] + 1, dp[i][j - 1] + 1, dp[i - 1][j - 1] + cost});
        }
    }
    return dp[a.size()][b.size()];
}

double similarity(const std::string& a, const std::string& b) {
    std::string x = normalizeText(a);
    std::string y = normalizeText(b);
    std::size_t m = std::max(x.size(), y.size());
    if (m == 0) return 1.0;
    return 1.0 - static_cast<double>(levenshtein(x, y)) / static_cast<double>(m);
}

double recordSimilarity(const std::map<std::string, std::string>& a,
                        const std::map<std::string, std::string>& b,
                        const std::vector<std::string>& columns) {
    if (columns.empty()) return 1.0;
    double sum = 0.0;
    for (const auto& c : columns) sum += similarity(a.at(c), b.at(c));
    return sum / columns.size();
}

std::string blockingKey(const std::map<std::string, std::string>& row, const std::vector<std::string>& columns) {
    std::ostringstream out;
    for (std::size_t i = 0; i < columns.size(); i++) {
        std::string t = normalizeText(row.at(columns[i]));
        t = std::regex_replace(t, std::regex(R"([^a-z0-9])"), "");
        if (t.size() > 4) t = t.substr(0, 4);
        if (i) out << "|";
        out << t;
    }
    return out.str();
}

struct DSU {
    std::vector<int> parent;
    std::vector<int> rank;

    explicit DSU(int n) : parent(n), rank(n, 0) {
        std::iota(parent.begin(), parent.end(), 0);
    }

    int find(int x) {
        if (parent[x] != x) parent[x] = find(parent[x]);
        return parent[x];
    }

    void unite(int a, int b) {
        int ra = find(a), rb = find(b);
        if (ra == rb) return;
        if (rank[ra] < rank[rb]) std::swap(ra, rb);
        parent[rb] = ra;
        if (rank[ra] == rank[rb]) rank[ra]++;
    }
};

FindResult findDuplicates(const std::vector<std::map<std::string, std::string>>& rows,
                          const std::vector<std::string>& columns,
                          double threshold) {
    DSU dsu(static_cast<int>(rows.size()));
    std::vector<json> pairScores;

    std::unordered_map<std::string, std::vector<int>> exactGroups;
    for (int i = 0; i < static_cast<int>(rows.size()); i++) {
        std::ostringstream key;
        for (const auto& c : columns) key << rows[i].at(c) << '\x01';
        exactGroups[key.str()].push_back(i);
    }
    for (const auto& [_, members] : exactGroups) {
        if (members.size() <= 1) continue;
        for (std::size_t i = 1; i < members.size(); i++) dsu.unite(members[0], members[i]);
        for (std::size_t i = 0; i < members.size(); i++) {
            for (std::size_t j = i + 1; j < members.size(); j++) {
                pairScores.push_back({{"i", members[i]}, {"j", members[j]}, {"score", 1.0}, {"type", "exact"}});
            }
        }
    }

    std::unordered_map<std::string, std::vector<int>> blocks;
    for (int i = 0; i < static_cast<int>(rows.size()); i++) blocks[blockingKey(rows[i], columns)].push_back(i);
    for (const auto& [_, bucket] : blocks) {
        if (bucket.size() < 2) continue;
        for (std::size_t a = 0; a < bucket.size(); a++) {
            for (std::size_t b = a + 1; b < bucket.size(); b++) {
                int i = bucket[a], j = bucket[b];
                double score = recordSimilarity(rows[i], rows[j], columns);
                if (score >= threshold) {
                    dsu.unite(i, j);
                    pairScores.push_back({{"i", i}, {"j", j}, {"score", score}, {"type", "fuzzy"}});
                }
            }
        }
    }

    std::unordered_map<int, std::vector<int>> groups;
    for (int i = 0; i < static_cast<int>(rows.size()); i++) groups[dsu.find(i)].push_back(i);
    std::vector<std::vector<int>> clusters;
    for (auto& [_, members] : groups) {
        if (members.size() > 1) {
            std::sort(members.begin(), members.end());
            clusters.push_back(members);
        }
    }
    std::sort(clusters.begin(), clusters.end(), [](const auto& a, const auto& b) { return a.front() < b.front(); });
    return {clusters, pairScores};
}

json buildReport(const std::vector<std::map<std::string, std::string>>& rows,
                 const std::vector<std::string>& headers,
                 const std::vector<std::string>& columns,
                 const FindResult& found,
                 double threshold) {
    json clusters = json::array();
    for (std::size_t i = 0; i < found.clusters.size(); i++) {
        const auto& members = found.clusters[i];
        json records = json::array();
        for (std::size_t j = 0; j < members.size(); j++) {
            int idx = members[j];
            json data = json::object();
            for (const auto& h : headers) data[h] = rows[idx].at(h);
            records.push_back({
                {"row_index", idx},
                {"role", j == 0 ? "primary" : "duplicate"},
                {"data", data}
            });
        }
        json pairs = json::array();
        for (std::size_t a = 0; a < members.size(); a++) {
            for (std::size_t b = a + 1; b < members.size(); b++) {
                int ia = members[a], ib = members[b];
                double score = recordSimilarity(rows[ia], rows[ib], columns);
                pairs.push_back({{"row_a", ia}, {"row_b", ib}, {"similarity", score}});
            }
        }
        clusters.push_back({
            {"cluster_id", "C" + pad4(static_cast<int>(i + 1))},
            {"primary_row_index", members[0]},
            {"size", members.size()},
            {"matching_columns", columns},
            {"records", records},
            {"pairwise_similarity", pairs}
        });
    }

    std::vector<double> allScores;
    for (const auto& c : clusters) for (const auto& p : c["pairwise_similarity"]) allScores.push_back(p["similarity"].get<double>());
    json breakdown = {
        {"0.85-0.90", std::count_if(allScores.begin(), allScores.end(), [](double s) { return s >= 0.85 && s < 0.90; })},
        {"0.90-0.95", std::count_if(allScores.begin(), allScores.end(), [](double s) { return s >= 0.90 && s < 0.95; })},
        {"0.95-1.00", std::count_if(allScores.begin(), allScores.end(), [](double s) { return s >= 0.95 && s <= 1.00; })}
    };

    return {
        {"metadata", {
            {"total_records", rows.size()},
            {"threshold", threshold},
            {"compared_columns", columns}
        }},
        {"summary", {
            {"duplicate_clusters", clusters.size()},
            {"total_duplicate_records", [&]() {
                int x = 0;
                for (const auto& c : clusters) x += static_cast<int>(c["size"].get<int>()) - 1;
                return x;
            }()},
            {"similarity_breakdown", breakdown}
        }},
        {"clusters", clusters}
    };
}

void writeDeduplicated(const std::filesystem::path& path,
                       const std::vector<std::string>& headers,
                       const std::vector<std::map<std::string, std::string>>& rows,
                       const std::vector<std::vector<int>>& clusters) {
    std::set<int> duplicates;
    for (const auto& c : clusters) for (std::size_t i = 1; i < c.size(); i++) duplicates.insert(c[i]);
    std::ofstream out(path);
    for (std::size_t i = 0; i < headers.size(); i++) {
        if (i) out << ",";
        out << headers[i];
    }
    out << "\n";
    for (int i = 0; i < static_cast<int>(rows.size()); i++) {
        if (duplicates.count(i)) continue;
        for (std::size_t j = 0; j < headers.size(); j++) {
            if (j) out << ",";
            out << csvEscape(rows[i].at(headers[j]));
        }
        out << "\n";
    }
}

void generateSample(const std::filesystem::path& path) {
    std::vector<std::string> firstNames = {"Alice", "Bob", "Carol", "David", "Eva", "Frank", "Grace", "Helen"};
    std::vector<std::string> lastNames = {"Smith", "Johnson", "Brown", "Wilson", "Taylor", "Miller", "Davis", "Moore"};
    std::vector<std::string> cities = {"Austin", "Boston", "Chicago", "Denver", "Seattle"};
    std::vector<std::map<std::string, std::string>> rows;
    for (int i = 0; i < 450; i++) {
        std::string fn = firstNames[i % firstNames.size()];
        std::string ln = lastNames[(i * 3) % lastNames.size()];
        std::string city = cities[i % cities.size()];
        rows.push_back({
            {"id", "R" + pad4(i + 1)},
            {"name", fn + " " + ln},
            {"email", fn + "." + ln + std::to_string(i) + "@example.com"},
            {"city", city},
            {"phone", "555-" + std::to_string(1000 + i)}
        });
        rows.back()["email"] = normalizeText(rows.back()["email"]);
    }
    for (int i = 0; i < 25; i++) {
        auto dup = rows[i];
        dup["id"] = "DUPX" + std::to_string(i);
        rows.push_back(dup);
    }
    for (int i = 0; i < 25; i++) {
        auto base = rows[100 + i];
        base["id"] = "DUPF" + std::to_string(i);
        base["name"] = std::regex_replace(base["name"], std::regex("Smith"), "Smiht");
        base["name"] = std::regex_replace(base["name"], std::regex("David"), "Davd");
        std::transform(base["city"].begin(), base["city"].end(), base["city"].begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
        base["email"] = std::regex_replace(base["email"], std::regex("@example\\.com"), "@example.co");
        rows.push_back(base);
    }

    std::ofstream out(path);
    std::vector<std::string> headers = {"id", "name", "email", "city", "phone"};
    for (std::size_t i = 0; i < headers.size(); i++) {
        if (i) out << ",";
        out << headers[i];
    }
    out << "\n";
    for (const auto& row : rows) {
        for (std::size_t i = 0; i < headers.size(); i++) {
            if (i) out << ",";
            out << csvEscape(row.at(headers[i]));
        }
        out << "\n";
    }
}

void printSummary(const json& report, const std::filesystem::path& outputPath,
                  const std::optional<std::filesystem::path>& dedupPath) {
    std::cout << "Duplicate Record Finder\n";
    std::cout << "=======================\n";
    std::cout << "Total records         : " << report["metadata"]["total_records"] << "\n";
    std::cout << "Duplicate clusters    : " << report["summary"]["duplicate_clusters"] << "\n";
    std::cout << "Total duplicate rows  : " << report["summary"]["total_duplicate_records"] << "\n";
    std::cout << "Similarity breakdown  :\n";
    for (auto it = report["summary"]["similarity_breakdown"].begin(); it != report["summary"]["similarity_breakdown"].end(); ++it) {
        std::cout << "  " << it.key() << ": " << it.value() << "\n";
    }
    std::cout << "JSON report saved     : " << outputPath.string() << "\n";
    if (dedupPath.has_value()) std::cout << "Deduplicated CSV saved: " << dedupPath->string() << "\n";
}

} // namespace

int main(int argc, char** argv) {
    ParsedArgs args = parseArgs(argc, argv);
    double threshold = 0.85;
    if (args.options.count("threshold")) threshold = std::clamp(std::stod(args.options["threshold"]), 0.0, 1.0);
    std::filesystem::path outputPath = std::filesystem::absolute(args.options.count("output") ? args.options["output"] : "duplicates_report.json");

    try {
        std::filesystem::path inputPath;
        if (args.positional.empty()) {
            inputPath = std::filesystem::absolute("sample_duplicates.csv");
            generateSample(inputPath);
            std::cout << "No input provided. Generated sample dataset: " << inputPath.string() << "\n";
        } else {
            inputPath = std::filesystem::absolute(args.positional[0]);
            if (!std::filesystem::exists(inputPath)) {
                std::cerr << "Input file not found: " << inputPath.string() << "\n";
                return 1;
            }
        }

        LoadResult loaded = loadCsv(inputPath);
        if (loaded.rows.empty()) {
            std::cerr << "No rows found.\n";
            return 1;
        }

        std::vector<std::string> columns;
        if (args.options.count("columns")) {
            std::stringstream ss(args.options["columns"]);
            std::string col;
            while (std::getline(ss, col, ',')) {
                col = std::regex_replace(col, std::regex(R"(^\s+|\s+$)"), "");
                if (!col.empty() && std::find(loaded.headers.begin(), loaded.headers.end(), col) != loaded.headers.end()) columns.push_back(col);
            }
        } else {
            columns = loaded.headers;
        }
        if (columns.empty()) {
            std::cerr << "No valid comparison columns found.\n";
            return 1;
        }

        FindResult found = findDuplicates(loaded.rows, columns, threshold);
        json report = buildReport(loaded.rows, loaded.headers, columns, found, threshold);
        std::ofstream out(outputPath);
        out << report.dump(2) << "\n";
        out.close();

        std::optional<std::filesystem::path> dedupPath;
        if (args.options.count("deduplicate")) {
            std::string val = args.options["deduplicate"];
            dedupPath = std::filesystem::absolute(val == "true" ? "deduplicated.csv" : val);
            writeDeduplicated(*dedupPath, loaded.headers, loaded.rows, found.clusters);
        }

        printSummary(report, outputPath, dedupPath);
        return 0;
    } catch (const std::exception& e) {
        std::cerr << "Failed: " << e.what() << "\n";
        return 1;
    }
}