← All tasks
cppclaude-code/cpp-t1 #10Not a task: already works

Duplicate Record Finder (cpp, written by Claude Code)

envgap__claude-code__cpp-t1-10

Written by a coding agent; not on GitHubWritten 2026-02-27

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/cpp-t1 #10 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Duplicate Record Finder

Write a program that identifies duplicate and near-duplicate records in tabular datasets using exact matching, fuzzy string matching, and configurable similarity thresholds.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Support exact duplicate detection: find rows where all specified columns match exactly
- Support fuzzy duplicate detection: find rows where string columns are similar above a configurable threshold (0.0 to 1.0) using edit distance or similar string similarity metric
- Accept a --columns flag to specify which columns to compare (default: all columns)
- Accept a --threshold flag for fuzzy matching sensitivity (default: 0.85)
- Group duplicates into clusters and assign each cluster an ID
- For each duplicate cluster, designate the first occurrence as the "primary" record and subsequent ones as duplicates
- Compute similarity scores for each pair within a cluster
- Print a summary to console: total records, number of duplicate clusters, total duplicate records, breakdown by similarity range
- Save a detailed report as JSON with --output flag (default: duplicates_report.json) containing each cluster with its records, similarity scores, and matching columns
- Export a deduplicated CSV (keeping only primary records) via --deduplicate flag
- If no input file is given, generate a sample CSV with 500 records including exact duplicates, near-duplicates (typos, case differences, abbreviations), and unique records, then analyze it
- Handle large files efficiently by using blocking/indexing strategies to avoid comparing every pair

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.20)
project(DuplicateRecordFinder VERSION 1.0.0 LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)

include(FetchContent)

# rapidfuzz-cpp (latest)
FetchContent_Declare(
    rapidfuzz
    GIT_REPOSITORY https://github.com/rapidfuzz/rapidfuzz-cpp.git
    GIT_TAG main
)
FetchContent_MakeAvailable(rapidfuzz)

# nlohmann/json 3.11.3
FetchContent_Declare(
    nlohmann_json
    GIT_REPOSITORY https://github.com/nlohmann/json.git
    GIT_TAG v3.11.3
)
FetchContent_MakeAvailable(nlohmann_json)

# fast-cpp-csv-parser (master)
FetchContent_Declare(
    fast_csv
    GIT_REPOSITORY https://github.com/ben-strasser/fast-cpp-csv-parser.git
    GIT_TAG master
)
FetchContent_MakeAvailable(fast_csv)

add_executable(finder main.cpp)

target_include_directories(finder PRIVATE ${fast_csv_SOURCE_DIR})

target_link_libraries(finder
    PRIVATE
    rapidfuzz::rapidfuzz
    nlohmann_json::nlohmann_json
)
main.cpp
/**
 * Duplicate Record Finder - C++ Trial 1
 * Uses rapidfuzz-cpp + nlohmann/json + fast-cpp-csv-parser.
 * Supports exact matching, fuzzy string matching with configurable thresholds,
 * and multiple comparison strategies (Levenshtein, Jaro-Winkler, Cosine).
 */

#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
#include <vector>
#include <map>
#include <set>
#include <cmath>
#include <algorithm>
#include <numeric>
#include <functional>
#include <iomanip>

#include <nlohmann/json.hpp>
#include <csv.h>
#include <rapidfuzz/fuzz.hpp>
#include <rapidfuzz/distance.hpp>

using json = nlohmann::json;

// --- Data Structures ---

struct Record {
    std::map<std::string, std::string> fields;
};

struct ColumnRule {
    std::string strategy;
    double weight;
    bool exact;
};

struct ColumnScore {
    double score;
    std::string strategy;
};

struct PairResult {
    double overall;
    std::map<std::string, ColumnScore> columns;
};

struct DuplicateGroup {
    std::vector<int> indices;
    std::map<std::string, PairResult> pairScores;
};

// --- Similarity Strategies ---

double levenshteinSimilarity(const std::string& s1, const std::string& s2) {
    if (s1.empty() && s2.empty()) return 1.0;
    double dist = static_cast<double>(rapidfuzz::levenshtein_distance(s1, s2));
    double maxLen = static_cast<double>(std::max(s1.size(), s2.size()));
    if (maxLen == 0.0) return 1.0;
    return 1.0 - dist / maxLen;
}

double jaroSimilarity(const std::string& s1, const std::string& s2) {
    if (s1 == s2) return 1.0;
    if (s1.empty() || s2.empty()) return 0.0;

    int maxDist = std::max(0, (int)(std::max(s1.size(), s2.size()) / 2) - 1);
    std::vector<bool> s1Matches(s1.size(), false);
    std::vector<bool> s2Matches(s2.size(), false);
    int matches = 0, transpositions = 0;

    for (int i = 0; i < (int)s1.size(); i++) {
        int start = std::max(0, i - maxDist);
        int end = std::min(i + maxDist + 1, (int)s2.size());
        for (int j = start; j < end; j++) {
            if (s2Matches[j] || s1[i] != s2[j]) continue;
            s1Matches[i] = true;
            s2Matches[j] = true;
            matches++;
            break;
        }
    }
    if (matches == 0) return 0.0;

    int k = 0;
    for (int i = 0; i < (int)s1.size(); i++) {
        if (!s1Matches[i]) continue;
        while (!s2Matches[k]) k++;
        if (s1[i] != s2[k]) transpositions++;
        k++;
    }

    return (static_cast<double>(matches) / s1.size()
          + static_cast<double>(matches) / s2.size()
          + static_cast<double>(matches - transpositions / 2.0) / matches) / 3.0;
}

double jaroWinklerSimilarity(const std::string& s1, const std::string& s2) {
    double jaro = jaroSimilarity(s1, s2);
    int prefix = 0;
    int limit = std::min(4, (int)std::min(s1.size(), s2.size()));
    for (int i = 0; i < limit; i++) {
        if (s1[i] == s2[i]) prefix++;
        else break;
    }
    return jaro + prefix * 0.1 * (1.0 - jaro);
}

double cosineSimilarity(const std::string& s1, const std::string& s2) {
    auto toLower = [](std::string s) {
        std::transform(s.begin(), s.end(), s.begin(), ::tolower);
        return s;
    };
    std::string a = toLower(s1);
    std::string b = toLower(s2);
    if (a.empty() || b.empty()) return 0.0;

    auto bigrams = [](const std::string& s) {
        std::map<std::string, int> bg;
        for (size_t i = 0; i + 1 < s.size(); i++) {
            bg[s.substr(i, 2)]++;
        }
        return bg;
    };

    auto bg1 = bigrams(a);
    auto bg2 = bigrams(b);
    if (bg1.empty() || bg2.empty()) return (a == b) ? 1.0 : 0.0;

    std::set<std::string> allKeys;
    for (auto& p : bg1) allKeys.insert(p.first);
    for (auto& p : bg2) allKeys.insert(p.first);

    double dot = 0, mag1 = 0, mag2 = 0;
    for (auto& k : allKeys) {
        double v1 = bg1.count(k) ? bg1[k] : 0;
        double v2 = bg2.count(k) ? bg2[k] : 0;
        dot += v1 * v2;
        mag1 += v1 * v1;
        mag2 += v2 * v2;
    }
    mag1 = std::sqrt(mag1);
    mag2 = std::sqrt(mag2);
    if (mag1 == 0 || mag2 == 0) return 0.0;
    return dot / (mag1 * mag2);
}

using SimilarityFn = std::function<double(const std::string&, const std::string&)>;

std::map<std::string, SimilarityFn> STRATEGIES = {
    {"levenshtein", levenshteinSimilarity},
    {"jaro-winkler", jaroWinklerSimilarity},
    {"cosine", cosineSimilarity},
};

// --- Core Logic ---

std::string trim(const std::string& s) {
    size_t start = s.find_first_not_of(" \t\r\n");
    size_t end = s.find_last_not_of(" \t\r\n");
    return (start == std::string::npos) ? "" : s.substr(start, end - start + 1);
}

std::string toLowerStr(const std::string& s) {
    std::string result = s;
    std::transform(result.begin(), result.end(), result.begin(), ::tolower);
    return result;
}

double roundTo4(double v) {
    return std::round(v * 10000.0) / 10000.0;
}

PairResult computeRecordSimilarity(const Record& rec1, const Record& rec2,
                                    const std::map<std::string, ColumnRule>& columnRules,
                                    const std::string& defaultStrategy) {
    PairResult result;
    double weightedScore = 0.0, weightsSum = 0.0;

    for (auto& [col, rule] : columnRules) {
        std::string strategyName = rule.strategy.empty() ? defaultStrategy : rule.strategy;
        double weight = rule.weight;

        std::string val1 = trim(rec1.fields.count(col) ? rec1.fields.at(col) : "");
        std::string val2 = trim(rec2.fields.count(col) ? rec2.fields.at(col) : "");

        double score;
        if (rule.exact) {
            score = (toLowerStr(val1) == toLowerStr(val2)) ? 1.0 : 0.0;
        } else {
            auto it = STRATEGIES.find(strategyName);
            SimilarityFn fn = (it != STRATEGIES.end()) ? it->second : levenshteinSimilarity;
            score = fn(val1, val2);
        }

        result.columns[col] = {roundTo4(score), strategyName};
        weightedScore += score * weight;
        weightsSum += weight;
    }

    result.overall = weightsSum > 0 ? roundTo4(weightedScore / weightsSum) : 0.0;
    return result;
}

std::vector<std::vector<int>> findExactDuplicates(const std::vector<Record>& records) {
    std::map<std::string, std::vector<int>> groups;
    for (int i = 0; i < (int)records.size(); i++) {
        std::string key;
        std::map<std::string, std::string> sorted(records[i].fields.begin(), records[i].fields.end());
        for (auto& [k, v] : sorted) {
            key += k + "=" + v + "|";
        }
        groups[key].push_back(i);
    }

    std::vector<std::vector<int>> result;
    for (auto& [key, indices] : groups) {
        if (indices.size() > 1) {
            result.push_back(indices);
        }
    }
    return result;
}

// Union-Find
class UnionFind {
public:
    std::vector<int> parent;
    UnionFind(int n) : parent(n) { std::iota(parent.begin(), parent.end(), 0); }
    int find(int x) {
        while (parent[x] != x) { parent[x] = parent[parent[x]]; x = parent[x]; }
        return x;
    }
    void unite(int a, int b) {
        int ra = find(a), rb = find(b);
        if (ra != rb) parent[ra] = rb;
    }
};

std::vector<DuplicateGroup> findFuzzyDuplicates(const std::vector<Record>& records,
                                                  const std::map<std::string, ColumnRule>& columnRules,
                                                  double threshold,
                                                  const std::string& defaultStrategy) {
    int n = (int)records.size();
    UnionFind uf(n);
    std::map<std::string, PairResult> pairScores;

    for (int i = 0; i < n; i++) {
        for (int j = i + 1; j < n; j++) {
            PairResult result = computeRecordSimilarity(records[i], records[j], columnRules, defaultStrategy);
            if (result.overall >= threshold) {
                uf.unite(i, j);
                pairScores[std::to_string(i) + "-" + std::to_string(j)] = result;
            }
        }
    }

    std::map<int, std::vector<int>> groupMap;
    for (int i = 0; i < n; i++) {
        groupMap[uf.find(i)].push_back(i);
    }

    std::vector<DuplicateGroup> duplicateGroups;
    for (auto& [root, members] : groupMap) {
        if (members.size() > 1) {
            DuplicateGroup group;
            group.indices = members;
            std::sort(group.indices.begin(), group.indices.end());
            for (size_t a = 0; a < members.size(); a++) {
                for (size_t b = a + 1; b < members.size(); b++) {
                    int lo = std::min(members[a], members[b]);
                    int hi = std::max(members[a], members[b]);
                    std::string key = std::to_string(lo) + "-" + std::to_string(hi);
                    if (pairScores.count(key)) {
                        group.pairScores[key] = pairScores[key];
                    }
                }
            }
            duplicateGroups.push_back(group);
        }
    }
    return duplicateGroups;
}

std::map<std::string, ColumnRule> buildColumnRules(const std::vector<std::string>& columns,
                                                    const std::string& strategy,
                                                    const std::vector<std::string>& exactCols) {
    std::map<std::string, ColumnRule> rules;
    std::set<std::string> exactSet(exactCols.begin(), exactCols.end());
    for (auto& col : columns) {
        std::string lower = toLowerStr(col);
        if (lower == "id" || lower == "index") continue;
        rules[col] = {strategy, 1.0, exactSet.count(col) > 0};
    }
    return rules;
}

// --- Sample Data ---

std::vector<Record> getSampleRecords() {
    return {
        {{{ "id","1"},{"first_name","John"},{"last_name","Smith"},{"email","john.smith@email.com"},{"phone","555-0101"},{"city","New York"}}},
        {{{ "id","2"},{"first_name","John"},{"last_name","Smith"},{"email","john.smith@email.com"},{"phone","555-0101"},{"city","New York"}}},
        {{{ "id","3"},{"first_name","Jon"},{"last_name","Smyth"},{"email","jon.smyth@email.com"},{"phone","555-0101"},{"city","New York"}}},
        {{{ "id","4"},{"first_name","Jane"},{"last_name","Doe"},{"email","jane.doe@email.com"},{"phone","555-0202"},{"city","Los Angeles"}}},
        {{{ "id","5"},{"first_name","Jane"},{"last_name","Doe"},{"email","jane.doe@email.com"},{"phone","555-0202"},{"city","Los Angeles"}}},
        {{{ "id","6"},{"first_name","Jayne"},{"last_name","Doe"},{"email","jayne.doe@email.com"},{"phone","555-0203"},{"city","Los Angeles"}}},
        {{{ "id","7"},{"first_name","Robert"},{"last_name","Johnson"},{"email","r.johnson@email.com"},{"phone","555-0303"},{"city","Chicago"}}},
        {{{ "id","8"},{"first_name","Bob"},{"last_name","Johnson"},{"email","bob.johnson@email.com"},{"phone","555-0304"},{"city","Chicago"}}},
        {{{ "id","9"},{"first_name","Alice"},{"last_name","Williams"},{"email","alice.w@email.com"},{"phone","555-0404"},{"city","Houston"}}},
        {{{ "id","10"},{"first_name","Alice"},{"last_name","Willams"},{"email","alice.w@email.com"},{"phone","555-0404"},{"city","Houston"}}},
        {{{ "id","11"},{"first_name","Michael"},{"last_name","Brown"},{"email","m.brown@email.com"},{"phone","555-0505"},{"city","Phoenix"}}},
        {{{ "id","12"},{"first_name","Emily"},{"last_name","Davis"},{"email","emily.d@email.com"},{"phone","555-0606"},{"city","Philadelphia"}}},
        {{{ "id","13"},{"first_name","Emilie"},{"last_name","Davis"},{"email","emilie.davis@email.com"},{"phone","555-0607"},{"city","Philadelphia"}}},
        {{{ "id","14"},{"first_name","David"},{"last_name","Garcia"},{"email","d.garcia@email.com"},{"phone","555-0707"},{"city","San Antonio"}}},
        {{{ "id","15"},{"first_name","David"},{"last_name","Garcia"},{"email","d.garcia@email.com"},{"phone","555-0707"},{"city","San Antonio"}}},
    };
}

void generateSampleCsv(const std::string& outputPath) {
    auto records = getSampleRecords();
    std::ofstream out(outputPath);
    if (!out.is_open()) {
        std::cerr << "Error: Cannot create " << outputPath << std::endl;
        return;
    }
    out << "id,first_name,last_name,email,phone,city\n";
    for (auto& rec : records) {
        out << rec.fields.at("id") << ","
            << rec.fields.at("first_name") << ","
            << rec.fields.at("last_name") << ","
            << rec.fields.at("email") << ","
            << rec.fields.at("phone") << ","
            << rec.fields.at("city") << "\n";
    }
    out.close();
    std::cout << "Sample dataset generated: " << outputPath << " (" << records.size() << " records)" << std::endl;
}

// --- CSV Loading ---

std::vector<Record> loadCsv(const std::string& filePath, std::vector<std::string>& columns) {
    std::vector<Record> records;

    // Read header first
    std::ifstream headerFile(filePath);
    std::string headerLine;
    std::getline(headerFile, headerLine);
    headerFile.close();

    std::stringstream ss(headerLine);
    std::string col;
    while (std::getline(ss, col, ',')) {
        col = trim(col);
        columns.push_back(col);
    }

    // Use fast-cpp-csv-parser with 6 columns
    io::CSVReader<6> reader(filePath);
    reader.read_header(io::ignore_extra_column, "id", "first_name", "last_name", "email", "phone", "city");

    std::string id, firstName, lastName, email, phone, city;
    while (reader.read_row(id, firstName, lastName, email, phone, city)) {
        Record rec;
        rec.fields["id"] = id;
        rec.fields["first_name"] = firstName;
        rec.fields["last_name"] = lastName;
        rec.fields["email"] = email;
        rec.fields["phone"] = phone;
        rec.fields["city"] = city;
        records.push_back(rec);
    }
    return records;
}

// --- Reporting ---

void printConsoleReport(const std::vector<Record>& records,
                        const std::vector<std::vector<int>>& exactGroups,
                        const std::vector<DuplicateGroup>& fuzzyGroups) {
    std::string sep(70, '=');
    std::cout << "\n" << sep << std::endl;
    std::cout << "  DUPLICATE RECORD FINDER - REPORT (rapidfuzz-cpp + nlohmann/json + fast-csv)" << std::endl;
    std::cout << sep << std::endl;
    std::cout << "\nTotal records analyzed: " << records.size() << std::endl;

    std::cout << "\n--- Exact Duplicates: " << exactGroups.size() << " group(s) ---" << std::endl;
    for (size_t i = 0; i < exactGroups.size(); i++) {
        std::cout << "\n  Group " << (i + 1) << " (" << exactGroups[i].size() << " records):" << std::endl;
        for (int idx : exactGroups[i]) {
            std::cout << "    Row " << idx << ": ";
            for (auto& [k, v] : records[idx].fields) {
                std::cout << k << "=" << v << " ";
            }
            std::cout << std::endl;
        }
    }

    std::cout << "\n--- Fuzzy Duplicate Groups: " << fuzzyGroups.size() << " group(s) ---" << std::endl;
    for (size_t i = 0; i < fuzzyGroups.size(); i++) {
        auto& group = fuzzyGroups[i];
        std::cout << "\n  Group " << (i + 1) << " (" << group.indices.size() << " records):" << std::endl;
        for (int idx : group.indices) {
            std::cout << "    Row " << idx << ": ";
            for (auto& [k, v] : records[idx].fields) {
                std::cout << k << "=" << v << " ";
            }
            std::cout << std::endl;
        }
        for (auto& [pairKey, scores] : group.pairScores) {
            std::cout << "    Pair " << pairKey << ": overall=" << scores.overall << std::endl;
            for (auto& [col, info] : scores.columns) {
                std::cout << "      " << col << ": " << info.score << " (" << info.strategy << ")" << std::endl;
            }
        }
    }
    std::cout << "\n" << sep << std::endl;
}

void generateJsonReport(const std::vector<Record>& records,
                         const std::vector<std::vector<int>>& exactGroups,
                         const std::vector<DuplicateGroup>& fuzzyGroups,
                         const std::string& outputPath) {
    json report;
    report["summary"]["total_records"] = records.size();
    report["summary"]["exact_duplicate_groups"] = exactGroups.size();
    report["summary"]["fuzzy_duplicate_groups"] = fuzzyGroups.size();

    json exactArr = json::array();
    for (auto& group : exactGroups) {
        json g;
        json recs = json::array();
        for (int idx : group) {
            json r;
            r["row_index"] = idx;
            r["data"] = records[idx].fields;
            recs.push_back(r);
        }
        g["records"] = recs;
        exactArr.push_back(g);
    }
    report["exact_duplicates"] = exactArr;

    json fuzzyArr = json::array();
    for (auto& group : fuzzyGroups) {
        json g;
        json recs = json::array();
        for (int idx : group.indices) {
            json r;
            r["row_index"] = idx;
            r["data"] = records[idx].fields;
            recs.push_back(r);
        }
        g["records"] = recs;
        json ps;
        for (auto& [key, pr] : group.pairScores) {
            json pj;
            pj["overall"] = pr.overall;
            json cols;
            for (auto& [col, cs] : pr.columns) {
                cols[col] = {{"score", cs.score}, {"strategy", cs.strategy}};
            }
            pj["columns"] = cols;
            ps[key] = pj;
        }
        g["pair_scores"] = ps;
        fuzzyArr.push_back(g);
    }
    report["fuzzy_duplicates"] = fuzzyArr;

    std::ofstream out(outputPath);
    out << report.dump(2) << std::endl;
    out.close();
    std::cout << "\nJSON report saved to: " << outputPath << std::endl;
}

// --- Argument Parsing ---

struct Config {
    std::string input;
    double threshold = 0.8;
    std::string strategy = "levenshtein";
    std::vector<std::string> exactCols;
    std::string output = "duplicates_report.json";
};

Config parseArgs(int argc, char* argv[]) {
    Config cfg;
    for (int i = 1; i < argc; i++) {
        std::string arg = argv[i];
        if ((arg == "--input" || arg == "-i") && i + 1 < argc) {
            cfg.input = argv[++i];
        } else if ((arg == "--threshold" || arg == "-t") && i + 1 < argc) {
            cfg.threshold = std::stod(argv[++i]);
        } else if ((arg == "--strategy" || arg == "-s") && i + 1 < argc) {
            cfg.strategy = argv[++i];
        } else if (arg == "--exact-cols") {
            while (i + 1 < argc && argv[i + 1][0] != '-') {
                cfg.exactCols.push_back(argv[++i]);
            }
        } else if ((arg == "--output" || arg == "-o") && i + 1 < argc) {
            cfg.output = argv[++i];
        }
    }
    return cfg;
}

// --- Main ---

int main(int argc, char* argv[]) {
    Config cfg = parseArgs(argc, argv);

    std::vector<Record> records;
    std::vector<std::string> columns;

    if (!cfg.input.empty()) {
        std::ifstream test(cfg.input);
        if (!test.good()) {
            std::cerr << "Error: File '" << cfg.input << "' not found." << std::endl;
            return 1;
        }
        test.close();
        std::cout << "Loading data from: " << cfg.input << std::endl;
        records = loadCsv(cfg.input, columns);
    } else {
        std::cout << "No input file specified. Generating sample dataset..." << std::endl;
        std::string samplePath = "sample_data.csv";
        generateSampleCsv(samplePath);
        records = loadCsv(samplePath, columns);
    }

    std::cout << "Loaded " << records.size() << " records with " << columns.size() << " columns" << std::endl;

    auto columnRules = buildColumnRules(columns, cfg.strategy, cfg.exactCols);
    std::cout << "Using strategy: " << cfg.strategy << ", threshold: " << cfg.threshold << std::endl;

    std::cout << "\nFinding exact duplicates..." << std::endl;
    auto exactGroups = findExactDuplicates(records);

    std::cout << "Finding fuzzy duplicates..." << std::endl;
    auto fuzzyGroups = findFuzzyDuplicates(records, columnRules, cfg.threshold, cfg.strategy);

    printConsoleReport(records, exactGroups, fuzzyGroups);
    generateJsonReport(records, exactGroups, fuzzyGroups, cfg.output);

    return 0;
}
README.md
# Duplicate Record Finder - C++ Trial 1

Identifies duplicate and near-duplicate records in CSV files using exact matching and fuzzy string matching with configurable similarity thresholds.

## Dependencies

- rapidfuzz-cpp (latest via FetchContent)
- nlohmann/json 3.11.3
- fast-cpp-csv-parser (master)

## Build & Run

```bash
mkdir build && cd build
cmake ..
cmake --build .
./finder

# With input file
./finder --input data.csv

# With custom threshold and strategy
./finder --input data.csv --threshold 0.85 --strategy jaro-winkler

# With exact match columns
./finder --input data.csv --exact-cols email phone

# Custom output path
./finder --input data.csv --output results.json
```

## Strategies

- `levenshtein` - Levenshtein distance-based similarity (via rapidfuzz)
- `jaro-winkler` - Jaro-Winkler similarity
- `cosine` - Cosine similarity using character bigrams

## Output

- Console report with duplicate groups and similarity scores
- `duplicates_report.json` with detailed findings