← All tasks
cppclaude-code/cpp-t1 #7Not a task: already works

Data Profiling Tool (cpp, written by Claude Code)

envgap__claude-code__cpp-t1-7

Written by a coding agent; not on GitHubWritten 2026-02-27

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/cpp-t1 #7 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Data Profiling Tool

Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV or JSON data file path as a command-line argument
- Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings)
- For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th)
- For string columns: compute min/max/average length, most common values (top 10), and unique count
- For all columns: count total values, missing/null values, missing percentage, and unique value count
- Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates
- Compute a pairwise correlation matrix for all numeric columns
- Print a formatted summary report to console showing key statistics per column
- Save the full profiling report as a JSON file with --output flag (default: data_profile.json)
- If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it
- Handle files with inconsistent delimiters or encoding issues gracefully

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.22)
project(data_profiler VERSION 1.0.0 LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)

# ---------------------------------------------------------------------------
# Dependencies
# ---------------------------------------------------------------------------
# fast-cpp-csv-parser (latest) — header-only CSV reader by Ben Strasser.
# nlohmann/json v3.11.3        — header-only JSON library.
# ---------------------------------------------------------------------------

include(FetchContent)

FetchContent_Declare(
    fast_csv
    GIT_REPOSITORY https://github.com/ben-strasser/fast-cpp-csv-parser.git
    GIT_TAG        master
    GIT_SHALLOW    TRUE
)

FetchContent_Declare(
    nlohmann_json
    GIT_REPOSITORY https://github.com/nlohmann/json.git
    GIT_TAG        v3.11.3
    GIT_SHALLOW    TRUE
)

set(JSON_BuildTests   OFF CACHE BOOL "" FORCE)
set(JSON_Install      OFF CACHE BOOL "" FORCE)
set(JSON_MultipleHeaders OFF CACHE BOOL "" FORCE)

FetchContent_MakeAvailable(fast_csv nlohmann_json)

# ---------------------------------------------------------------------------
# Executable
# ---------------------------------------------------------------------------

add_executable(data_profiler main.cpp)

# fast-cpp-csv-parser is header-only with no CMake target; add its source dir
# as an include path directly.
target_include_directories(data_profiler PRIVATE ${fast_csv_SOURCE_DIR})
target_link_libraries(data_profiler PRIVATE nlohmann_json::nlohmann_json)

if(CMAKE_CXX_COMPILER_ID MATCHES "GNU|Clang")
    target_compile_options(data_profiler PRIVATE -Wall -Wextra -Wpedantic)
endif()
main.cpp
#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
#include <vector>
#include <map>
#include <set>
#include <algorithm>
#include <cmath>
#include <numeric>
#include <iomanip>
#include <random>
#include <limits>
#include <regex>
#include <chrono>

#include <csv.h>
#include <nlohmann/json.hpp>

using json = nlohmann::json;

// ---------------------------------------------------------------------------
// Utility helpers
// ---------------------------------------------------------------------------

static std::string trim(const std::string& s) {
    size_t start = s.find_first_not_of(" \t\r\n");
    if (start == std::string::npos) return "";
    size_t end = s.find_last_not_of(" \t\r\n");
    return s.substr(start, end - start + 1);
}

static std::string toLower(const std::string& s) {
    std::string r = s;
    std::transform(r.begin(), r.end(), r.begin(),
                   [](unsigned char c) { return std::tolower(c); });
    return r;
}

static const std::set<std::string> NULL_TOKENS = {
    "", "null", "none", "na", "n/a", "nan", "undefined", "missing", "."
};

static bool isNull(const std::string& val) {
    return NULL_TOKENS.count(toLower(trim(val))) > 0;
}

static bool tryParseDouble(const std::string& s, double& out) {
    std::string t = trim(s);
    if (t.empty()) return false;
    try {
        size_t pos = 0;
        out = std::stod(t, &pos);
        return pos == t.size();
    } catch (...) {
        return false;
    }
}

static bool isDatetime(const std::string& s) {
    static const std::regex date_re(
        R"(\d{4}[-/]\d{1,2}[-/]\d{1,2})"
        R"((\s+\d{1,2}:\d{2}(:\d{2})?)?)"
    );
    return std::regex_match(trim(s), date_re);
}

static bool isBoolean(const std::string& s) {
    static const std::set<std::string> bools = {
        "true", "false", "yes", "no", "1", "0", "y", "n"
    };
    return bools.count(toLower(trim(s))) > 0;
}

static double roundVal(double v, int places) {
    double f = std::pow(10.0, places);
    return std::round(v * f) / f;
}

static std::string fmtDbl(double v, int prec = 4) {
    if (std::isnan(v)) return "N/A";
    std::ostringstream oss;
    oss << std::fixed << std::setprecision(prec) << v;
    return oss.str();
}

// ---------------------------------------------------------------------------
// Column type enum
// ---------------------------------------------------------------------------

enum class ColType { INTEGER, FLOAT, NUMERIC_MIXED, CATEGORICAL, BOOLEAN, DATETIME, TEXT, EMPTY };

static std::string colTypeStr(ColType t) {
    switch (t) {
        case ColType::INTEGER:       return "integer";
        case ColType::FLOAT:         return "float";
        case ColType::NUMERIC_MIXED: return "numeric_mixed";
        case ColType::CATEGORICAL:   return "categorical";
        case ColType::BOOLEAN:       return "boolean";
        case ColType::DATETIME:      return "datetime";
        case ColType::TEXT:          return "text";
        case ColType::EMPTY:         return "empty";
    }
    return "unknown";
}

static bool isNumericType(ColType t) {
    return t == ColType::INTEGER || t == ColType::FLOAT || t == ColType::NUMERIC_MIXED;
}

// ---------------------------------------------------------------------------
// CSV reading via fast-cpp-csv-parser (generic fallback for variable columns)
// fast-cpp-csv-parser is column-count-fixed, so we use its io::LineReader
// for flexibility and parse fields ourselves.
// ---------------------------------------------------------------------------

static std::vector<std::string> splitCSVLine(const std::string& line) {
    std::vector<std::string> fields;
    std::string field;
    bool in_quotes = false;
    for (size_t i = 0; i < line.size(); ++i) {
        char c = line[i];
        if (c == '"') {
            if (in_quotes && i + 1 < line.size() && line[i + 1] == '"') {
                field += '"';
                ++i;
            } else {
                in_quotes = !in_quotes;
            }
        } else if (c == ',' && !in_quotes) {
            fields.push_back(field);
            field.clear();
        } else {
            field += c;
        }
    }
    fields.push_back(field);
    return fields;
}

struct Dataset {
    std::vector<std::string> headers;
    std::map<std::string, std::vector<std::string>> columns;
    int row_count = 0;
};

static Dataset readCSV(const std::string& filename) {
    Dataset ds;

    // Use fast-csv-parser's LineReader for efficient line-by-line I/O
    io::LineReader reader(filename);

    // Header
    char* header_line = reader.next_line();
    if (!header_line) {
        throw std::runtime_error("File is empty: " + filename);
    }

    std::string hdr(header_line);
    // Strip UTF-8 BOM
    if (hdr.size() >= 3 &&
        static_cast<unsigned char>(hdr[0]) == 0xEF &&
        static_cast<unsigned char>(hdr[1]) == 0xBB &&
        static_cast<unsigned char>(hdr[2]) == 0xBF) {
        hdr = hdr.substr(3);
    }

    ds.headers = splitCSVLine(hdr);
    for (auto& h : ds.headers) {
        h = trim(h);
        ds.columns[h] = {};
    }

    // Data rows
    while (char* line = reader.next_line()) {
        std::string row(line);
        if (trim(row).empty()) continue;
        auto fields = splitCSVLine(row);
        fields.resize(ds.headers.size(), "");
        for (size_t i = 0; i < ds.headers.size(); ++i) {
            ds.columns[ds.headers[i]].push_back(trim(fields[i]));
        }
        ds.row_count++;
    }

    return ds;
}

// ---------------------------------------------------------------------------
// Type detection
// ---------------------------------------------------------------------------

static ColType detectColumnType(const std::vector<std::string>& values) {
    std::vector<std::string> clean;
    for (const auto& v : values) {
        if (!isNull(v)) clean.push_back(trim(v));
    }
    if (clean.empty()) return ColType::EMPTY;

    // Boolean check
    bool all_bool = true;
    for (const auto& v : clean) {
        if (!isBoolean(v)) { all_bool = false; break; }
    }
    if (all_bool) return ColType::BOOLEAN;

    // Datetime check
    int dt_count = 0;
    for (const auto& v : clean) {
        if (isDatetime(v)) dt_count++;
    }
    if (static_cast<double>(dt_count) / clean.size() > 0.8) return ColType::DATETIME;

    // Numeric check
    int num_count = 0;
    bool all_ints = true;
    for (const auto& v : clean) {
        double d;
        if (tryParseDouble(v, d)) {
            num_count++;
            if (v.find('.') != std::string::npos || d != std::floor(d)) {
                all_ints = false;
            }
        } else {
            all_ints = false;
        }
    }
    double num_ratio = static_cast<double>(num_count) / clean.size();
    if (num_ratio > 0.8) {
        if (num_count == static_cast<int>(clean.size()) && all_ints) return ColType::INTEGER;
        if (num_count == static_cast<int>(clean.size())) return ColType::FLOAT;
        return ColType::NUMERIC_MIXED;
    }

    // Categorical vs text
    std::set<std::string> unique(clean.begin(), clean.end());
    if (static_cast<double>(unique.size()) / clean.size() < 0.5 || unique.size() <= 20) {
        return ColType::CATEGORICAL;
    }
    return ColType::TEXT;
}

// ---------------------------------------------------------------------------
// Percentile (linear interpolation)
// ---------------------------------------------------------------------------

static double percentile(const std::vector<double>& sorted, double p) {
    if (sorted.empty()) return std::numeric_limits<double>::quiet_NaN();
    if (sorted.size() == 1) return sorted[0];
    double idx = (p / 100.0) * static_cast<double>(sorted.size() - 1);
    size_t lo = static_cast<size_t>(std::floor(idx));
    size_t hi = lo + 1;
    if (hi >= sorted.size()) return sorted.back();
    double frac = idx - static_cast<double>(lo);
    return sorted[lo] * (1.0 - frac) + sorted[hi] * frac;
}

// ---------------------------------------------------------------------------
// Numeric statistics
// ---------------------------------------------------------------------------

static json computeNumericStats(const std::vector<std::string>& values) {
    std::vector<double> nums;
    for (const auto& v : values) {
        double d;
        if (!isNull(v) && tryParseDouble(v, d)) nums.push_back(d);
    }

    json stats;
    stats["count"] = static_cast<int>(nums.size());
    if (nums.empty()) return stats;

    std::vector<double> sorted = nums;
    std::sort(sorted.begin(), sorted.end());

    double sum = std::accumulate(nums.begin(), nums.end(), 0.0);
    double mean = sum / nums.size();

    double sq = 0.0;
    for (double v : nums) { double d = v - mean; sq += d * d; }
    double variance = (nums.size() > 1) ? sq / static_cast<double>(nums.size() - 1) : 0.0;
    double std_dev = std::sqrt(variance);

    double q1 = percentile(sorted, 25.0);
    double q3 = percentile(sorted, 75.0);
    double iqr = q3 - q1;

    // Skewness and kurtosis
    double skewness = 0.0, kurtosis = 0.0;
    if (nums.size() > 2 && std_dev > 0) {
        double n = static_cast<double>(nums.size());
        double m3 = 0.0, m4 = 0.0;
        for (double v : nums) {
            double d = (v - mean) / std_dev;
            m3 += d * d * d;
            m4 += d * d * d * d;
        }
        skewness = (n / ((n - 1) * (n - 2))) * m3;
        kurtosis = ((n * (n + 1)) / ((n - 1) * (n - 2) * (n - 3))) * m4
                   - (3.0 * (n - 1) * (n - 1)) / ((n - 2) * (n - 3));
    }

    stats["mean"]     = roundVal(mean, 4);
    stats["median"]   = roundVal(percentile(sorted, 50.0), 4);
    stats["std"]      = roundVal(std_dev, 4);
    stats["min"]      = roundVal(sorted.front(), 4);
    stats["max"]      = roundVal(sorted.back(), 4);
    stats["q1"]       = roundVal(q1, 4);
    stats["q3"]       = roundVal(q3, 4);
    stats["skewness"] = roundVal(skewness, 4);
    stats["kurtosis"] = roundVal(kurtosis, 4);

    // Outliers (IQR method)
    double lower = q1 - 1.5 * iqr;
    double upper = q3 + 1.5 * iqr;
    int outlier_count = 0;
    int zeros = 0, negatives = 0;
    for (double v : nums) {
        if (v < lower || v > upper) outlier_count++;
        if (v == 0.0) zeros++;
        if (v < 0.0) negatives++;
    }
    stats["outlier_count"] = outlier_count;
    stats["zeros"]         = zeros;
    stats["negatives"]     = negatives;

    return stats;
}

// ---------------------------------------------------------------------------
// Categorical statistics
// ---------------------------------------------------------------------------

static json computeCategoricalStats(const std::vector<std::string>& values) {
    std::map<std::string, int> freq;
    int valid_count = 0;
    for (const auto& v : values) {
        if (!isNull(v)) {
            freq[trim(v)]++;
            valid_count++;
        }
    }

    json stats;
    stats["valid_count"]  = valid_count;
    stats["unique_count"] = static_cast<int>(freq.size());

    // Sort by frequency descending
    std::vector<std::pair<std::string, int>> sorted(freq.begin(), freq.end());
    std::sort(sorted.begin(), sorted.end(),
              [](const auto& a, const auto& b) { return a.second > b.second; });

    if (!sorted.empty()) {
        stats["mode"]           = sorted[0].first;
        stats["mode_frequency"] = sorted[0].second;
    }

    json top = json::object();
    int limit = std::min(static_cast<int>(sorted.size()), 10);
    for (int i = 0; i < limit; ++i) {
        top[sorted[i].first] = sorted[i].second;
    }
    stats["top_values"] = top;

    return stats;
}

// ---------------------------------------------------------------------------
// Distribution computation
// ---------------------------------------------------------------------------

static json computeDistribution(const std::vector<std::string>& values, ColType type) {
    json dist;

    if (isNumericType(type)) {
        std::vector<double> nums;
        for (const auto& v : values) {
            double d;
            if (!isNull(v) && tryParseDouble(v, d)) nums.push_back(d);
        }
        if (nums.empty()) return dist;

        std::sort(nums.begin(), nums.end());
        double mn = nums.front(), mx = nums.back();
        int bin_count = 10;
        double bin_width = (mx - mn) / bin_count;
        if (bin_width == 0) bin_width = 1;

        json bins = json::array();
        std::vector<int> counts(bin_count, 0);
        for (int i = 0; i <= bin_count; ++i) {
            bins.push_back(roundVal(mn + i * bin_width, 4));
        }
        for (double n : nums) {
            int idx = static_cast<int>(std::floor((n - mn) / bin_width));
            if (idx >= bin_count) idx = bin_count - 1;
            if (idx < 0) idx = 0;
            counts[idx]++;
        }

        dist["type"]   = "histogram";
        dist["bins"]   = bins;
        dist["counts"] = counts;
    } else {
        std::map<std::string, int> freq;
        for (const auto& v : values) {
            if (!isNull(v)) freq[trim(v)]++;
        }
        std::vector<std::pair<std::string, int>> sorted(freq.begin(), freq.end());
        std::sort(sorted.begin(), sorted.end(),
                  [](const auto& a, const auto& b) { return a.second > b.second; });

        json vals = json::object();
        int limit = std::min(static_cast<int>(sorted.size()), 20);
        for (int i = 0; i < limit; ++i) {
            vals[sorted[i].first] = sorted[i].second;
        }
        dist["type"]   = "frequency";
        dist["values"] = vals;
    }

    return dist;
}

// ---------------------------------------------------------------------------
// Correlation matrix
// ---------------------------------------------------------------------------

static json computeCorrelationMatrix(const Dataset& ds,
                                      const std::map<std::string, ColType>& types) {
    std::vector<std::string> num_cols;
    for (const auto& h : ds.headers) {
        auto it = types.find(h);
        if (it != types.end() && isNumericType(it->second)) {
            num_cols.push_back(h);
        }
    }

    json matrix = json::object();
    if (num_cols.size() < 2) return matrix;

    // Pre-parse numeric columns
    std::map<std::string, std::vector<double>> parsed;
    for (const auto& col : num_cols) {
        auto& vals = ds.columns.at(col);
        auto& pv = parsed[col];
        pv.resize(ds.row_count, std::numeric_limits<double>::quiet_NaN());
        for (int i = 0; i < ds.row_count; ++i) {
            double d;
            if (!isNull(vals[i]) && tryParseDouble(vals[i], d)) {
                pv[i] = d;
            }
        }
    }

    for (const auto& colA : num_cols) {
        matrix[colA] = json::object();
        for (const auto& colB : num_cols) {
            // Collect paired non-NaN values
            std::vector<double> xs, ys;
            for (int i = 0; i < ds.row_count; ++i) {
                double a = parsed[colA][i];
                double b = parsed[colB][i];
                if (!std::isnan(a) && !std::isnan(b)) {
                    xs.push_back(a);
                    ys.push_back(b);
                }
            }

            if (xs.size() < 2) {
                matrix[colA][colB] = nullptr;
                continue;
            }

            double n = static_cast<double>(xs.size());
            double sx = std::accumulate(xs.begin(), xs.end(), 0.0);
            double sy = std::accumulate(ys.begin(), ys.end(), 0.0);
            double sxx = 0, syy = 0, sxy = 0;
            for (size_t i = 0; i < xs.size(); ++i) {
                sxx += xs[i] * xs[i];
                syy += ys[i] * ys[i];
                sxy += xs[i] * ys[i];
            }
            double denom = std::sqrt((n * sxx - sx * sx) * (n * syy - sy * sy));
            if (denom == 0.0) {
                matrix[colA][colB] = nullptr;
            } else {
                double r = (n * sxy - sx * sy) / denom;
                matrix[colA][colB] = roundVal(r, 4);
            }
        }
    }

    return matrix;
}

// ---------------------------------------------------------------------------
// Data quality score
// ---------------------------------------------------------------------------

static double computeDataQualityScore(const Dataset& ds,
                                       const std::map<std::string, ColType>& types) {
    int total_cells = ds.row_count * static_cast<int>(ds.headers.size());
    if (total_cells == 0) return 0.0;

    int null_count = 0;
    for (const auto& h : ds.headers) {
        for (const auto& v : ds.columns.at(h)) {
            if (isNull(v)) null_count++;
        }
    }
    double completeness = 1.0 - static_cast<double>(null_count) / total_cells;

    double uniqueness_sum = 0.0;
    int uniqueness_count = 0;
    for (const auto& h : ds.headers) {
        std::vector<std::string> clean;
        for (const auto& v : ds.columns.at(h)) {
            if (!isNull(v)) clean.push_back(trim(v));
        }
        if (!clean.empty()) {
            std::set<std::string> unique(clean.begin(), clean.end());
            uniqueness_sum += static_cast<double>(unique.size()) / clean.size();
            uniqueness_count++;
        }
    }
    double avg_uniqueness = uniqueness_count > 0 ? uniqueness_sum / uniqueness_count : 0.0;

    double consistency_sum = 0.0;
    int consistency_count = 0;
    for (const auto& h : ds.headers) {
        std::vector<std::string> clean;
        for (const auto& v : ds.columns.at(h)) {
            if (!isNull(v)) clean.push_back(trim(v));
        }
        if (!clean.empty()) {
            int num = 0;
            for (const auto& v : clean) {
                double d;
                if (tryParseDouble(v, d)) num++;
            }
            double ratio = static_cast<double>(num) / clean.size();
            consistency_sum += std::max(ratio, 1.0 - ratio);
            consistency_count++;
        }
    }
    double avg_consistency = consistency_count > 0 ? consistency_sum / consistency_count : 0.0;

    return roundVal((completeness * 0.5 + avg_consistency * 0.3 +
                     std::min(avg_uniqueness, 1.0) * 0.2) * 100, 2);
}

// ---------------------------------------------------------------------------
// Sample dataset generator
// ---------------------------------------------------------------------------

static void generateSampleDataset(const std::string& filename) {
    std::mt19937 rng(42);

    const std::vector<std::string> names = {
        "Alice", "Bob", "Charlie", "Diana", "Eve", "", "Frank", "Grace", "", "Heidi"
    };
    const std::vector<std::string> depts = {
        "Engineering", "Sales", "HR", "Marketing", "", "Engineering", "Sales", ""
    };
    const std::vector<std::string> dates = {
        "2020-01-15", "2021-06-30", "2019-12-01", "", "invalid-date", "2022-03-10"
    };
    const std::vector<std::string> actives = {"true", "false", "", "yes", "no", "1", "0"};

    std::ofstream out(filename);
    if (!out) {
        std::cerr << "[ERROR] Cannot create " << filename << "\n";
        return;
    }

    out << "id,name,age,salary,department,score,join_date,is_active\n";

    std::uniform_int_distribution<int> age_d(18, 80);
    std::uniform_real_distribution<double> salary_d(25000.0, 150000.0);
    std::uniform_real_distribution<double> score_d(0.0, 100.0);
    std::uniform_int_distribution<int> miss_d(0, 99);

    for (int i = 0; i < 200; ++i) {
        out << (i + 1) << ",";

        // name
        out << names[rng() % names.size()] << ",";

        // age
        if (miss_d(rng) < 8) out << ",";
        else if (i == 50)    out << "999,";
        else if (i == 100)   out << "-1,";
        else                 out << age_d(rng) << ",";

        // salary
        if (miss_d(rng) < 8) out << ",";
        else {
            std::ostringstream oss;
            oss << std::fixed << std::setprecision(2);
            if (i == 75) oss << -500.0;
            else         oss << salary_d(rng);
            out << oss.str() << ",";
        }

        // department
        out << depts[rng() % depts.size()] << ",";

        // score
        if (miss_d(rng) < 10) out << ",";
        else {
            std::ostringstream oss;
            oss << std::fixed << std::setprecision(1) << score_d(rng);
            out << oss.str() << ",";
        }

        // join_date
        out << dates[rng() % dates.size()] << ",";

        // is_active
        out << actives[rng() % actives.size()] << "\n";
    }

    std::cout << "[INFO] Generated sample dataset: " << filename << " (200 rows)\n";
}

// ---------------------------------------------------------------------------
// Console report
// ---------------------------------------------------------------------------

static void printReport(const json& report) {
    std::cout << "\n" << std::string(60, '=') << "\n"
              << "  DATA PROFILING TOOL - Trial 1 (fast-cpp-csv-parser + nlohmann/json)\n"
              << std::string(60, '=') << "\n\n";

    auto& overview = report["dataset_overview"];
    std::cout << "--- Dataset Overview ---\n"
              << "  Rows:            " << overview["rows"] << "\n"
              << "  Columns:         " << overview["columns"] << "\n"
              << "  Total cells:     " << overview["total_cells"] << "\n"
              << "  Total missing:   " << overview["total_missing"]
              << " (" << overview["total_missing_pct"] << "%)\n"
              << "\n  Data Quality Score: " << report["data_quality_score"] << " / 100\n\n";

    std::cout << "--- Column Profiles ---\n";
    std::cout << std::left
              << std::setw(16) << "  Column"
              << std::setw(16) << "Type"
              << std::setw(10) << "Missing"
              << std::setw(10) << "Miss%"
              << std::setw(10) << "Unique"
              << "\n  " << std::string(60, '-') << "\n";

    for (auto& [col_name, profile] : report["column_profiles"].items()) {
        std::cout << "  " << std::left
                  << std::setw(14) << col_name.substr(0, 13)
                  << std::setw(16) << profile["detected_type"].get<std::string>()
                  << std::setw(10) << profile["missing_count"]
                  << std::setw(10) << (std::to_string(profile["missing_percentage"].get<double>()) + "%").substr(0, 8)
                  << std::setw(10) << profile["unique_count"]
                  << "\n";
    }

    // Detailed per-column
    for (auto& [col_name, profile] : report["column_profiles"].items()) {
        std::cout << "\n  [" << col_name << "]\n"
                  << "    Type:          " << profile["detected_type"] << "\n"
                  << "    Missing:       " << profile["missing_count"]
                  << " (" << profile["missing_percentage"] << "%)\n"
                  << "    Unique:        " << profile["unique_count"] << "\n";

        if (profile.contains("numeric_stats") && !profile["numeric_stats"].empty()) {
            auto& ns = profile["numeric_stats"];
            for (auto& [k, v] : ns.items()) {
                std::cout << "    " << std::left << std::setw(15) << (k + ":")
                          << " " << v << "\n";
            }
        }

        if (profile.contains("categorical_stats") && !profile["categorical_stats"].empty()) {
            auto& cs = profile["categorical_stats"];
            if (cs.contains("mode"))
                std::cout << "    Mode:          " << cs["mode"]
                          << " (freq: " << cs["mode_frequency"] << ")\n";
            if (cs.contains("top_values")) {
                std::cout << "    Top values:\n";
                int cnt = 0;
                for (auto& [val, freq] : cs["top_values"].items()) {
                    if (cnt++ >= 5) break;
                    std::cout << "      " << val << ": " << freq << "\n";
                }
            }
        }
    }

    // Correlation matrix
    if (report.contains("correlation_matrix") && !report["correlation_matrix"].empty()) {
        std::cout << "\n--- Correlation Matrix ---\n";
        auto& corr = report["correlation_matrix"];
        std::vector<std::string> cols;
        for (auto& [k, _] : corr.items()) cols.push_back(k);

        std::cout << "  " << std::left << std::setw(15) << "";
        for (const auto& c : cols) std::cout << std::setw(14) << c.substr(0, 12);
        std::cout << "\n";

        for (const auto& r : cols) {
            std::cout << "  " << std::left << std::setw(15) << r.substr(0, 14);
            for (const auto& c : cols) {
                if (corr[r][c].is_null())
                    std::cout << std::setw(14) << "N/A";
                else
                    std::cout << std::setw(14) << fmtDbl(corr[r][c].get<double>());
            }
            std::cout << "\n";
        }
    }
}

// ---------------------------------------------------------------------------
// main
// ---------------------------------------------------------------------------

int main(int argc, char* argv[]) {
    std::string filepath;

    if (argc < 2) {
        std::cout << "[INFO] No input file provided. Generating sample dataset...\n";
        filepath = "sample_data.csv";
        generateSampleDataset(filepath);
    } else {
        filepath = argv[1];
    }

    std::ifstream check(filepath);
    if (!check.good()) {
        std::cerr << "[ERROR] File not found: " << filepath << "\n";
        return 1;
    }
    check.close();

    try {
        Dataset ds = readCSV(filepath);

        if (ds.row_count == 0) {
            std::cout << "[WARN] File contains only a header row.\n";
            return 0;
        }

        std::cout << "[INFO] Loaded file: " << filepath << "\n"
                  << "[INFO] Shape: " << ds.row_count << " rows x "
                  << ds.headers.size() << " columns\n";

        // Detect column types
        std::map<std::string, ColType> col_types;
        for (const auto& h : ds.headers) {
            col_types[h] = detectColumnType(ds.columns[h]);
        }

        // Build column profiles
        json column_profiles = json::object();
        for (const auto& h : ds.headers) {
            json profile;
            profile["name"]          = h;
            profile["detected_type"] = colTypeStr(col_types[h]);
            profile["total_count"]   = static_cast<int>(ds.columns[h].size());

            int missing = 0;
            for (const auto& v : ds.columns[h]) {
                if (isNull(v)) missing++;
            }
            profile["missing_count"]      = missing;
            profile["missing_percentage"] = roundVal(
                ds.row_count > 0 ? static_cast<double>(missing) / ds.row_count * 100 : 0, 2);

            std::set<std::string> unique;
            for (const auto& v : ds.columns[h]) {
                if (!isNull(v)) unique.insert(trim(v));
            }
            profile["unique_count"] = static_cast<int>(unique.size());

            if (isNumericType(col_types[h])) {
                profile["numeric_stats"] = computeNumericStats(ds.columns[h]);
            } else {
                profile["categorical_stats"] = computeCategoricalStats(ds.columns[h]);
            }

            profile["distribution"] = computeDistribution(ds.columns[h], col_types[h]);
            column_profiles[h] = profile;
        }

        // Correlation matrix
        json corr_matrix = computeCorrelationMatrix(ds, col_types);

        // Quality score
        double quality = computeDataQualityScore(ds, col_types);

        // Total missing
        int total_missing = 0;
        int total_cells = ds.row_count * static_cast<int>(ds.headers.size());
        for (const auto& h : ds.headers) {
            for (const auto& v : ds.columns[h]) {
                if (isNull(v)) total_missing++;
            }
        }

        // Build final report
        json report;
        report["file"]         = filepath;
        report["generated_at"] = []() {
            auto now = std::chrono::system_clock::now();
            auto t = std::chrono::system_clock::to_time_t(now);
            std::ostringstream oss;
            oss << std::put_time(std::localtime(&t), "%Y-%m-%dT%H:%M:%S");
            return oss.str();
        }();

        json overview;
        overview["rows"]              = ds.row_count;
        overview["columns"]           = static_cast<int>(ds.headers.size());
        overview["total_cells"]       = total_cells;
        overview["total_missing"]     = total_missing;
        overview["total_missing_pct"] = roundVal(
            total_cells > 0 ? static_cast<double>(total_missing) / total_cells * 100 : 0, 2);

        report["dataset_overview"]   = overview;
        report["data_quality_score"] = quality;
        report["column_profiles"]    = column_profiles;
        report["correlation_matrix"] = corr_matrix;

        // Console output
        printReport(report);

        // Save JSON
        std::ofstream out("profile_report.json");
        if (out) {
            out << report.dump(2) << "\n";
            std::cout << "\n[INFO] Profile report saved: profile_report.json\n";
        } else {
            std::cerr << "[WARN] Could not write profile_report.json\n";
        }

    } catch (const std::exception& e) {
        std::cerr << "[ERROR] " << e.what() << "\n";
        return 1;
    }

    return 0;
}
README.md
# Data Profiling Tool - C++ Trial 1

Profiles tabular datasets: column types, distributions, missing values,
correlations, and quality issues. Reads CSV input, auto-detects column types
(numeric, categorical, datetime, boolean), computes per-column statistics,
detects missing values, outliers, and unique counts. Outputs profiling report
as JSON.

## Dependencies

| Library | Version | Purpose |
|---------|---------|---------|
| [fast-cpp-csv-parser](https://github.com/ben-strasser/fast-cpp-csv-parser) | latest | High-performance CSV line reading |
| [nlohmann/json](https://github.com/nlohmann/json) | 3.11.3 | JSON report serialisation |

Both libraries are header-only. CMake's `FetchContent` downloads them
automatically on the first configure step.

## Build Instructions

```bash
# Configure (downloads dependencies on first run)
cmake -B build -DCMAKE_BUILD_TYPE=Release

# Compile
cmake --build build --parallel

# Binary is at: build/data_profiler
```

### System requirements

| Tool | Minimum version |
|------|----------------|
| C++ compiler (g++ / clang++) | C++17 support |
| CMake | 3.22 |
| git | any (for FetchContent) |

## Usage

```bash
# Profile your own CSV file
./build/data_profiler path/to/data.csv

# Generate and profile a sample dataset
./build/data_profiler
```

## Output

- Console report with column profiles, distributions, and correlations
- `profile_report.json` with full profiling results