Data Profiling Tool (cpp, written by Claude Code)
envgap__claude-code__cpp-t1-7
Written by a coding agent; not on GitHubWritten 2026-02-27
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
claude-code/cpp-t1 #7 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Data Profiling Tool Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report. FUNCTIONAL REQUIREMENTS: - Accept a CSV or JSON data file path as a command-line argument - Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings) - For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th) - For string columns: compute min/max/average length, most common values (top 10), and unique count - For all columns: count total values, missing/null values, missing percentage, and unique value count - Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates - Compute a pairwise correlation matrix for all numeric columns - Print a formatted summary report to console showing key statistics per column - Save the full profiling report as a JSON file with --output flag (default: data_profile.json) - If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it - Handle files with inconsistent delimiters or encoding issues gracefully Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.22)
project(data_profiler VERSION 1.0.0 LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)
# ---------------------------------------------------------------------------
# Dependencies
# ---------------------------------------------------------------------------
# fast-cpp-csv-parser (latest) — header-only CSV reader by Ben Strasser.
# nlohmann/json v3.11.3 — header-only JSON library.
# ---------------------------------------------------------------------------
include(FetchContent)
FetchContent_Declare(
fast_csv
GIT_REPOSITORY https://github.com/ben-strasser/fast-cpp-csv-parser.git
GIT_TAG master
GIT_SHALLOW TRUE
)
FetchContent_Declare(
nlohmann_json
GIT_REPOSITORY https://github.com/nlohmann/json.git
GIT_TAG v3.11.3
GIT_SHALLOW TRUE
)
set(JSON_BuildTests OFF CACHE BOOL "" FORCE)
set(JSON_Install OFF CACHE BOOL "" FORCE)
set(JSON_MultipleHeaders OFF CACHE BOOL "" FORCE)
FetchContent_MakeAvailable(fast_csv nlohmann_json)
# ---------------------------------------------------------------------------
# Executable
# ---------------------------------------------------------------------------
add_executable(data_profiler main.cpp)
# fast-cpp-csv-parser is header-only with no CMake target; add its source dir
# as an include path directly.
target_include_directories(data_profiler PRIVATE ${fast_csv_SOURCE_DIR})
target_link_libraries(data_profiler PRIVATE nlohmann_json::nlohmann_json)
if(CMAKE_CXX_COMPILER_ID MATCHES "GNU|Clang")
target_compile_options(data_profiler PRIVATE -Wall -Wextra -Wpedantic)
endif()
main.cpp
#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
#include <vector>
#include <map>
#include <set>
#include <algorithm>
#include <cmath>
#include <numeric>
#include <iomanip>
#include <random>
#include <limits>
#include <regex>
#include <chrono>
#include <csv.h>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
// ---------------------------------------------------------------------------
// Utility helpers
// ---------------------------------------------------------------------------
static std::string trim(const std::string& s) {
size_t start = s.find_first_not_of(" \t\r\n");
if (start == std::string::npos) return "";
size_t end = s.find_last_not_of(" \t\r\n");
return s.substr(start, end - start + 1);
}
static std::string toLower(const std::string& s) {
std::string r = s;
std::transform(r.begin(), r.end(), r.begin(),
[](unsigned char c) { return std::tolower(c); });
return r;
}
static const std::set<std::string> NULL_TOKENS = {
"", "null", "none", "na", "n/a", "nan", "undefined", "missing", "."
};
static bool isNull(const std::string& val) {
return NULL_TOKENS.count(toLower(trim(val))) > 0;
}
static bool tryParseDouble(const std::string& s, double& out) {
std::string t = trim(s);
if (t.empty()) return false;
try {
size_t pos = 0;
out = std::stod(t, &pos);
return pos == t.size();
} catch (...) {
return false;
}
}
static bool isDatetime(const std::string& s) {
static const std::regex date_re(
R"(\d{4}[-/]\d{1,2}[-/]\d{1,2})"
R"((\s+\d{1,2}:\d{2}(:\d{2})?)?)"
);
return std::regex_match(trim(s), date_re);
}
static bool isBoolean(const std::string& s) {
static const std::set<std::string> bools = {
"true", "false", "yes", "no", "1", "0", "y", "n"
};
return bools.count(toLower(trim(s))) > 0;
}
static double roundVal(double v, int places) {
double f = std::pow(10.0, places);
return std::round(v * f) / f;
}
static std::string fmtDbl(double v, int prec = 4) {
if (std::isnan(v)) return "N/A";
std::ostringstream oss;
oss << std::fixed << std::setprecision(prec) << v;
return oss.str();
}
// ---------------------------------------------------------------------------
// Column type enum
// ---------------------------------------------------------------------------
enum class ColType { INTEGER, FLOAT, NUMERIC_MIXED, CATEGORICAL, BOOLEAN, DATETIME, TEXT, EMPTY };
static std::string colTypeStr(ColType t) {
switch (t) {
case ColType::INTEGER: return "integer";
case ColType::FLOAT: return "float";
case ColType::NUMERIC_MIXED: return "numeric_mixed";
case ColType::CATEGORICAL: return "categorical";
case ColType::BOOLEAN: return "boolean";
case ColType::DATETIME: return "datetime";
case ColType::TEXT: return "text";
case ColType::EMPTY: return "empty";
}
return "unknown";
}
static bool isNumericType(ColType t) {
return t == ColType::INTEGER || t == ColType::FLOAT || t == ColType::NUMERIC_MIXED;
}
// ---------------------------------------------------------------------------
// CSV reading via fast-cpp-csv-parser (generic fallback for variable columns)
// fast-cpp-csv-parser is column-count-fixed, so we use its io::LineReader
// for flexibility and parse fields ourselves.
// ---------------------------------------------------------------------------
static std::vector<std::string> splitCSVLine(const std::string& line) {
std::vector<std::string> fields;
std::string field;
bool in_quotes = false;
for (size_t i = 0; i < line.size(); ++i) {
char c = line[i];
if (c == '"') {
if (in_quotes && i + 1 < line.size() && line[i + 1] == '"') {
field += '"';
++i;
} else {
in_quotes = !in_quotes;
}
} else if (c == ',' && !in_quotes) {
fields.push_back(field);
field.clear();
} else {
field += c;
}
}
fields.push_back(field);
return fields;
}
struct Dataset {
std::vector<std::string> headers;
std::map<std::string, std::vector<std::string>> columns;
int row_count = 0;
};
static Dataset readCSV(const std::string& filename) {
Dataset ds;
// Use fast-csv-parser's LineReader for efficient line-by-line I/O
io::LineReader reader(filename);
// Header
char* header_line = reader.next_line();
if (!header_line) {
throw std::runtime_error("File is empty: " + filename);
}
std::string hdr(header_line);
// Strip UTF-8 BOM
if (hdr.size() >= 3 &&
static_cast<unsigned char>(hdr[0]) == 0xEF &&
static_cast<unsigned char>(hdr[1]) == 0xBB &&
static_cast<unsigned char>(hdr[2]) == 0xBF) {
hdr = hdr.substr(3);
}
ds.headers = splitCSVLine(hdr);
for (auto& h : ds.headers) {
h = trim(h);
ds.columns[h] = {};
}
// Data rows
while (char* line = reader.next_line()) {
std::string row(line);
if (trim(row).empty()) continue;
auto fields = splitCSVLine(row);
fields.resize(ds.headers.size(), "");
for (size_t i = 0; i < ds.headers.size(); ++i) {
ds.columns[ds.headers[i]].push_back(trim(fields[i]));
}
ds.row_count++;
}
return ds;
}
// ---------------------------------------------------------------------------
// Type detection
// ---------------------------------------------------------------------------
static ColType detectColumnType(const std::vector<std::string>& values) {
std::vector<std::string> clean;
for (const auto& v : values) {
if (!isNull(v)) clean.push_back(trim(v));
}
if (clean.empty()) return ColType::EMPTY;
// Boolean check
bool all_bool = true;
for (const auto& v : clean) {
if (!isBoolean(v)) { all_bool = false; break; }
}
if (all_bool) return ColType::BOOLEAN;
// Datetime check
int dt_count = 0;
for (const auto& v : clean) {
if (isDatetime(v)) dt_count++;
}
if (static_cast<double>(dt_count) / clean.size() > 0.8) return ColType::DATETIME;
// Numeric check
int num_count = 0;
bool all_ints = true;
for (const auto& v : clean) {
double d;
if (tryParseDouble(v, d)) {
num_count++;
if (v.find('.') != std::string::npos || d != std::floor(d)) {
all_ints = false;
}
} else {
all_ints = false;
}
}
double num_ratio = static_cast<double>(num_count) / clean.size();
if (num_ratio > 0.8) {
if (num_count == static_cast<int>(clean.size()) && all_ints) return ColType::INTEGER;
if (num_count == static_cast<int>(clean.size())) return ColType::FLOAT;
return ColType::NUMERIC_MIXED;
}
// Categorical vs text
std::set<std::string> unique(clean.begin(), clean.end());
if (static_cast<double>(unique.size()) / clean.size() < 0.5 || unique.size() <= 20) {
return ColType::CATEGORICAL;
}
return ColType::TEXT;
}
// ---------------------------------------------------------------------------
// Percentile (linear interpolation)
// ---------------------------------------------------------------------------
static double percentile(const std::vector<double>& sorted, double p) {
if (sorted.empty()) return std::numeric_limits<double>::quiet_NaN();
if (sorted.size() == 1) return sorted[0];
double idx = (p / 100.0) * static_cast<double>(sorted.size() - 1);
size_t lo = static_cast<size_t>(std::floor(idx));
size_t hi = lo + 1;
if (hi >= sorted.size()) return sorted.back();
double frac = idx - static_cast<double>(lo);
return sorted[lo] * (1.0 - frac) + sorted[hi] * frac;
}
// ---------------------------------------------------------------------------
// Numeric statistics
// ---------------------------------------------------------------------------
static json computeNumericStats(const std::vector<std::string>& values) {
std::vector<double> nums;
for (const auto& v : values) {
double d;
if (!isNull(v) && tryParseDouble(v, d)) nums.push_back(d);
}
json stats;
stats["count"] = static_cast<int>(nums.size());
if (nums.empty()) return stats;
std::vector<double> sorted = nums;
std::sort(sorted.begin(), sorted.end());
double sum = std::accumulate(nums.begin(), nums.end(), 0.0);
double mean = sum / nums.size();
double sq = 0.0;
for (double v : nums) { double d = v - mean; sq += d * d; }
double variance = (nums.size() > 1) ? sq / static_cast<double>(nums.size() - 1) : 0.0;
double std_dev = std::sqrt(variance);
double q1 = percentile(sorted, 25.0);
double q3 = percentile(sorted, 75.0);
double iqr = q3 - q1;
// Skewness and kurtosis
double skewness = 0.0, kurtosis = 0.0;
if (nums.size() > 2 && std_dev > 0) {
double n = static_cast<double>(nums.size());
double m3 = 0.0, m4 = 0.0;
for (double v : nums) {
double d = (v - mean) / std_dev;
m3 += d * d * d;
m4 += d * d * d * d;
}
skewness = (n / ((n - 1) * (n - 2))) * m3;
kurtosis = ((n * (n + 1)) / ((n - 1) * (n - 2) * (n - 3))) * m4
- (3.0 * (n - 1) * (n - 1)) / ((n - 2) * (n - 3));
}
stats["mean"] = roundVal(mean, 4);
stats["median"] = roundVal(percentile(sorted, 50.0), 4);
stats["std"] = roundVal(std_dev, 4);
stats["min"] = roundVal(sorted.front(), 4);
stats["max"] = roundVal(sorted.back(), 4);
stats["q1"] = roundVal(q1, 4);
stats["q3"] = roundVal(q3, 4);
stats["skewness"] = roundVal(skewness, 4);
stats["kurtosis"] = roundVal(kurtosis, 4);
// Outliers (IQR method)
double lower = q1 - 1.5 * iqr;
double upper = q3 + 1.5 * iqr;
int outlier_count = 0;
int zeros = 0, negatives = 0;
for (double v : nums) {
if (v < lower || v > upper) outlier_count++;
if (v == 0.0) zeros++;
if (v < 0.0) negatives++;
}
stats["outlier_count"] = outlier_count;
stats["zeros"] = zeros;
stats["negatives"] = negatives;
return stats;
}
// ---------------------------------------------------------------------------
// Categorical statistics
// ---------------------------------------------------------------------------
static json computeCategoricalStats(const std::vector<std::string>& values) {
std::map<std::string, int> freq;
int valid_count = 0;
for (const auto& v : values) {
if (!isNull(v)) {
freq[trim(v)]++;
valid_count++;
}
}
json stats;
stats["valid_count"] = valid_count;
stats["unique_count"] = static_cast<int>(freq.size());
// Sort by frequency descending
std::vector<std::pair<std::string, int>> sorted(freq.begin(), freq.end());
std::sort(sorted.begin(), sorted.end(),
[](const auto& a, const auto& b) { return a.second > b.second; });
if (!sorted.empty()) {
stats["mode"] = sorted[0].first;
stats["mode_frequency"] = sorted[0].second;
}
json top = json::object();
int limit = std::min(static_cast<int>(sorted.size()), 10);
for (int i = 0; i < limit; ++i) {
top[sorted[i].first] = sorted[i].second;
}
stats["top_values"] = top;
return stats;
}
// ---------------------------------------------------------------------------
// Distribution computation
// ---------------------------------------------------------------------------
static json computeDistribution(const std::vector<std::string>& values, ColType type) {
json dist;
if (isNumericType(type)) {
std::vector<double> nums;
for (const auto& v : values) {
double d;
if (!isNull(v) && tryParseDouble(v, d)) nums.push_back(d);
}
if (nums.empty()) return dist;
std::sort(nums.begin(), nums.end());
double mn = nums.front(), mx = nums.back();
int bin_count = 10;
double bin_width = (mx - mn) / bin_count;
if (bin_width == 0) bin_width = 1;
json bins = json::array();
std::vector<int> counts(bin_count, 0);
for (int i = 0; i <= bin_count; ++i) {
bins.push_back(roundVal(mn + i * bin_width, 4));
}
for (double n : nums) {
int idx = static_cast<int>(std::floor((n - mn) / bin_width));
if (idx >= bin_count) idx = bin_count - 1;
if (idx < 0) idx = 0;
counts[idx]++;
}
dist["type"] = "histogram";
dist["bins"] = bins;
dist["counts"] = counts;
} else {
std::map<std::string, int> freq;
for (const auto& v : values) {
if (!isNull(v)) freq[trim(v)]++;
}
std::vector<std::pair<std::string, int>> sorted(freq.begin(), freq.end());
std::sort(sorted.begin(), sorted.end(),
[](const auto& a, const auto& b) { return a.second > b.second; });
json vals = json::object();
int limit = std::min(static_cast<int>(sorted.size()), 20);
for (int i = 0; i < limit; ++i) {
vals[sorted[i].first] = sorted[i].second;
}
dist["type"] = "frequency";
dist["values"] = vals;
}
return dist;
}
// ---------------------------------------------------------------------------
// Correlation matrix
// ---------------------------------------------------------------------------
static json computeCorrelationMatrix(const Dataset& ds,
const std::map<std::string, ColType>& types) {
std::vector<std::string> num_cols;
for (const auto& h : ds.headers) {
auto it = types.find(h);
if (it != types.end() && isNumericType(it->second)) {
num_cols.push_back(h);
}
}
json matrix = json::object();
if (num_cols.size() < 2) return matrix;
// Pre-parse numeric columns
std::map<std::string, std::vector<double>> parsed;
for (const auto& col : num_cols) {
auto& vals = ds.columns.at(col);
auto& pv = parsed[col];
pv.resize(ds.row_count, std::numeric_limits<double>::quiet_NaN());
for (int i = 0; i < ds.row_count; ++i) {
double d;
if (!isNull(vals[i]) && tryParseDouble(vals[i], d)) {
pv[i] = d;
}
}
}
for (const auto& colA : num_cols) {
matrix[colA] = json::object();
for (const auto& colB : num_cols) {
// Collect paired non-NaN values
std::vector<double> xs, ys;
for (int i = 0; i < ds.row_count; ++i) {
double a = parsed[colA][i];
double b = parsed[colB][i];
if (!std::isnan(a) && !std::isnan(b)) {
xs.push_back(a);
ys.push_back(b);
}
}
if (xs.size() < 2) {
matrix[colA][colB] = nullptr;
continue;
}
double n = static_cast<double>(xs.size());
double sx = std::accumulate(xs.begin(), xs.end(), 0.0);
double sy = std::accumulate(ys.begin(), ys.end(), 0.0);
double sxx = 0, syy = 0, sxy = 0;
for (size_t i = 0; i < xs.size(); ++i) {
sxx += xs[i] * xs[i];
syy += ys[i] * ys[i];
sxy += xs[i] * ys[i];
}
double denom = std::sqrt((n * sxx - sx * sx) * (n * syy - sy * sy));
if (denom == 0.0) {
matrix[colA][colB] = nullptr;
} else {
double r = (n * sxy - sx * sy) / denom;
matrix[colA][colB] = roundVal(r, 4);
}
}
}
return matrix;
}
// ---------------------------------------------------------------------------
// Data quality score
// ---------------------------------------------------------------------------
static double computeDataQualityScore(const Dataset& ds,
const std::map<std::string, ColType>& types) {
int total_cells = ds.row_count * static_cast<int>(ds.headers.size());
if (total_cells == 0) return 0.0;
int null_count = 0;
for (const auto& h : ds.headers) {
for (const auto& v : ds.columns.at(h)) {
if (isNull(v)) null_count++;
}
}
double completeness = 1.0 - static_cast<double>(null_count) / total_cells;
double uniqueness_sum = 0.0;
int uniqueness_count = 0;
for (const auto& h : ds.headers) {
std::vector<std::string> clean;
for (const auto& v : ds.columns.at(h)) {
if (!isNull(v)) clean.push_back(trim(v));
}
if (!clean.empty()) {
std::set<std::string> unique(clean.begin(), clean.end());
uniqueness_sum += static_cast<double>(unique.size()) / clean.size();
uniqueness_count++;
}
}
double avg_uniqueness = uniqueness_count > 0 ? uniqueness_sum / uniqueness_count : 0.0;
double consistency_sum = 0.0;
int consistency_count = 0;
for (const auto& h : ds.headers) {
std::vector<std::string> clean;
for (const auto& v : ds.columns.at(h)) {
if (!isNull(v)) clean.push_back(trim(v));
}
if (!clean.empty()) {
int num = 0;
for (const auto& v : clean) {
double d;
if (tryParseDouble(v, d)) num++;
}
double ratio = static_cast<double>(num) / clean.size();
consistency_sum += std::max(ratio, 1.0 - ratio);
consistency_count++;
}
}
double avg_consistency = consistency_count > 0 ? consistency_sum / consistency_count : 0.0;
return roundVal((completeness * 0.5 + avg_consistency * 0.3 +
std::min(avg_uniqueness, 1.0) * 0.2) * 100, 2);
}
// ---------------------------------------------------------------------------
// Sample dataset generator
// ---------------------------------------------------------------------------
static void generateSampleDataset(const std::string& filename) {
std::mt19937 rng(42);
const std::vector<std::string> names = {
"Alice", "Bob", "Charlie", "Diana", "Eve", "", "Frank", "Grace", "", "Heidi"
};
const std::vector<std::string> depts = {
"Engineering", "Sales", "HR", "Marketing", "", "Engineering", "Sales", ""
};
const std::vector<std::string> dates = {
"2020-01-15", "2021-06-30", "2019-12-01", "", "invalid-date", "2022-03-10"
};
const std::vector<std::string> actives = {"true", "false", "", "yes", "no", "1", "0"};
std::ofstream out(filename);
if (!out) {
std::cerr << "[ERROR] Cannot create " << filename << "\n";
return;
}
out << "id,name,age,salary,department,score,join_date,is_active\n";
std::uniform_int_distribution<int> age_d(18, 80);
std::uniform_real_distribution<double> salary_d(25000.0, 150000.0);
std::uniform_real_distribution<double> score_d(0.0, 100.0);
std::uniform_int_distribution<int> miss_d(0, 99);
for (int i = 0; i < 200; ++i) {
out << (i + 1) << ",";
// name
out << names[rng() % names.size()] << ",";
// age
if (miss_d(rng) < 8) out << ",";
else if (i == 50) out << "999,";
else if (i == 100) out << "-1,";
else out << age_d(rng) << ",";
// salary
if (miss_d(rng) < 8) out << ",";
else {
std::ostringstream oss;
oss << std::fixed << std::setprecision(2);
if (i == 75) oss << -500.0;
else oss << salary_d(rng);
out << oss.str() << ",";
}
// department
out << depts[rng() % depts.size()] << ",";
// score
if (miss_d(rng) < 10) out << ",";
else {
std::ostringstream oss;
oss << std::fixed << std::setprecision(1) << score_d(rng);
out << oss.str() << ",";
}
// join_date
out << dates[rng() % dates.size()] << ",";
// is_active
out << actives[rng() % actives.size()] << "\n";
}
std::cout << "[INFO] Generated sample dataset: " << filename << " (200 rows)\n";
}
// ---------------------------------------------------------------------------
// Console report
// ---------------------------------------------------------------------------
static void printReport(const json& report) {
std::cout << "\n" << std::string(60, '=') << "\n"
<< " DATA PROFILING TOOL - Trial 1 (fast-cpp-csv-parser + nlohmann/json)\n"
<< std::string(60, '=') << "\n\n";
auto& overview = report["dataset_overview"];
std::cout << "--- Dataset Overview ---\n"
<< " Rows: " << overview["rows"] << "\n"
<< " Columns: " << overview["columns"] << "\n"
<< " Total cells: " << overview["total_cells"] << "\n"
<< " Total missing: " << overview["total_missing"]
<< " (" << overview["total_missing_pct"] << "%)\n"
<< "\n Data Quality Score: " << report["data_quality_score"] << " / 100\n\n";
std::cout << "--- Column Profiles ---\n";
std::cout << std::left
<< std::setw(16) << " Column"
<< std::setw(16) << "Type"
<< std::setw(10) << "Missing"
<< std::setw(10) << "Miss%"
<< std::setw(10) << "Unique"
<< "\n " << std::string(60, '-') << "\n";
for (auto& [col_name, profile] : report["column_profiles"].items()) {
std::cout << " " << std::left
<< std::setw(14) << col_name.substr(0, 13)
<< std::setw(16) << profile["detected_type"].get<std::string>()
<< std::setw(10) << profile["missing_count"]
<< std::setw(10) << (std::to_string(profile["missing_percentage"].get<double>()) + "%").substr(0, 8)
<< std::setw(10) << profile["unique_count"]
<< "\n";
}
// Detailed per-column
for (auto& [col_name, profile] : report["column_profiles"].items()) {
std::cout << "\n [" << col_name << "]\n"
<< " Type: " << profile["detected_type"] << "\n"
<< " Missing: " << profile["missing_count"]
<< " (" << profile["missing_percentage"] << "%)\n"
<< " Unique: " << profile["unique_count"] << "\n";
if (profile.contains("numeric_stats") && !profile["numeric_stats"].empty()) {
auto& ns = profile["numeric_stats"];
for (auto& [k, v] : ns.items()) {
std::cout << " " << std::left << std::setw(15) << (k + ":")
<< " " << v << "\n";
}
}
if (profile.contains("categorical_stats") && !profile["categorical_stats"].empty()) {
auto& cs = profile["categorical_stats"];
if (cs.contains("mode"))
std::cout << " Mode: " << cs["mode"]
<< " (freq: " << cs["mode_frequency"] << ")\n";
if (cs.contains("top_values")) {
std::cout << " Top values:\n";
int cnt = 0;
for (auto& [val, freq] : cs["top_values"].items()) {
if (cnt++ >= 5) break;
std::cout << " " << val << ": " << freq << "\n";
}
}
}
}
// Correlation matrix
if (report.contains("correlation_matrix") && !report["correlation_matrix"].empty()) {
std::cout << "\n--- Correlation Matrix ---\n";
auto& corr = report["correlation_matrix"];
std::vector<std::string> cols;
for (auto& [k, _] : corr.items()) cols.push_back(k);
std::cout << " " << std::left << std::setw(15) << "";
for (const auto& c : cols) std::cout << std::setw(14) << c.substr(0, 12);
std::cout << "\n";
for (const auto& r : cols) {
std::cout << " " << std::left << std::setw(15) << r.substr(0, 14);
for (const auto& c : cols) {
if (corr[r][c].is_null())
std::cout << std::setw(14) << "N/A";
else
std::cout << std::setw(14) << fmtDbl(corr[r][c].get<double>());
}
std::cout << "\n";
}
}
}
// ---------------------------------------------------------------------------
// main
// ---------------------------------------------------------------------------
int main(int argc, char* argv[]) {
std::string filepath;
if (argc < 2) {
std::cout << "[INFO] No input file provided. Generating sample dataset...\n";
filepath = "sample_data.csv";
generateSampleDataset(filepath);
} else {
filepath = argv[1];
}
std::ifstream check(filepath);
if (!check.good()) {
std::cerr << "[ERROR] File not found: " << filepath << "\n";
return 1;
}
check.close();
try {
Dataset ds = readCSV(filepath);
if (ds.row_count == 0) {
std::cout << "[WARN] File contains only a header row.\n";
return 0;
}
std::cout << "[INFO] Loaded file: " << filepath << "\n"
<< "[INFO] Shape: " << ds.row_count << " rows x "
<< ds.headers.size() << " columns\n";
// Detect column types
std::map<std::string, ColType> col_types;
for (const auto& h : ds.headers) {
col_types[h] = detectColumnType(ds.columns[h]);
}
// Build column profiles
json column_profiles = json::object();
for (const auto& h : ds.headers) {
json profile;
profile["name"] = h;
profile["detected_type"] = colTypeStr(col_types[h]);
profile["total_count"] = static_cast<int>(ds.columns[h].size());
int missing = 0;
for (const auto& v : ds.columns[h]) {
if (isNull(v)) missing++;
}
profile["missing_count"] = missing;
profile["missing_percentage"] = roundVal(
ds.row_count > 0 ? static_cast<double>(missing) / ds.row_count * 100 : 0, 2);
std::set<std::string> unique;
for (const auto& v : ds.columns[h]) {
if (!isNull(v)) unique.insert(trim(v));
}
profile["unique_count"] = static_cast<int>(unique.size());
if (isNumericType(col_types[h])) {
profile["numeric_stats"] = computeNumericStats(ds.columns[h]);
} else {
profile["categorical_stats"] = computeCategoricalStats(ds.columns[h]);
}
profile["distribution"] = computeDistribution(ds.columns[h], col_types[h]);
column_profiles[h] = profile;
}
// Correlation matrix
json corr_matrix = computeCorrelationMatrix(ds, col_types);
// Quality score
double quality = computeDataQualityScore(ds, col_types);
// Total missing
int total_missing = 0;
int total_cells = ds.row_count * static_cast<int>(ds.headers.size());
for (const auto& h : ds.headers) {
for (const auto& v : ds.columns[h]) {
if (isNull(v)) total_missing++;
}
}
// Build final report
json report;
report["file"] = filepath;
report["generated_at"] = []() {
auto now = std::chrono::system_clock::now();
auto t = std::chrono::system_clock::to_time_t(now);
std::ostringstream oss;
oss << std::put_time(std::localtime(&t), "%Y-%m-%dT%H:%M:%S");
return oss.str();
}();
json overview;
overview["rows"] = ds.row_count;
overview["columns"] = static_cast<int>(ds.headers.size());
overview["total_cells"] = total_cells;
overview["total_missing"] = total_missing;
overview["total_missing_pct"] = roundVal(
total_cells > 0 ? static_cast<double>(total_missing) / total_cells * 100 : 0, 2);
report["dataset_overview"] = overview;
report["data_quality_score"] = quality;
report["column_profiles"] = column_profiles;
report["correlation_matrix"] = corr_matrix;
// Console output
printReport(report);
// Save JSON
std::ofstream out("profile_report.json");
if (out) {
out << report.dump(2) << "\n";
std::cout << "\n[INFO] Profile report saved: profile_report.json\n";
} else {
std::cerr << "[WARN] Could not write profile_report.json\n";
}
} catch (const std::exception& e) {
std::cerr << "[ERROR] " << e.what() << "\n";
return 1;
}
return 0;
}
README.md
# Data Profiling Tool - C++ Trial 1 Profiles tabular datasets: column types, distributions, missing values, correlations, and quality issues. Reads CSV input, auto-detects column types (numeric, categorical, datetime, boolean), computes per-column statistics, detects missing values, outliers, and unique counts. Outputs profiling report as JSON. ## Dependencies | Library | Version | Purpose | |---------|---------|---------| | [fast-cpp-csv-parser](https://github.com/ben-strasser/fast-cpp-csv-parser) | latest | High-performance CSV line reading | | [nlohmann/json](https://github.com/nlohmann/json) | 3.11.3 | JSON report serialisation | Both libraries are header-only. CMake's `FetchContent` downloads them automatically on the first configure step. ## Build Instructions ```bash # Configure (downloads dependencies on first run) cmake -B build -DCMAKE_BUILD_TYPE=Release # Compile cmake --build build --parallel # Binary is at: build/data_profiler ``` ### System requirements | Tool | Minimum version | |------|----------------| | C++ compiler (g++ / clang++) | C++17 support | | CMake | 3.22 | | git | any (for FetchContent) | ## Usage ```bash # Profile your own CSV file ./build/data_profiler path/to/data.csv # Generate and profile a sample dataset ./build/data_profiler ``` ## Output - Console report with column profiles, distributions, and correlations - `profile_report.json` with full profiling results