Data Profiling Tool (cpp, written by Codex)
envgap__codex__cpp-t1-7
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
Default argument given twice for formatDouble precision parameter
Not a benchmark task.
- Its repair changed source code, so it is not an environment task.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/cpp-t1 #7 · read the task the agent was given
Codex wrote this cpp project from the task below. It does not run on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Data Profiling Tool Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report. FUNCTIONAL REQUIREMENTS: - Accept a CSV or JSON data file path as a command-line argument - Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings) - For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th) - For string columns: compute min/max/average length, most common values (top 10), and unique count - For all columns: count total values, missing/null values, missing percentage, and unique value count - Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates - Compute a pairwise correlation matrix for all numeric columns - Print a formatted summary report to console showing key statistics per column - Save the full profiling report as a JSON file with --output flag (default: data_profile.json) - If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it - Handle files with inconsistent delimiters or encoding issues gracefully Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.22) project(data_profiling_tool_cpp VERSION 1.0.0 LANGUAGES CXX) set(CMAKE_CXX_STANDARD 20) set(CMAKE_CXX_STANDARD_REQUIRED ON) set(CMAKE_CXX_EXTENSIONS OFF) include(FetchContent) # Pinned nlohmann/json release FetchContent_Declare( nlohmann_json URL https://github.com/nlohmann/json/releases/download/v3.11.3/json.tar.xz ) FetchContent_MakeAvailable(nlohmann_json) add_executable(data_profiler src/main.cpp) target_link_libraries(data_profiler PRIVATE nlohmann_json::nlohmann_json)
README.md
# Data Profiling Tool (C++) Profiles CSV/JSON datasets for: - column data type detection (integer/float/boolean/date/categorical/string) - missingness and uniqueness - numeric distribution stats and percentiles - outlier detection - numeric correlation matrix - data quality issues ## Requirements - Ubuntu 22.04 - G++ 12+ - CMake 3.22+ - Network access during CMake configure (for dependency fetch) ## Dependencies - Direct: - `nlohmann/json` pinned to `v3.11.3` via CMake FetchContent URL - Transitive: - none ## Build ```bash cmake -S . -B build cmake --build build ``` ## Run With input: ```bash ./build/data_profiler /path/to/data.csv --output data_profile.json ``` JSON input: ```bash ./build/data_profiler /path/to/data.json --output data_profile.json ``` No input (generates sample data): ```bash ./build/data_profiler ``` ## Output - Console summary table - Full JSON report (`data_profile.json` default)
src/main.cpp
#include <algorithm>
#include <chrono>
#include <cmath>
#include <ctime>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <map>
#include <optional>
#include <numeric>
#include <regex>
#include <set>
#include <sstream>
#include <string>
#include <unordered_map>
#include <vector>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
namespace {
struct ParsedArgs {
std::map<std::string, std::string> options;
std::vector<std::string> positional;
};
struct LoadResult {
std::vector<std::map<std::string, std::string>> rows;
std::string encoding;
std::string delimiter;
};
std::string formatDouble(double value, int precision = 2);
ParsedArgs parseArgs(int argc, char** argv) {
ParsedArgs out;
for (int i = 1; i < argc; i++) {
std::string token = argv[i];
if (token.rfind("--", 0) == 0) {
std::string key = token.substr(2);
if (i + 1 < argc && std::string(argv[i + 1]).rfind("--", 0) != 0) {
out.options[key] = argv[++i];
} else {
out.options[key] = "true";
}
} else {
out.positional.push_back(token);
}
}
return out;
}
std::string readFileBinary(const std::filesystem::path& path) {
std::ifstream in(path, std::ios::binary);
if (!in.is_open()) throw std::runtime_error("Failed to open input file: " + path.string());
std::ostringstream ss;
ss << in.rdbuf();
return ss.str();
}
bool looksUtf8(const std::string& text) {
int continuation = 0;
for (unsigned char c : text) {
if (continuation == 0) {
if ((c >> 5) == 0x6) continuation = 1;
else if ((c >> 4) == 0xE) continuation = 2;
else if ((c >> 3) == 0x1E) continuation = 3;
else if ((c >> 7) == 0) continuation = 0;
else return false;
} else {
if ((c >> 6) != 0x2) return false;
continuation--;
}
}
return continuation == 0;
}
std::vector<std::string> parseCsvLine(const std::string& line, char delimiter) {
std::vector<std::string> out;
std::string current;
bool inQuotes = false;
for (std::size_t i = 0; i < line.size(); i++) {
char ch = line[i];
if (ch == '"') {
if (inQuotes && i + 1 < line.size() && line[i + 1] == '"') {
current += '"';
i++;
} else {
inQuotes = !inQuotes;
}
} else if (ch == delimiter && !inQuotes) {
out.push_back(current);
current.clear();
} else {
current += ch;
}
}
out.push_back(current);
return out;
}
char detectDelimiter(const std::vector<std::string>& lines) {
std::vector<char> cands = {',', ';', '\t', '|'};
char best = ',';
int bestScore = -1;
for (char c : cands) {
int score = 0;
for (const auto& line : lines) score += static_cast<int>(std::count(line.begin(), line.end(), c));
if (score > bestScore) {
bestScore = score;
best = c;
}
}
return best;
}
LoadResult loadCsv(const std::filesystem::path& path) {
std::string content = readFileBinary(path);
std::string encoding = looksUtf8(content) ? "utf-8" : "latin1";
std::vector<std::string> lines;
std::istringstream ss(content);
std::string line;
while (std::getline(ss, line)) {
if (!line.empty() && line.back() == '\r') line.pop_back();
if (!line.empty()) lines.push_back(line);
}
if (lines.empty()) return LoadResult{{}, encoding, ","};
std::vector<std::string> sample(lines.begin(), lines.begin() + std::min<std::size_t>(5, lines.size()));
char delimiter = detectDelimiter(sample);
std::vector<std::string> headers = parseCsvLine(lines[0], delimiter);
for (auto& h : headers) {
h = std::regex_replace(h, std::regex(R"(^\s+|\s+$)"), "");
}
std::vector<std::map<std::string, std::string>> rows;
for (std::size_t i = 1; i < lines.size(); i++) {
auto values = parseCsvLine(lines[i], delimiter);
std::map<std::string, std::string> row;
for (std::size_t j = 0; j < headers.size(); j++) {
row[headers[j]] = j < values.size() ? values[j] : "";
}
rows.push_back(row);
}
return LoadResult{rows, encoding, std::string(1, delimiter)};
}
std::map<std::string, std::string> jsonObjToRow(const json& obj) {
std::map<std::string, std::string> row;
for (auto it = obj.begin(); it != obj.end(); ++it) {
if (it.value().is_null()) row[it.key()] = "";
else if (it.value().is_string()) row[it.key()] = it.value().get<std::string>();
else row[it.key()] = it.value().dump();
}
return row;
}
LoadResult loadJson(const std::filesystem::path& path) {
std::string content = readFileBinary(path);
std::string encoding = looksUtf8(content) ? "utf-8" : "latin1";
std::vector<std::map<std::string, std::string>> rows;
try {
json root = json::parse(content);
if (root.is_array()) {
for (const auto& item : root) {
if (item.is_object()) rows.push_back(jsonObjToRow(item));
}
} else if (root.is_object()) {
rows.push_back(jsonObjToRow(root));
}
} catch (...) {
std::istringstream ss(content);
std::string line;
while (std::getline(ss, line)) {
if (line.empty()) continue;
json item = json::parse(line);
if (item.is_object()) rows.push_back(jsonObjToRow(item));
}
}
return LoadResult{rows, encoding, ""};
}
bool isMissing(const std::string& value) {
std::string s = value;
std::transform(s.begin(), s.end(), s.begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
s = std::regex_replace(s, std::regex(R"(^\s+|\s+$)"), "");
return s.empty() || s == "null" || s == "na" || s == "n/a" || s == "none";
}
std::optional<double> parseNumber(const std::string& value) {
std::string s = std::regex_replace(value, std::regex(R"(^\s+|\s+$)"), "");
if (s.empty()) return std::nullopt;
if (!std::regex_match(s, std::regex(R"(^[-+]?\d+(\.\d+)?$)"))) return std::nullopt;
try {
return std::stod(s);
} catch (...) {
return std::nullopt;
}
}
std::optional<bool> parseBoolean(const std::string& value) {
std::string s = value;
std::transform(s.begin(), s.end(), s.begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
s = std::regex_replace(s, std::regex(R"(^\s+|\s+$)"), "");
if (s == "true" || s == "1" || s == "yes" || s == "y") return true;
if (s == "false" || s == "0" || s == "no" || s == "n") return false;
return std::nullopt;
}
std::optional<double> parseDateEpoch(const std::string& value) {
std::string s = std::regex_replace(value, std::regex(R"(^\s+|\s+$)"), "");
if (s.empty()) return std::nullopt;
std::tm tm{};
std::istringstream ss1(s);
ss1 >> std::get_time(&tm, "%Y-%m-%d");
if (!ss1.fail()) return static_cast<double>(std::mktime(&tm));
std::istringstream ss2(s);
ss2 >> std::get_time(&tm, "%Y-%m-%d %H:%M:%S");
if (!ss2.fail()) return static_cast<double>(std::mktime(&tm));
if (s.size() >= 19 && s[10] == 'T') {
std::string copy = s.substr(0, 19);
std::istringstream ss3(copy);
ss3 >> std::get_time(&tm, "%Y-%m-%dT%H:%M:%S");
if (!ss3.fail()) return static_cast<double>(std::mktime(&tm));
}
return std::nullopt;
}
double mean(const std::vector<double>& values) {
if (values.empty()) return 0.0;
double sum = 0.0;
for (double v : values) sum += v;
return sum / values.size();
}
double stddev(const std::vector<double>& values, double m) {
if (values.size() <= 1) return 0.0;
double acc = 0.0;
for (double v : values) acc += (v - m) * (v - m);
return std::sqrt(acc / values.size());
}
double skewness(const std::vector<double>& values, double m, double sd) {
if (values.size() < 3 || sd == 0.0) return 0.0;
double acc = 0.0;
for (double v : values) acc += std::pow(v - m, 3);
acc /= values.size();
return acc / std::pow(sd, 3);
}
std::optional<double> percentile(const std::vector<double>& sorted, double p) {
if (sorted.empty()) return std::nullopt;
if (sorted.size() == 1) return sorted[0];
double pos = (sorted.size() - 1) * p;
std::size_t lo = static_cast<std::size_t>(std::floor(pos));
std::size_t hi = static_cast<std::size_t>(std::ceil(pos));
if (lo == hi) return sorted[lo];
double w = pos - lo;
return sorted[lo] + (sorted[hi] - sorted[lo]) * w;
}
std::string detectType(const std::vector<std::string>& values) {
if (values.empty()) return "string";
int boolCount = 0, numCount = 0, dateCount = 0;
for (const auto& v : values) {
if (parseBoolean(v).has_value()) boolCount++;
if (parseNumber(v).has_value()) numCount++;
if (parseDateEpoch(v).has_value()) dateCount++;
}
if (boolCount == static_cast<int>(values.size())) return "boolean";
if (numCount == static_cast<int>(values.size())) {
bool hasFloat = std::any_of(values.begin(), values.end(), [](const std::string& s) { return s.find('.') != std::string::npos; });
return hasFloat ? "float" : "integer";
}
if (dateCount >= std::max(3, static_cast<int>(std::floor(values.size() * 0.9)))) return "date";
std::set<std::string> unique(values.begin(), values.end());
double ratio = values.empty() ? 0.0 : static_cast<double>(unique.size()) / values.size();
if (unique.size() <= 20 || ratio <= 0.1) return "categorical";
return "string";
}
std::vector<std::pair<std::string, int>> topValues(const std::vector<std::string>& values, int n) {
std::map<std::string, int> freq;
for (const auto& v : values) freq[v]++;
std::vector<std::pair<std::string, int>> out(freq.begin(), freq.end());
std::sort(out.begin(), out.end(), [](const auto& a, const auto& b) {
if (a.second != b.second) return a.second > b.second;
return a.first < b.first;
});
if (static_cast<int>(out.size()) > n) out.resize(n);
return out;
}
std::optional<double> correlation(const std::vector<std::optional<double>>& xs, const std::vector<std::optional<double>>& ys) {
std::vector<double> x, y;
std::size_t n = std::min(xs.size(), ys.size());
for (std::size_t i = 0; i < n; i++) {
if (!xs[i].has_value() || !ys[i].has_value()) continue;
x.push_back(*xs[i]);
y.push_back(*ys[i]);
}
if (x.size() < 2) return std::nullopt;
double mx = mean(x), my = mean(y);
double sx = stddev(x, mx), sy = stddev(y, my);
if (sx == 0.0 || sy == 0.0) return 0.0;
double cov = 0.0;
for (std::size_t i = 0; i < x.size(); i++) cov += (x[i] - mx) * (y[i] - my);
cov /= x.size();
return cov / (sx * sy);
}
json profile(const std::vector<std::map<std::string, std::string>>& rows) {
std::set<std::string> columns;
for (const auto& row : rows) {
for (const auto& [k, _] : row) columns.insert(k);
}
json columnProfiles = json::object();
json issues = json::array();
std::vector<std::string> numericCols;
std::map<std::string, std::vector<std::optional<double>>> numericSeries;
for (const auto& col : columns) {
std::vector<std::string> raw;
raw.reserve(rows.size());
for (const auto& row : rows) {
auto it = row.find(col);
raw.push_back(it != row.end() ? it->second : "");
}
int missing = 0;
std::vector<std::string> nonMissing;
for (const auto& v : raw) {
if (isMissing(v)) missing++;
else nonMissing.push_back(v);
}
std::set<std::string> unique(nonMissing.begin(), nonMissing.end());
std::string type = detectType(nonMissing);
double missingPct = raw.empty() ? 0.0 : static_cast<double>(missing) * 100.0 / raw.size();
json cp = {
{"type", type},
{"total_values", raw.size()},
{"missing_values", missing},
{"missing_percentage", std::round(missingPct * 10000.0) / 10000.0},
{"unique_values", unique.size()},
{"quality_issues", json::array()}
};
if (missing == static_cast<int>(raw.size())) {
cp["quality_issues"].push_back("entirely_null");
issues.push_back({{"column", col}, {"issue", "entirely_null"}});
}
if (unique.size() == 1 && !nonMissing.empty()) {
cp["quality_issues"].push_back("single_unique_value");
issues.push_back({{"column", col}, {"issue", "single_unique_value"}});
}
if (type == "integer" || type == "float") {
std::vector<double> nums;
for (const auto& v : nonMissing) {
auto n = parseNumber(v);
if (n.has_value()) nums.push_back(*n);
}
std::sort(nums.begin(), nums.end());
double mu = nums.empty() ? 0.0 : mean(nums);
double sd = nums.empty() ? 0.0 : stddev(nums, mu);
std::vector<double> outliers;
for (double x : nums) if (sd > 0.0 && std::abs(x - mu) > 4 * sd) outliers.push_back(x);
if (!outliers.empty()) {
cp["quality_issues"].push_back("extreme_outliers");
issues.push_back({{"column", col}, {"issue", "extreme_outliers"}, {"count", outliers.size()}});
}
cp["numeric_stats"] = {
{"min", nums.empty() ? json(nullptr) : json(nums.front())},
{"max", nums.empty() ? json(nullptr) : json(nums.back())},
{"mean", nums.empty() ? json(nullptr) : json(mu)},
{"median", percentile(nums, 0.5).has_value() ? json(*percentile(nums, 0.5)) : json(nullptr)},
{"stddev", sd},
{"skewness", skewness(nums, mu, sd)},
{"percentiles", {
{"p25", percentile(nums, 0.25).has_value() ? json(*percentile(nums, 0.25)) : json(nullptr)},
{"p50", percentile(nums, 0.50).has_value() ? json(*percentile(nums, 0.50)) : json(nullptr)},
{"p75", percentile(nums, 0.75).has_value() ? json(*percentile(nums, 0.75)) : json(nullptr)},
{"p95", percentile(nums, 0.95).has_value() ? json(*percentile(nums, 0.95)) : json(nullptr)},
{"p99", percentile(nums, 0.99).has_value() ? json(*percentile(nums, 0.99)) : json(nullptr)}
}},
{"outliers_beyond_4std", outliers}
};
numericCols.push_back(col);
std::vector<std::optional<double>> series;
for (const auto& v : raw) series.push_back(isMissing(v) ? std::nullopt : parseNumber(v));
numericSeries[col] = series;
} else if (type == "string" || type == "categorical") {
std::vector<int> lengths;
int numericLike = 0, dateLike = 0;
for (const auto& v : nonMissing) {
lengths.push_back(static_cast<int>(v.size()));
if (parseNumber(v).has_value()) numericLike++;
if (parseDateEpoch(v).has_value()) dateLike++;
}
if (!nonMissing.empty() && static_cast<double>(numericLike) / nonMissing.size() >= 0.8) {
cp["quality_issues"].push_back("string_looks_numeric");
issues.push_back({{"column", col}, {"issue", "string_looks_numeric"}});
}
if (!nonMissing.empty() && static_cast<double>(dateLike) / nonMissing.size() >= 0.8) {
cp["quality_issues"].push_back("string_looks_date");
issues.push_back({{"column", col}, {"issue", "string_looks_date"}});
}
double avgLen = lengths.empty() ? 0.0 : std::accumulate(lengths.begin(), lengths.end(), 0.0) / lengths.size();
json top = json::array();
for (const auto& [val, count] : topValues(nonMissing, 10)) top.push_back({{"value", val}, {"count", count}});
cp["string_stats"] = {
{"min_length", lengths.empty() ? json(nullptr) : json(*std::min_element(lengths.begin(), lengths.end()))},
{"max_length", lengths.empty() ? json(nullptr) : json(*std::max_element(lengths.begin(), lengths.end()))},
{"avg_length", lengths.empty() ? json(nullptr) : json(avgLen)},
{"unique_count", unique.size()},
{"top_values", top}
};
}
columnProfiles[col] = cp;
}
json corr = json::object();
for (const auto& c1 : numericCols) {
corr[c1] = json::object();
for (const auto& c2 : numericCols) {
if (c1 == c2) corr[c1][c2] = 1.0;
else {
auto v = correlation(numericSeries[c1], numericSeries[c2]);
corr[c1][c2] = v.has_value() ? json(*v) : json(nullptr);
}
}
}
return {
{"row_count", rows.size()},
{"column_count", columns.size()},
{"columns", columnProfiles},
{"correlations", corr},
{"issues", issues}
};
}
void printSummary(const json& report) {
std::vector<std::string> header = {"Column", "Type", "Total", "Missing%", "Unique", "Notes"};
std::vector<std::vector<std::string>> rows;
for (auto it = report["columns"].begin(); it != report["columns"].end(); ++it) {
const std::string col = it.key();
const json& info = it.value();
std::string notes;
for (std::size_t i = 0; i < info["quality_issues"].size(); i++) {
if (i) notes += ",";
notes += info["quality_issues"][i].get<std::string>();
}
rows.push_back({
col,
info["type"].get<std::string>(),
std::to_string(info["total_values"].get<int>()),
formatDouble(info["missing_percentage"].get<double>(), 2),
std::to_string(info["unique_values"].get<int>()),
notes
});
}
std::vector<std::size_t> widths(header.size(), 0);
for (std::size_t i = 0; i < header.size(); i++) widths[i] = header[i].size();
for (const auto& row : rows) for (std::size_t i = 0; i < row.size(); i++) widths[i] = std::max(widths[i], row[i].size());
auto printSep = [&]() {
std::cout << "+-";
for (std::size_t i = 0; i < widths.size(); i++) {
std::cout << std::string(widths[i], '-');
std::cout << (i + 1 < widths.size() ? "-+-" : "-+");
}
std::cout << "\n";
};
auto printRow = [&](const std::vector<std::string>& row) {
std::cout << "| ";
for (std::size_t i = 0; i < row.size(); i++) {
std::cout << row[i] << std::string(widths[i] > row[i].size() ? widths[i] - row[i].size() : 0, ' ');
std::cout << (i + 1 < row.size() ? " | " : " |\n");
}
};
std::cout << "Data Profiling Summary\n";
std::cout << "======================\n";
std::cout << "Rows: " << report["row_count"] << "\n";
std::cout << "Columns: " << report["column_count"] << "\n";
printSep();
printRow(header);
printSep();
for (const auto& row : rows) printRow(row);
printSep();
}
std::vector<std::map<std::string, std::string>> generateSampleRows() {
std::vector<std::map<std::string, std::string>> rows;
std::vector<std::string> countries = {"US", "CA", "GB", "DE", "IN"};
for (int i = 0; i < 1000; i++) {
int age = i % 200 == 0 ? 140 : 20 + (i % 45);
double income = i % 150 == 0 ? 250000 : 30000 + (i % 120) * 800 + (i % 7) * 13.5;
std::ostringstream incomeStr;
incomeStr << std::fixed << std::setprecision(2) << income;
std::ostringstream signupStr;
signupStr << "2025-"
<< std::setw(2) << std::setfill('0') << ((i % 12) + 1)
<< "-"
<< std::setw(2) << std::setfill('0') << ((i % 28) + 1)
<< "T12:00:00Z";
rows.push_back({
{"id", std::to_string(i + 1)},
{"age", std::to_string(age)},
{"income", incomeStr.str()},
{"is_active", i % 2 == 0 ? "true" : "false"},
{"signup_date", signupStr.str()},
{"country", countries[i % countries.size()]},
{"status_code_str", std::to_string(1000 + (i % 4))},
{"comment", i % 25 == 0 ? "" : "note_" + std::to_string(i % 17)},
{"constant_col", "CONST"},
{"all_null_col", ""}
});
}
return rows;
}
std::string escapeCsv(const std::string& value) {
if (value.find(',') != std::string::npos || value.find('"') != std::string::npos || value.find('\n') != std::string::npos) {
std::string out = "\"";
for (char c : value) {
if (c == '"') out += "\"\"";
else out += c;
}
out += "\"";
return out;
}
return value;
}
void writeSampleCsv(const std::vector<std::map<std::string, std::string>>& rows, const std::filesystem::path& path) {
if (rows.empty()) return;
std::vector<std::string> headers;
for (const auto& [k, _] : rows[0]) headers.push_back(k);
std::ofstream out(path);
for (std::size_t i = 0; i < headers.size(); i++) {
if (i) out << ",";
out << headers[i];
}
out << "\n";
for (const auto& row : rows) {
for (std::size_t i = 0; i < headers.size(); i++) {
if (i) out << ",";
auto it = row.find(headers[i]);
out << escapeCsv(it != row.end() ? it->second : "");
}
out << "\n";
}
}
std::string isoNowUtc() {
auto now = std::chrono::system_clock::now();
std::time_t t = std::chrono::system_clock::to_time_t(now);
std::tm tm{};
#ifdef _WIN32
gmtime_s(&tm, &t);
#else
gmtime_r(&t, &tm);
#endif
char buf[32];
std::strftime(buf, sizeof(buf), "%Y-%m-%dT%H:%M:%SZ", &tm);
return buf;
}
std::string formatDouble(double value, int precision = 2) {
std::ostringstream out;
out << std::fixed << std::setprecision(precision) << value;
return out.str();
}
} // namespace
int main(int argc, char** argv) {
ParsedArgs args = parseArgs(argc, argv);
std::filesystem::path outputPath = std::filesystem::absolute(args.options.count("output") ? args.options["output"] : "data_profile.json");
std::vector<std::map<std::string, std::string>> rows;
json metadata = json::object();
try {
if (args.positional.empty()) {
rows = generateSampleRows();
std::filesystem::path sample = std::filesystem::absolute("sample_profile_data.csv");
writeSampleCsv(rows, sample);
metadata["generated_sample"] = sample.string();
std::cout << "No input file provided. Generated sample dataset: " << sample.string() << "\n";
} else {
std::filesystem::path input = std::filesystem::absolute(args.positional[0]);
if (!std::filesystem::exists(input)) {
std::cerr << "Input file not found: " << input.string() << "\n";
return 1;
}
std::string ext = input.extension().string();
std::transform(ext.begin(), ext.end(), ext.begin(), [](unsigned char c) { return static_cast<char>(std::tolower(c)); });
LoadResult loaded;
if (ext == ".json" || ext == ".jsonl" || ext == ".ndjson") {
loaded = loadJson(input);
metadata["input_format"] = "json";
} else {
loaded = loadCsv(input);
metadata["input_format"] = "csv";
metadata["delimiter"] = loaded.delimiter;
}
rows = std::move(loaded.rows);
metadata["encoding"] = loaded.encoding;
metadata["input_file"] = input.string();
}
json report = profile(rows);
report["generated_at"] = isoNowUtc();
for (auto it = metadata.begin(); it != metadata.end(); ++it) report[it.key()] = it.value();
printSummary(report);
std::ofstream out(outputPath);
if (!out.is_open()) throw std::runtime_error("Failed to write output file: " + outputPath.string());
out << report.dump(2) << "\n";
out.close();
std::cout << "Saved JSON profile: " << outputPath.string() << "\n";
return 0;
} catch (const std::exception& e) {
std::cerr << "Failed to profile data: " << e.what() << "\n";
return 1;
}
}