CSV Statistical Analyzer (cpp, written by Codex)
envgap__codex__cpp-t1-1
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/cpp-t1 #1 · read the task the agent was given
Codex wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.16) project(csv_statistical_analyzer VERSION 1.0.0 LANGUAGES CXX) set(CMAKE_CXX_STANDARD 20) set(CMAKE_CXX_STANDARD_REQUIRED ON) set(CMAKE_CXX_EXTENSIONS OFF) add_executable(csv_statistical_analyzer src/main.cpp)
README.md
# CSV Statistical Analyzer (C++) A robust C++ CLI that reads CSV data, auto-detects numeric/categorical columns, computes comprehensive statistics, detects IQR outliers, prints aligned console tables, and writes a full `report.json`. ## Features - Accepts a CSV file path as a command-line argument. - If no path is provided, generates `sample.csv` (250 rows, 5 numeric + 2 categorical columns) and analyzes it. - Auto-detects numeric vs categorical columns from data. - Numeric statistics per column: - mean, median, standard deviation, variance - min, max - percentiles (25th, 50th, 75th) - non-missing value count - outliers via IQR method (`Q1 - 1.5*IQR`, `Q3 + 1.5*IQR`) - Categorical statistics per column: - unique count - most frequent value - top 10 value frequencies - Handles messy real-world CSV data: - missing values - mixed types - malformed rows (length mismatch; padded or truncated and recorded) - empty files and header-only files - quoted fields containing commas - Writes machine-readable JSON report to `report.json`. ## Project Structure - `src/main.cpp` - full analyzer implementation. - `CMakeLists.txt` - CMake build config. ## Requirements - Ubuntu 22.04 - C++20 compiler (`g++` 11+ or `clang++` 14+) - Optional: CMake 3.16+ ## Dependencies (Pinned) This implementation uses only the C++ standard library. - Direct runtime dependencies: none - Transitive runtime dependencies: none ## Build and Run ### Option A: g++ ```bash cd TMLR/code_generation/codex_generated/p_01/cpp g++ -std=c++20 -O2 -Wall -Wextra -pedantic -o csv_statistical_analyzer src/main.cpp ./csv_statistical_analyzer ./your_file.csv ``` Run with no input to auto-generate sample data: ```bash ./csv_statistical_analyzer ``` ### Option B: CMake ```bash cd TMLR/code_generation/codex_generated/p_01/cpp cmake -S . -B build -DCMAKE_BUILD_TYPE=Release cmake --build build --config Release ./build/csv_statistical_analyzer ./your_file.csv ``` Run with no input file: ```bash ./build/csv_statistical_analyzer ``` ## Output - Console report: aligned summary tables for numeric and categorical stats and outlier details. - `report.json`: full structured analysis with: - metadata - column type classifications - numeric and categorical column stats - malformed row records - warning messages - If no input file is provided: - `sample.csv` is generated in the working directory. ## Statistical Notes - Variance and standard deviation are sample-based (`n - 1`) when `n > 1`. - Percentiles use linear interpolation on sorted values. - Numeric type detection classifies a column as numeric when at least 80% of non-missing values parse as valid numbers.
src/main.cpp
#include <algorithm>
#include <chrono>
#include <cmath>
#include <cstddef>
#include <ctime>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <limits>
#include <map>
#include <optional>
#include <random>
#include <regex>
#include <set>
#include <sstream>
#include <stdexcept>
#include <string>
#include <unordered_map>
#include <utility>
#include <vector>
namespace {
constexpr std::size_t kDefaultSampleRows = 250;
const std::set<std::string> kMissingTokens = {
"", "na", "n/a", "null", "none", "nan", "undefined", "-"
};
const std::regex kNumberPattern(R"(^[-+]?(?:\d+(?:\.\d+)?|\.\d+)(?:[eE][-+]?\d+)?$)");
const std::regex kCommaNumberPattern(R"(^[-+]?\d{1,3}(?:,\d{3})+(?:\.\d+)?$)");
struct ColumnEntry {
std::size_t rowNumber;
std::string value;
};
struct RowData {
std::size_t rowNumber;
std::vector<std::string> values;
};
struct MalformedRow {
std::size_t rowNumber;
std::size_t expectedColumns;
std::size_t actualColumns;
std::string action;
};
struct NormalizedCsv {
std::vector<std::string> headers;
std::vector<RowData> dataRows;
std::vector<MalformedRow> malformedRows;
std::size_t skippedBlankRows = 0;
std::size_t sourceRowCount = 0;
};
struct TypeInfo {
std::string type;
std::size_t nonMissingCount = 0;
std::size_t numericCount = 0;
double numericRatio = 0.0;
bool allValuesMissing = true;
};
struct TypeClassification {
std::string type;
std::size_t nonMissingCount = 0;
std::size_t numericValueCount = 0;
double numericRatio = 0.0;
bool allValuesMissing = true;
};
struct NumericOutlier {
std::size_t rowNumber;
double value;
};
struct NumericStats {
std::string column;
std::string type = "numeric";
double typeConfidence = 0.0;
std::size_t nonMissingValueCount = 0;
std::size_t missingValueCount = 0;
std::size_t invalidNumericValueCount = 0;
std::optional<double> mean;
std::optional<double> median;
std::optional<double> standardDeviation;
std::optional<double> variance;
std::optional<double> minimum;
std::optional<double> maximum;
std::optional<double> p25;
std::optional<double> p50;
std::optional<double> p75;
std::optional<double> lowerBound;
std::optional<double> upperBound;
std::optional<double> iqr;
std::vector<NumericOutlier> outliers;
};
struct CategoricalFrequency {
std::string value;
std::size_t count;
};
struct CategoricalStats {
std::string column;
std::string type = "categorical";
double typeConfidence = 0.0;
std::size_t nonMissingValueCount = 0;
std::size_t missingValueCount = 0;
std::size_t uniqueCount = 0;
std::optional<std::string> mostFrequentValue;
std::vector<CategoricalFrequency> topFrequencies;
};
struct Metadata {
std::string generatedAt;
std::string inputFile;
std::size_t totalColumns = 0;
std::size_t sourceRowCount = 0;
std::size_t dataRowCount = 0;
std::size_t malformedRowCount = 0;
std::size_t skippedBlankRowCount = 0;
};
struct Report {
Metadata metadata;
std::map<std::string, TypeClassification> columnClassifications;
std::map<std::string, NumericStats> numericColumns;
std::map<std::string, CategoricalStats> categoricalColumns;
std::vector<MalformedRow> malformedRows;
std::vector<std::string> warnings;
};
struct AnalysisResult {
std::map<std::string, TypeClassification> columnClassifications;
std::map<std::string, NumericStats> numericColumns;
std::map<std::string, CategoricalStats> categoricalColumns;
};
std::string toLower(std::string text) {
std::transform(text.begin(), text.end(), text.begin(), [](unsigned char c) {
return static_cast<char>(std::tolower(c));
});
return text;
}
std::string trim(const std::string& input) {
std::size_t start = 0;
while (start < input.size() && std::isspace(static_cast<unsigned char>(input[start]))) {
++start;
}
std::size_t end = input.size();
while (end > start && std::isspace(static_cast<unsigned char>(input[end - 1]))) {
--end;
}
return input.substr(start, end - start);
}
std::string stripBom(const std::string& text) {
if (text.size() >= 3 && static_cast<unsigned char>(text[0]) == 0xEF &&
static_cast<unsigned char>(text[1]) == 0xBB && static_cast<unsigned char>(text[2]) == 0xBF) {
return text.substr(3);
}
return text;
}
bool isMissing(const std::string& value) {
return kMissingTokens.count(toLower(trim(value))) > 0;
}
std::optional<double> parsePotentialNumber(const std::string& value) {
std::string text = trim(value);
if (text.empty()) {
return std::nullopt;
}
try {
if (std::regex_match(text, kNumberPattern)) {
double parsed = std::stod(text);
if (std::isfinite(parsed)) {
return parsed;
}
return std::nullopt;
}
if (std::regex_match(text, kCommaNumberPattern)) {
text.erase(std::remove(text.begin(), text.end(), ','), text.end());
double parsed = std::stod(text);
if (std::isfinite(parsed)) {
return parsed;
}
return std::nullopt;
}
} catch (const std::exception&) {
return std::nullopt;
}
return std::nullopt;
}
double roundTo(double value, int decimals) {
double factor = std::pow(10.0, static_cast<double>(decimals));
return std::round(value * factor) / factor;
}
std::vector<std::vector<std::string>> parseCsv(const std::string& content) {
std::vector<std::vector<std::string>> rows;
std::vector<std::string> currentRow;
std::string currentField;
bool inQuotes = false;
for (std::size_t i = 0; i < content.size(); ++i) {
char ch = content[i];
if (inQuotes) {
if (ch == '"') {
if (i + 1 < content.size() && content[i + 1] == '"') {
currentField.push_back('"');
++i;
} else {
inQuotes = false;
}
} else {
currentField.push_back(ch);
}
continue;
}
if (ch == '"') {
if (currentField.empty()) {
inQuotes = true;
} else {
currentField.push_back(ch);
}
continue;
}
if (ch == ',') {
currentRow.push_back(currentField);
currentField.clear();
continue;
}
if (ch == '\n') {
currentRow.push_back(currentField);
currentField.clear();
rows.push_back(currentRow);
currentRow.clear();
continue;
}
if (ch == '\r') {
currentRow.push_back(currentField);
currentField.clear();
rows.push_back(currentRow);
currentRow.clear();
if (i + 1 < content.size() && content[i + 1] == '\n') {
++i;
}
continue;
}
currentField.push_back(ch);
}
if (!currentField.empty() || !currentRow.empty() || (!content.empty() && content.back() == ',')) {
currentRow.push_back(currentField);
rows.push_back(currentRow);
}
return rows;
}
std::vector<std::string> uniqueHeaders(const std::vector<std::string>& rawHeaders) {
std::vector<std::string> headers;
std::unordered_map<std::string, std::size_t> seen;
for (std::size_t i = 0; i < rawHeaders.size(); ++i) {
std::string cleaned = trim(stripBom(rawHeaders[i]));
if (cleaned.empty()) {
cleaned = "column_" + std::to_string(i + 1);
}
std::size_t count = seen[cleaned];
seen[cleaned] = count + 1;
if (count == 0) {
headers.push_back(cleaned);
} else {
headers.push_back(cleaned + "_" + std::to_string(count + 1));
}
}
return headers;
}
NormalizedCsv normalizeRows(const std::vector<std::vector<std::string>>& rows) {
NormalizedCsv normalized;
if (rows.empty()) {
return normalized;
}
normalized.headers = uniqueHeaders(rows.front());
normalized.sourceRowCount = rows.size();
const std::size_t expectedColumns = normalized.headers.size();
for (std::size_t i = 1; i < rows.size(); ++i) {
const auto& current = rows[i];
bool allBlank = current.empty();
if (!allBlank) {
allBlank = true;
for (const auto& value : current) {
if (!isMissing(value)) {
allBlank = false;
break;
}
}
}
if (allBlank) {
++normalized.skippedBlankRows;
continue;
}
if (current.size() != expectedColumns) {
MalformedRow malformed;
malformed.rowNumber = i + 1;
malformed.expectedColumns = expectedColumns;
malformed.actualColumns = current.size();
malformed.action = current.size() < expectedColumns ? "padded_with_missing" : "truncated_extra_columns";
normalized.malformedRows.push_back(malformed);
}
RowData rowData;
rowData.rowNumber = i + 1;
rowData.values.assign(current.begin(), current.begin() + std::min(expectedColumns, current.size()));
while (rowData.values.size() < expectedColumns) {
rowData.values.push_back("");
}
normalized.dataRows.push_back(std::move(rowData));
}
return normalized;
}
std::optional<double> percentile(const std::vector<double>& sortedValues, double p) {
if (sortedValues.empty()) {
return std::nullopt;
}
if (sortedValues.size() == 1) {
return sortedValues[0];
}
double rank = (p / 100.0) * static_cast<double>(sortedValues.size() - 1);
std::size_t lower = static_cast<std::size_t>(std::floor(rank));
std::size_t upper = static_cast<std::size_t>(std::ceil(rank));
if (lower == upper) {
return sortedValues[lower];
}
double weight = rank - static_cast<double>(lower);
return sortedValues[lower] * (1.0 - weight) + sortedValues[upper] * weight;
}
TypeInfo detectColumnType(const std::vector<std::string>& values) {
TypeInfo info;
std::size_t nonMissing = 0;
std::size_t numericCount = 0;
for (const auto& value : values) {
if (isMissing(value)) {
continue;
}
++nonMissing;
if (parsePotentialNumber(value).has_value()) {
++numericCount;
}
}
info.nonMissingCount = nonMissing;
info.numericCount = numericCount;
info.numericRatio = nonMissing == 0 ? 0.0 : static_cast<double>(numericCount) / static_cast<double>(nonMissing);
info.type = (nonMissing > 0 && info.numericRatio >= 0.8) ? "numeric" : "categorical";
info.allValuesMissing = nonMissing == 0;
return info;
}
NumericStats analyzeNumericColumn(const std::string& name, const std::vector<ColumnEntry>& entries, const TypeInfo& typeInfo) {
NumericStats stats;
stats.column = name;
stats.typeConfidence = roundTo(typeInfo.numericRatio, 4);
std::size_t missingCount = 0;
std::size_t invalidCount = 0;
std::vector<NumericOutlier> numericEntries;
std::vector<double> numericValues;
for (const auto& entry : entries) {
if (isMissing(entry.value)) {
++missingCount;
continue;
}
auto parsed = parsePotentialNumber(entry.value);
if (parsed.has_value()) {
numericEntries.push_back({entry.rowNumber, *parsed});
numericValues.push_back(*parsed);
} else {
++invalidCount;
}
}
stats.nonMissingValueCount = numericValues.size();
stats.missingValueCount = missingCount;
stats.invalidNumericValueCount = invalidCount;
if (numericValues.empty()) {
return stats;
}
std::sort(numericValues.begin(), numericValues.end());
const std::size_t n = numericValues.size();
double sum = 0.0;
for (double value : numericValues) {
sum += value;
}
double mean = sum / static_cast<double>(n);
auto q1 = percentile(numericValues, 25.0);
auto q2 = percentile(numericValues, 50.0);
auto q3 = percentile(numericValues, 75.0);
double variance = 0.0;
if (n > 1) {
for (double value : numericValues) {
double diff = value - mean;
variance += diff * diff;
}
variance /= static_cast<double>(n - 1);
}
stats.mean = mean;
stats.median = q2;
stats.standardDeviation = std::sqrt(variance);
stats.variance = variance;
stats.minimum = numericValues.front();
stats.maximum = numericValues.back();
stats.p25 = q1;
stats.p50 = q2;
stats.p75 = q3;
if (q1.has_value() && q3.has_value()) {
double iqr = *q3 - *q1;
double lowerBound = *q1 - 1.5 * iqr;
double upperBound = *q3 + 1.5 * iqr;
stats.iqr = iqr;
stats.lowerBound = lowerBound;
stats.upperBound = upperBound;
for (const auto& entry : numericEntries) {
if (entry.value < lowerBound || entry.value > upperBound) {
stats.outliers.push_back(entry);
}
}
}
return stats;
}
CategoricalStats analyzeCategoricalColumn(const std::string& name, const std::vector<ColumnEntry>& entries, const TypeInfo& typeInfo) {
CategoricalStats stats;
stats.column = name;
stats.typeConfidence = roundTo(1.0 - typeInfo.numericRatio, 4);
std::size_t missingCount = 0;
std::unordered_map<std::string, std::size_t> frequencies;
for (const auto& entry : entries) {
if (isMissing(entry.value)) {
++missingCount;
continue;
}
std::string cleaned = trim(entry.value);
++frequencies[cleaned];
}
stats.missingValueCount = missingCount;
stats.nonMissingValueCount = entries.size() - missingCount;
stats.uniqueCount = frequencies.size();
std::vector<CategoricalFrequency> sorted;
sorted.reserve(frequencies.size());
for (const auto& item : frequencies) {
sorted.push_back({item.first, item.second});
}
std::sort(sorted.begin(), sorted.end(), [](const CategoricalFrequency& a, const CategoricalFrequency& b) {
if (a.count != b.count) {
return a.count > b.count;
}
return a.value < b.value;
});
if (!sorted.empty()) {
stats.mostFrequentValue = sorted.front().value;
}
for (std::size_t i = 0; i < std::min<std::size_t>(10, sorted.size()); ++i) {
stats.topFrequencies.push_back(sorted[i]);
}
return stats;
}
AnalysisResult analyzeData(const std::vector<std::string>& headers, const std::vector<RowData>& dataRows) {
std::map<std::string, std::vector<ColumnEntry>> entriesByColumn;
for (const auto& header : headers) {
entriesByColumn[header] = {};
}
for (const auto& row : dataRows) {
for (std::size_t i = 0; i < headers.size(); ++i) {
entriesByColumn[headers[i]].push_back({row.rowNumber, row.values[i]});
}
}
AnalysisResult result;
for (const auto& header : headers) {
const auto& entries = entriesByColumn[header];
std::vector<std::string> values;
values.reserve(entries.size());
for (const auto& entry : entries) {
values.push_back(entry.value);
}
TypeInfo typeInfo = detectColumnType(values);
TypeClassification classification;
classification.type = typeInfo.type;
classification.nonMissingCount = typeInfo.nonMissingCount;
classification.numericValueCount = typeInfo.numericCount;
classification.numericRatio = roundTo(typeInfo.numericRatio, 4);
classification.allValuesMissing = typeInfo.allValuesMissing;
result.columnClassifications[header] = classification;
if (typeInfo.type == "numeric") {
result.numericColumns[header] = analyzeNumericColumn(header, entries, typeInfo);
} else {
result.categoricalColumns[header] = analyzeCategoricalColumn(header, entries, typeInfo);
}
}
return result;
}
std::string formatNumber(double value) {
if (!std::isfinite(value)) {
return "N/A";
}
if (value == 0.0) {
return "0";
}
double abs = std::fabs(value);
std::ostringstream oss;
if (abs >= 1'000'000.0 || abs < 0.0001) {
oss << std::scientific << std::setprecision(4) << value;
} else if (abs >= 1000.0) {
oss << std::fixed << std::setprecision(2) << value;
} else {
oss << std::fixed << std::setprecision(4) << value;
}
std::string out = oss.str();
auto expPos = out.find('e');
if (expPos == std::string::npos) {
while (!out.empty() && out.back() == '0') {
out.pop_back();
}
if (!out.empty() && out.back() == '.') {
out.pop_back();
}
if (out == "-0") {
out = "0";
}
}
return out;
}
std::string formatNumber(const std::optional<double>& value) {
return value.has_value() ? formatNumber(*value) : "N/A";
}
std::string padRight(const std::string& text, std::size_t width) {
if (text.size() >= width) {
return text;
}
return text + std::string(width - text.size(), ' ');
}
void printTable(const std::string& title, const std::vector<std::string>& columns,
const std::vector<std::vector<std::string>>& rows) {
std::cout << "\n" << title << "\n";
if (rows.empty()) {
std::cout << " (none)\n";
return;
}
std::vector<std::size_t> widths(columns.size(), 0);
for (std::size_t i = 0; i < columns.size(); ++i) {
widths[i] = columns[i].size();
}
for (const auto& row : rows) {
for (std::size_t i = 0; i < row.size(); ++i) {
widths[i] = std::max(widths[i], row[i].size());
}
}
auto renderRow = [&](const std::vector<std::string>& cells) {
std::ostringstream line;
for (std::size_t i = 0; i < widths.size(); ++i) {
if (i > 0) {
line << " | ";
}
std::string value = i < cells.size() ? cells[i] : "";
line << padRight(value, widths[i]);
}
return line.str();
};
std::cout << renderRow(columns) << "\n";
for (std::size_t i = 0; i < widths.size(); ++i) {
if (i > 0) {
std::cout << "-+-";
}
std::cout << std::string(widths[i], '-');
}
std::cout << "\n";
for (const auto& row : rows) {
std::cout << renderRow(row) << "\n";
}
}
void summarizeForConsole(const Report& report) {
std::cout << "CSV Statistical Analyzer\n";
std::cout << "========================\n";
std::cout << "Input File: " << report.metadata.inputFile << "\n";
std::cout << "Generated : " << report.metadata.generatedAt << "\n";
std::cout << "Columns : " << report.metadata.totalColumns << "\n";
std::cout << "Data Rows : " << report.metadata.dataRowCount << "\n";
std::cout << "Malformed : " << report.metadata.malformedRowCount << "\n";
if (!report.warnings.empty()) {
std::cout << "\nWarnings:\n";
for (const auto& warning : report.warnings) {
std::cout << "- " << warning << "\n";
}
}
std::vector<std::vector<std::string>> numericRows;
for (const auto& item : report.numericColumns) {
const auto& col = item.second;
numericRows.push_back({
col.column,
std::to_string(col.nonMissingValueCount),
formatNumber(col.mean),
formatNumber(col.median),
formatNumber(col.standardDeviation),
formatNumber(col.variance),
formatNumber(col.minimum),
formatNumber(col.p25),
formatNumber(col.p50),
formatNumber(col.p75),
formatNumber(col.maximum),
std::to_string(col.outliers.size())
});
}
printTable("Numeric Columns",
{"Column", "Count", "Mean", "Median", "StdDev", "Variance", "Min", "P25", "P50", "P75", "Max", "Outliers"},
numericRows);
std::vector<std::vector<std::string>> categoricalRows;
for (const auto& item : report.categoricalColumns) {
const auto& col = item.second;
std::vector<std::string> preview;
for (std::size_t i = 0; i < std::min<std::size_t>(3, col.topFrequencies.size()); ++i) {
preview.push_back(col.topFrequencies[i].value + " (" + std::to_string(col.topFrequencies[i].count) + ")");
}
std::string previewText;
for (std::size_t i = 0; i < preview.size(); ++i) {
if (i > 0) {
previewText += ", ";
}
previewText += preview[i];
}
categoricalRows.push_back({
col.column,
std::to_string(col.nonMissingValueCount),
std::to_string(col.uniqueCount),
col.mostFrequentValue.has_value() ? *col.mostFrequentValue : "N/A",
previewText.empty() ? "N/A" : previewText
});
}
printTable("Categorical Columns",
{"Column", "NonMissing", "Unique", "Most Frequent", "Top Values (up to 3 shown)"},
categoricalRows);
std::cout << "\nOutlier Details (IQR Method)\n";
if (report.numericColumns.empty()) {
std::cout << " No numeric columns were detected.\n";
return;
}
for (const auto& item : report.numericColumns) {
const auto& col = item.second;
if (col.outliers.empty()) {
std::cout << "- " << col.column << ": none\n";
continue;
}
std::ostringstream details;
for (std::size_t i = 0; i < col.outliers.size(); ++i) {
if (i > 0) {
details << ", ";
}
details << "row " << col.outliers[i].rowNumber << " -> " << formatNumber(col.outliers[i].value);
}
std::cout << "- " << col.column << ": " << details.str() << "\n";
}
}
std::string jsonEscape(const std::string& input) {
std::ostringstream escaped;
for (char ch : input) {
switch (ch) {
case '\\': escaped << "\\\\"; break;
case '"': escaped << "\\\""; break;
case '\b': escaped << "\\b"; break;
case '\f': escaped << "\\f"; break;
case '\n': escaped << "\\n"; break;
case '\r': escaped << "\\r"; break;
case '\t': escaped << "\\t"; break;
default:
if (static_cast<unsigned char>(ch) < 0x20) {
escaped << "\\u" << std::hex << std::setw(4) << std::setfill('0')
<< static_cast<int>(static_cast<unsigned char>(ch)) << std::dec;
} else {
escaped << ch;
}
}
}
return escaped.str();
}
std::string jsonNumber(double value) {
if (!std::isfinite(value)) {
return "null";
}
std::ostringstream oss;
oss << std::fixed << std::setprecision(15) << value;
std::string out = oss.str();
while (!out.empty() && out.back() == '0') {
out.pop_back();
}
if (!out.empty() && out.back() == '.') {
out.pop_back();
}
if (out.empty() || out == "-0") {
out = "0";
}
return out;
}
std::string jsonMaybeNumber(const std::optional<double>& value) {
return value.has_value() ? jsonNumber(*value) : "null";
}
std::string jsonMaybeString(const std::optional<std::string>& value) {
if (!value.has_value()) {
return "null";
}
return "\"" + jsonEscape(*value) + "\"";
}
std::string nowIso8601Utc() {
auto now = std::chrono::system_clock::now();
std::time_t nowTime = std::chrono::system_clock::to_time_t(now);
std::tm tmUtc{};
#ifdef _WIN32
gmtime_s(&tmUtc, &nowTime);
#else
gmtime_r(&nowTime, &tmUtc);
#endif
std::ostringstream oss;
oss << std::put_time(&tmUtc, "%Y-%m-%dT%H:%M:%SZ");
return oss.str();
}
std::string csvEscape(const std::string& value) {
if (value.find_first_of(",\"\n\r") != std::string::npos) {
std::string escaped = value;
std::size_t pos = 0;
while ((pos = escaped.find('"', pos)) != std::string::npos) {
escaped.insert(pos, 1, '"');
pos += 2;
}
return "\"" + escaped + "\"";
}
return value;
}
std::string toFixed(double value, int precision) {
std::ostringstream oss;
oss << std::fixed << std::setprecision(precision) << value;
return oss.str();
}
void generateSampleCsv(const std::filesystem::path& filePath, std::size_t rowCount) {
const std::vector<std::string> headers = {
"temperature_c", "revenue_usd", "latency_ms", "quality_score", "units_sold", "region", "segment"
};
const std::vector<std::string> regions = {"North", "South", "East", "West", "New York, NY"};
const std::vector<std::string> segments = {"Consumer", "Enterprise", "SMB", "Public Sector"};
std::mt19937 rng(42);
auto normal = [&](double mean, double stdDev) {
std::normal_distribution<double> dist(mean, stdDev);
return dist(rng);
};
std::ofstream out(filePath);
for (std::size_t i = 0; i < headers.size(); ++i) {
if (i > 0) {
out << ",";
}
out << headers[i];
}
out << "\n";
for (std::size_t i = 0; i < rowCount; ++i) {
double temperature = normal(20.0, 7.0);
double revenue = normal(15000.0, 3500.0);
double latencyVal = normal(120.0, 35.0);
double qualityVal = std::clamp(normal(82.0, 9.0), 0.0, 100.0);
long long unitsVal = std::max(0LL, static_cast<long long>(std::llround(normal(240.0, 80.0))));
if (i % 57 == 0) {
revenue *= 3.5;
}
std::string latency = (i % 41 == 0) ? "timeout" : toFixed(latencyVal, 3);
std::string quality = (i % 23 == 0) ? "" : toFixed(qualityVal, 2);
std::string units = (i % 37 == 0) ? "1,250" : std::to_string(unitsVal);
std::vector<std::string> row = {
toFixed(temperature, 3),
toFixed(revenue, 2),
latency,
quality,
units,
regions[i % regions.size()],
segments[(i * 3) % segments.size()]
};
for (std::size_t c = 0; c < row.size(); ++c) {
if (c > 0) {
out << ",";
}
out << csvEscape(row[c]);
}
out << "\n";
}
}
std::string reportToJson(const Report& report) {
std::ostringstream os;
auto indent = [](int level) {
return std::string(level * 2, ' ');
};
os << "{\n";
os << indent(1) << "\"metadata\": {\n";
os << indent(2) << "\"generatedAt\": \"" << jsonEscape(report.metadata.generatedAt) << "\",\n";
os << indent(2) << "\"inputFile\": \"" << jsonEscape(report.metadata.inputFile) << "\",\n";
os << indent(2) << "\"totalColumns\": " << report.metadata.totalColumns << ",\n";
os << indent(2) << "\"sourceRowCount\": " << report.metadata.sourceRowCount << ",\n";
os << indent(2) << "\"dataRowCount\": " << report.metadata.dataRowCount << ",\n";
os << indent(2) << "\"malformedRowCount\": " << report.metadata.malformedRowCount << ",\n";
os << indent(2) << "\"skippedBlankRowCount\": " << report.metadata.skippedBlankRowCount << "\n";
os << indent(1) << "},\n";
os << indent(1) << "\"columnClassifications\": {\n";
for (auto it = report.columnClassifications.begin(); it != report.columnClassifications.end(); ++it) {
const auto& name = it->first;
const auto& info = it->second;
os << indent(2) << "\"" << jsonEscape(name) << "\": {\n";
os << indent(3) << "\"type\": \"" << jsonEscape(info.type) << "\",\n";
os << indent(3) << "\"nonMissingCount\": " << info.nonMissingCount << ",\n";
os << indent(3) << "\"numericValueCount\": " << info.numericValueCount << ",\n";
os << indent(3) << "\"numericRatio\": " << jsonNumber(info.numericRatio) << ",\n";
os << indent(3) << "\"allValuesMissing\": " << (info.allValuesMissing ? "true" : "false") << "\n";
os << indent(2) << "}";
if (std::next(it) != report.columnClassifications.end()) {
os << ",";
}
os << "\n";
}
os << indent(1) << "},\n";
os << indent(1) << "\"numericColumns\": {\n";
for (auto it = report.numericColumns.begin(); it != report.numericColumns.end(); ++it) {
const auto& name = it->first;
const auto& col = it->second;
os << indent(2) << "\"" << jsonEscape(name) << "\": {\n";
os << indent(3) << "\"column\": \"" << jsonEscape(col.column) << "\",\n";
os << indent(3) << "\"type\": \"numeric\",\n";
os << indent(3) << "\"typeConfidence\": " << jsonNumber(col.typeConfidence) << ",\n";
os << indent(3) << "\"nonMissingValueCount\": " << col.nonMissingValueCount << ",\n";
os << indent(3) << "\"missingValueCount\": " << col.missingValueCount << ",\n";
os << indent(3) << "\"invalidNumericValueCount\": " << col.invalidNumericValueCount << ",\n";
os << indent(3) << "\"mean\": " << jsonMaybeNumber(col.mean) << ",\n";
os << indent(3) << "\"median\": " << jsonMaybeNumber(col.median) << ",\n";
os << indent(3) << "\"standardDeviation\": " << jsonMaybeNumber(col.standardDeviation) << ",\n";
os << indent(3) << "\"variance\": " << jsonMaybeNumber(col.variance) << ",\n";
os << indent(3) << "\"min\": " << jsonMaybeNumber(col.minimum) << ",\n";
os << indent(3) << "\"max\": " << jsonMaybeNumber(col.maximum) << ",\n";
os << indent(3) << "\"percentiles\": {\n";
os << indent(4) << "\"p25\": " << jsonMaybeNumber(col.p25) << ",\n";
os << indent(4) << "\"p50\": " << jsonMaybeNumber(col.p50) << ",\n";
os << indent(4) << "\"p75\": " << jsonMaybeNumber(col.p75) << "\n";
os << indent(3) << "},\n";
os << indent(3) << "\"outliers\": {\n";
os << indent(4) << "\"lowerBound\": " << jsonMaybeNumber(col.lowerBound) << ",\n";
os << indent(4) << "\"upperBound\": " << jsonMaybeNumber(col.upperBound) << ",\n";
os << indent(4) << "\"iqr\": " << jsonMaybeNumber(col.iqr) << ",\n";
os << indent(4) << "\"count\": " << col.outliers.size() << ",\n";
os << indent(4) << "\"values\": [\n";
for (std::size_t i = 0; i < col.outliers.size(); ++i) {
os << indent(5) << "{\"rowNumber\": " << col.outliers[i].rowNumber
<< ", \"value\": " << jsonNumber(col.outliers[i].value) << "}";
if (i + 1 < col.outliers.size()) {
os << ",";
}
os << "\n";
}
os << indent(4) << "]\n";
os << indent(3) << "}\n";
os << indent(2) << "}";
if (std::next(it) != report.numericColumns.end()) {
os << ",";
}
os << "\n";
}
os << indent(1) << "},\n";
os << indent(1) << "\"categoricalColumns\": {\n";
for (auto it = report.categoricalColumns.begin(); it != report.categoricalColumns.end(); ++it) {
const auto& name = it->first;
const auto& col = it->second;
os << indent(2) << "\"" << jsonEscape(name) << "\": {\n";
os << indent(3) << "\"column\": \"" << jsonEscape(col.column) << "\",\n";
os << indent(3) << "\"type\": \"categorical\",\n";
os << indent(3) << "\"typeConfidence\": " << jsonNumber(col.typeConfidence) << ",\n";
os << indent(3) << "\"nonMissingValueCount\": " << col.nonMissingValueCount << ",\n";
os << indent(3) << "\"missingValueCount\": " << col.missingValueCount << ",\n";
os << indent(3) << "\"uniqueCount\": " << col.uniqueCount << ",\n";
os << indent(3) << "\"mostFrequentValue\": " << jsonMaybeString(col.mostFrequentValue) << ",\n";
os << indent(3) << "\"topFrequencies\": [\n";
for (std::size_t i = 0; i < col.topFrequencies.size(); ++i) {
os << indent(4) << "{\"value\": \"" << jsonEscape(col.topFrequencies[i].value)
<< "\", \"count\": " << col.topFrequencies[i].count << "}";
if (i + 1 < col.topFrequencies.size()) {
os << ",";
}
os << "\n";
}
os << indent(3) << "]\n";
os << indent(2) << "}";
if (std::next(it) != report.categoricalColumns.end()) {
os << ",";
}
os << "\n";
}
os << indent(1) << "},\n";
os << indent(1) << "\"malformedRows\": [\n";
for (std::size_t i = 0; i < report.malformedRows.size(); ++i) {
const auto& row = report.malformedRows[i];
os << indent(2) << "{\"rowNumber\": " << row.rowNumber
<< ", \"expectedColumns\": " << row.expectedColumns
<< ", \"actualColumns\": " << row.actualColumns
<< ", \"action\": \"" << jsonEscape(row.action) << "\"}";
if (i + 1 < report.malformedRows.size()) {
os << ",";
}
os << "\n";
}
os << indent(1) << "],\n";
os << indent(1) << "\"warnings\": [\n";
for (std::size_t i = 0; i < report.warnings.size(); ++i) {
os << indent(2) << "\"" << jsonEscape(report.warnings[i]) << "\"";
if (i + 1 < report.warnings.size()) {
os << ",";
}
os << "\n";
}
os << indent(1) << "]\n";
os << "}";
return os.str();
}
std::string readFile(const std::filesystem::path& path) {
std::ifstream input(path, std::ios::binary);
if (!input) {
throw std::runtime_error("Failed to read input file");
}
std::ostringstream buffer;
buffer << input.rdbuf();
return buffer.str();
}
void writeFile(const std::filesystem::path& path, const std::string& content) {
std::ofstream output(path, std::ios::binary);
if (!output) {
throw std::runtime_error("Failed to write output file");
}
output << content;
}
bool isWhitespaceOnly(const std::string& text) {
for (char ch : text) {
if (!std::isspace(static_cast<unsigned char>(ch))) {
return false;
}
}
return true;
}
} // namespace
int main(int argc, char** argv) {
try {
std::filesystem::path inputPath;
if (argc < 2) {
inputPath = std::filesystem::absolute("sample.csv");
generateSampleCsv(inputPath, kDefaultSampleRows);
std::cout << "No input file provided. Generated sample dataset at " << inputPath.string() << "\n";
} else {
inputPath = std::filesystem::absolute(argv[1]);
}
std::string content;
try {
content = readFile(inputPath);
} catch (const std::exception&) {
std::cerr << "Failed to read input file: " << inputPath.string() << "\n";
return 1;
}
Report report;
report.metadata.generatedAt = nowIso8601Utc();
report.metadata.inputFile = inputPath.string();
if (isWhitespaceOnly(content)) {
report.warnings.push_back("The input file is empty. No rows were available for analysis.");
} else {
auto parsedRows = parseCsv(content);
auto normalized = normalizeRows(parsedRows);
report.metadata.totalColumns = normalized.headers.size();
report.metadata.sourceRowCount = normalized.sourceRowCount;
report.metadata.dataRowCount = normalized.dataRows.size();
report.metadata.malformedRowCount = normalized.malformedRows.size();
report.metadata.skippedBlankRowCount = normalized.skippedBlankRows;
report.malformedRows = normalized.malformedRows;
if (normalized.headers.empty()) {
report.warnings.push_back("No header row was detected.");
} else if (normalized.dataRows.empty()) {
report.warnings.push_back("Header detected, but no data rows were found.");
}
auto analyzed = analyzeData(normalized.headers, normalized.dataRows);
report.columnClassifications = std::move(analyzed.columnClassifications);
report.numericColumns = std::move(analyzed.numericColumns);
report.categoricalColumns = std::move(analyzed.categoricalColumns);
}
summarizeForConsole(report);
std::filesystem::path outputPath = std::filesystem::absolute("report.json");
writeFile(outputPath, reportToJson(report) + "\n");
std::cout << "\nSaved JSON report to " << outputPath.string() << "\n";
return 0;
} catch (const std::exception& ex) {
std::cerr << "Unexpected failure while analyzing CSV data.\n";
std::cerr << ex.what() << "\n";
return 1;
}
}