← All tasks
cppcodex/cpp-t1 #1Not a task: already works

CSV Statistical Analyzer (cpp, written by Codex)

envgap__codex__cpp-t1-1

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/cpp-t1 #1 · read the task the agent was given
Codex wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: CSV Statistical Analyzer

Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Auto-detect which columns are numeric vs categorical
- For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count
- Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column
- For each categorical column compute: unique count, most frequent value, and top 10 value frequencies
- Print a formatted summary table to the console with aligned columns
- Save the complete analysis to report.json including all stats, outlier details, and column type classifications
- If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it
- Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.16)
project(csv_statistical_analyzer VERSION 1.0.0 LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 20)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)

add_executable(csv_statistical_analyzer src/main.cpp)
README.md
# CSV Statistical Analyzer (C++)

A robust C++ CLI that reads CSV data, auto-detects numeric/categorical columns, computes comprehensive statistics, detects IQR outliers, prints aligned console tables, and writes a full `report.json`.

## Features

- Accepts a CSV file path as a command-line argument.
- If no path is provided, generates `sample.csv` (250 rows, 5 numeric + 2 categorical columns) and analyzes it.
- Auto-detects numeric vs categorical columns from data.
- Numeric statistics per column:
  - mean, median, standard deviation, variance
  - min, max
  - percentiles (25th, 50th, 75th)
  - non-missing value count
  - outliers via IQR method (`Q1 - 1.5*IQR`, `Q3 + 1.5*IQR`)
- Categorical statistics per column:
  - unique count
  - most frequent value
  - top 10 value frequencies
- Handles messy real-world CSV data:
  - missing values
  - mixed types
  - malformed rows (length mismatch; padded or truncated and recorded)
  - empty files and header-only files
  - quoted fields containing commas
- Writes machine-readable JSON report to `report.json`.

## Project Structure

- `src/main.cpp` - full analyzer implementation.
- `CMakeLists.txt` - CMake build config.

## Requirements

- Ubuntu 22.04
- C++20 compiler (`g++` 11+ or `clang++` 14+)
- Optional: CMake 3.16+

## Dependencies (Pinned)

This implementation uses only the C++ standard library.

- Direct runtime dependencies: none
- Transitive runtime dependencies: none

## Build and Run

### Option A: g++

```bash
cd TMLR/code_generation/codex_generated/p_01/cpp
g++ -std=c++20 -O2 -Wall -Wextra -pedantic -o csv_statistical_analyzer src/main.cpp
./csv_statistical_analyzer ./your_file.csv
```

Run with no input to auto-generate sample data:

```bash
./csv_statistical_analyzer
```

### Option B: CMake

```bash
cd TMLR/code_generation/codex_generated/p_01/cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release
./build/csv_statistical_analyzer ./your_file.csv
```

Run with no input file:

```bash
./build/csv_statistical_analyzer
```

## Output

- Console report: aligned summary tables for numeric and categorical stats and outlier details.
- `report.json`: full structured analysis with:
  - metadata
  - column type classifications
  - numeric and categorical column stats
  - malformed row records
  - warning messages
- If no input file is provided:
  - `sample.csv` is generated in the working directory.

## Statistical Notes

- Variance and standard deviation are sample-based (`n - 1`) when `n > 1`.
- Percentiles use linear interpolation on sorted values.
- Numeric type detection classifies a column as numeric when at least 80% of non-missing values parse as valid numbers.
src/main.cpp
#include <algorithm>
#include <chrono>
#include <cmath>
#include <cstddef>
#include <ctime>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <limits>
#include <map>
#include <optional>
#include <random>
#include <regex>
#include <set>
#include <sstream>
#include <stdexcept>
#include <string>
#include <unordered_map>
#include <utility>
#include <vector>

namespace {

constexpr std::size_t kDefaultSampleRows = 250;

const std::set<std::string> kMissingTokens = {
    "", "na", "n/a", "null", "none", "nan", "undefined", "-"
};

const std::regex kNumberPattern(R"(^[-+]?(?:\d+(?:\.\d+)?|\.\d+)(?:[eE][-+]?\d+)?$)");
const std::regex kCommaNumberPattern(R"(^[-+]?\d{1,3}(?:,\d{3})+(?:\.\d+)?$)");

struct ColumnEntry {
    std::size_t rowNumber;
    std::string value;
};

struct RowData {
    std::size_t rowNumber;
    std::vector<std::string> values;
};

struct MalformedRow {
    std::size_t rowNumber;
    std::size_t expectedColumns;
    std::size_t actualColumns;
    std::string action;
};

struct NormalizedCsv {
    std::vector<std::string> headers;
    std::vector<RowData> dataRows;
    std::vector<MalformedRow> malformedRows;
    std::size_t skippedBlankRows = 0;
    std::size_t sourceRowCount = 0;
};

struct TypeInfo {
    std::string type;
    std::size_t nonMissingCount = 0;
    std::size_t numericCount = 0;
    double numericRatio = 0.0;
    bool allValuesMissing = true;
};

struct TypeClassification {
    std::string type;
    std::size_t nonMissingCount = 0;
    std::size_t numericValueCount = 0;
    double numericRatio = 0.0;
    bool allValuesMissing = true;
};

struct NumericOutlier {
    std::size_t rowNumber;
    double value;
};

struct NumericStats {
    std::string column;
    std::string type = "numeric";
    double typeConfidence = 0.0;
    std::size_t nonMissingValueCount = 0;
    std::size_t missingValueCount = 0;
    std::size_t invalidNumericValueCount = 0;
    std::optional<double> mean;
    std::optional<double> median;
    std::optional<double> standardDeviation;
    std::optional<double> variance;
    std::optional<double> minimum;
    std::optional<double> maximum;
    std::optional<double> p25;
    std::optional<double> p50;
    std::optional<double> p75;
    std::optional<double> lowerBound;
    std::optional<double> upperBound;
    std::optional<double> iqr;
    std::vector<NumericOutlier> outliers;
};

struct CategoricalFrequency {
    std::string value;
    std::size_t count;
};

struct CategoricalStats {
    std::string column;
    std::string type = "categorical";
    double typeConfidence = 0.0;
    std::size_t nonMissingValueCount = 0;
    std::size_t missingValueCount = 0;
    std::size_t uniqueCount = 0;
    std::optional<std::string> mostFrequentValue;
    std::vector<CategoricalFrequency> topFrequencies;
};

struct Metadata {
    std::string generatedAt;
    std::string inputFile;
    std::size_t totalColumns = 0;
    std::size_t sourceRowCount = 0;
    std::size_t dataRowCount = 0;
    std::size_t malformedRowCount = 0;
    std::size_t skippedBlankRowCount = 0;
};

struct Report {
    Metadata metadata;
    std::map<std::string, TypeClassification> columnClassifications;
    std::map<std::string, NumericStats> numericColumns;
    std::map<std::string, CategoricalStats> categoricalColumns;
    std::vector<MalformedRow> malformedRows;
    std::vector<std::string> warnings;
};

struct AnalysisResult {
    std::map<std::string, TypeClassification> columnClassifications;
    std::map<std::string, NumericStats> numericColumns;
    std::map<std::string, CategoricalStats> categoricalColumns;
};

std::string toLower(std::string text) {
    std::transform(text.begin(), text.end(), text.begin(), [](unsigned char c) {
        return static_cast<char>(std::tolower(c));
    });
    return text;
}

std::string trim(const std::string& input) {
    std::size_t start = 0;
    while (start < input.size() && std::isspace(static_cast<unsigned char>(input[start]))) {
        ++start;
    }

    std::size_t end = input.size();
    while (end > start && std::isspace(static_cast<unsigned char>(input[end - 1]))) {
        --end;
    }

    return input.substr(start, end - start);
}

std::string stripBom(const std::string& text) {
    if (text.size() >= 3 && static_cast<unsigned char>(text[0]) == 0xEF &&
        static_cast<unsigned char>(text[1]) == 0xBB && static_cast<unsigned char>(text[2]) == 0xBF) {
        return text.substr(3);
    }
    return text;
}

bool isMissing(const std::string& value) {
    return kMissingTokens.count(toLower(trim(value))) > 0;
}

std::optional<double> parsePotentialNumber(const std::string& value) {
    std::string text = trim(value);
    if (text.empty()) {
        return std::nullopt;
    }

    try {
        if (std::regex_match(text, kNumberPattern)) {
            double parsed = std::stod(text);
            if (std::isfinite(parsed)) {
                return parsed;
            }
            return std::nullopt;
        }

        if (std::regex_match(text, kCommaNumberPattern)) {
            text.erase(std::remove(text.begin(), text.end(), ','), text.end());
            double parsed = std::stod(text);
            if (std::isfinite(parsed)) {
                return parsed;
            }
            return std::nullopt;
        }
    } catch (const std::exception&) {
        return std::nullopt;
    }

    return std::nullopt;
}

double roundTo(double value, int decimals) {
    double factor = std::pow(10.0, static_cast<double>(decimals));
    return std::round(value * factor) / factor;
}

std::vector<std::vector<std::string>> parseCsv(const std::string& content) {
    std::vector<std::vector<std::string>> rows;
    std::vector<std::string> currentRow;
    std::string currentField;
    bool inQuotes = false;

    for (std::size_t i = 0; i < content.size(); ++i) {
        char ch = content[i];

        if (inQuotes) {
            if (ch == '"') {
                if (i + 1 < content.size() && content[i + 1] == '"') {
                    currentField.push_back('"');
                    ++i;
                } else {
                    inQuotes = false;
                }
            } else {
                currentField.push_back(ch);
            }
            continue;
        }

        if (ch == '"') {
            if (currentField.empty()) {
                inQuotes = true;
            } else {
                currentField.push_back(ch);
            }
            continue;
        }

        if (ch == ',') {
            currentRow.push_back(currentField);
            currentField.clear();
            continue;
        }

        if (ch == '\n') {
            currentRow.push_back(currentField);
            currentField.clear();
            rows.push_back(currentRow);
            currentRow.clear();
            continue;
        }

        if (ch == '\r') {
            currentRow.push_back(currentField);
            currentField.clear();
            rows.push_back(currentRow);
            currentRow.clear();
            if (i + 1 < content.size() && content[i + 1] == '\n') {
                ++i;
            }
            continue;
        }

        currentField.push_back(ch);
    }

    if (!currentField.empty() || !currentRow.empty() || (!content.empty() && content.back() == ',')) {
        currentRow.push_back(currentField);
        rows.push_back(currentRow);
    }

    return rows;
}

std::vector<std::string> uniqueHeaders(const std::vector<std::string>& rawHeaders) {
    std::vector<std::string> headers;
    std::unordered_map<std::string, std::size_t> seen;

    for (std::size_t i = 0; i < rawHeaders.size(); ++i) {
        std::string cleaned = trim(stripBom(rawHeaders[i]));
        if (cleaned.empty()) {
            cleaned = "column_" + std::to_string(i + 1);
        }

        std::size_t count = seen[cleaned];
        seen[cleaned] = count + 1;
        if (count == 0) {
            headers.push_back(cleaned);
        } else {
            headers.push_back(cleaned + "_" + std::to_string(count + 1));
        }
    }

    return headers;
}

NormalizedCsv normalizeRows(const std::vector<std::vector<std::string>>& rows) {
    NormalizedCsv normalized;
    if (rows.empty()) {
        return normalized;
    }

    normalized.headers = uniqueHeaders(rows.front());
    normalized.sourceRowCount = rows.size();
    const std::size_t expectedColumns = normalized.headers.size();

    for (std::size_t i = 1; i < rows.size(); ++i) {
        const auto& current = rows[i];
        bool allBlank = current.empty();
        if (!allBlank) {
            allBlank = true;
            for (const auto& value : current) {
                if (!isMissing(value)) {
                    allBlank = false;
                    break;
                }
            }
        }

        if (allBlank) {
            ++normalized.skippedBlankRows;
            continue;
        }

        if (current.size() != expectedColumns) {
            MalformedRow malformed;
            malformed.rowNumber = i + 1;
            malformed.expectedColumns = expectedColumns;
            malformed.actualColumns = current.size();
            malformed.action = current.size() < expectedColumns ? "padded_with_missing" : "truncated_extra_columns";
            normalized.malformedRows.push_back(malformed);
        }

        RowData rowData;
        rowData.rowNumber = i + 1;
        rowData.values.assign(current.begin(), current.begin() + std::min(expectedColumns, current.size()));
        while (rowData.values.size() < expectedColumns) {
            rowData.values.push_back("");
        }
        normalized.dataRows.push_back(std::move(rowData));
    }

    return normalized;
}

std::optional<double> percentile(const std::vector<double>& sortedValues, double p) {
    if (sortedValues.empty()) {
        return std::nullopt;
    }

    if (sortedValues.size() == 1) {
        return sortedValues[0];
    }

    double rank = (p / 100.0) * static_cast<double>(sortedValues.size() - 1);
    std::size_t lower = static_cast<std::size_t>(std::floor(rank));
    std::size_t upper = static_cast<std::size_t>(std::ceil(rank));

    if (lower == upper) {
        return sortedValues[lower];
    }

    double weight = rank - static_cast<double>(lower);
    return sortedValues[lower] * (1.0 - weight) + sortedValues[upper] * weight;
}

TypeInfo detectColumnType(const std::vector<std::string>& values) {
    TypeInfo info;
    std::size_t nonMissing = 0;
    std::size_t numericCount = 0;

    for (const auto& value : values) {
        if (isMissing(value)) {
            continue;
        }

        ++nonMissing;
        if (parsePotentialNumber(value).has_value()) {
            ++numericCount;
        }
    }

    info.nonMissingCount = nonMissing;
    info.numericCount = numericCount;
    info.numericRatio = nonMissing == 0 ? 0.0 : static_cast<double>(numericCount) / static_cast<double>(nonMissing);
    info.type = (nonMissing > 0 && info.numericRatio >= 0.8) ? "numeric" : "categorical";
    info.allValuesMissing = nonMissing == 0;
    return info;
}
NumericStats analyzeNumericColumn(const std::string& name, const std::vector<ColumnEntry>& entries, const TypeInfo& typeInfo) {
    NumericStats stats;
    stats.column = name;
    stats.typeConfidence = roundTo(typeInfo.numericRatio, 4);

    std::size_t missingCount = 0;
    std::size_t invalidCount = 0;
    std::vector<NumericOutlier> numericEntries;
    std::vector<double> numericValues;

    for (const auto& entry : entries) {
        if (isMissing(entry.value)) {
            ++missingCount;
            continue;
        }

        auto parsed = parsePotentialNumber(entry.value);
        if (parsed.has_value()) {
            numericEntries.push_back({entry.rowNumber, *parsed});
            numericValues.push_back(*parsed);
        } else {
            ++invalidCount;
        }
    }

    stats.nonMissingValueCount = numericValues.size();
    stats.missingValueCount = missingCount;
    stats.invalidNumericValueCount = invalidCount;

    if (numericValues.empty()) {
        return stats;
    }

    std::sort(numericValues.begin(), numericValues.end());
    const std::size_t n = numericValues.size();

    double sum = 0.0;
    for (double value : numericValues) {
        sum += value;
    }
    double mean = sum / static_cast<double>(n);

    auto q1 = percentile(numericValues, 25.0);
    auto q2 = percentile(numericValues, 50.0);
    auto q3 = percentile(numericValues, 75.0);

    double variance = 0.0;
    if (n > 1) {
        for (double value : numericValues) {
            double diff = value - mean;
            variance += diff * diff;
        }
        variance /= static_cast<double>(n - 1);
    }

    stats.mean = mean;
    stats.median = q2;
    stats.standardDeviation = std::sqrt(variance);
    stats.variance = variance;
    stats.minimum = numericValues.front();
    stats.maximum = numericValues.back();
    stats.p25 = q1;
    stats.p50 = q2;
    stats.p75 = q3;

    if (q1.has_value() && q3.has_value()) {
        double iqr = *q3 - *q1;
        double lowerBound = *q1 - 1.5 * iqr;
        double upperBound = *q3 + 1.5 * iqr;

        stats.iqr = iqr;
        stats.lowerBound = lowerBound;
        stats.upperBound = upperBound;

        for (const auto& entry : numericEntries) {
            if (entry.value < lowerBound || entry.value > upperBound) {
                stats.outliers.push_back(entry);
            }
        }
    }

    return stats;
}

CategoricalStats analyzeCategoricalColumn(const std::string& name, const std::vector<ColumnEntry>& entries, const TypeInfo& typeInfo) {
    CategoricalStats stats;
    stats.column = name;
    stats.typeConfidence = roundTo(1.0 - typeInfo.numericRatio, 4);

    std::size_t missingCount = 0;
    std::unordered_map<std::string, std::size_t> frequencies;

    for (const auto& entry : entries) {
        if (isMissing(entry.value)) {
            ++missingCount;
            continue;
        }

        std::string cleaned = trim(entry.value);
        ++frequencies[cleaned];
    }

    stats.missingValueCount = missingCount;
    stats.nonMissingValueCount = entries.size() - missingCount;
    stats.uniqueCount = frequencies.size();

    std::vector<CategoricalFrequency> sorted;
    sorted.reserve(frequencies.size());
    for (const auto& item : frequencies) {
        sorted.push_back({item.first, item.second});
    }

    std::sort(sorted.begin(), sorted.end(), [](const CategoricalFrequency& a, const CategoricalFrequency& b) {
        if (a.count != b.count) {
            return a.count > b.count;
        }
        return a.value < b.value;
    });

    if (!sorted.empty()) {
        stats.mostFrequentValue = sorted.front().value;
    }

    for (std::size_t i = 0; i < std::min<std::size_t>(10, sorted.size()); ++i) {
        stats.topFrequencies.push_back(sorted[i]);
    }

    return stats;
}

AnalysisResult analyzeData(const std::vector<std::string>& headers, const std::vector<RowData>& dataRows) {
    std::map<std::string, std::vector<ColumnEntry>> entriesByColumn;
    for (const auto& header : headers) {
        entriesByColumn[header] = {};
    }

    for (const auto& row : dataRows) {
        for (std::size_t i = 0; i < headers.size(); ++i) {
            entriesByColumn[headers[i]].push_back({row.rowNumber, row.values[i]});
        }
    }

    AnalysisResult result;

    for (const auto& header : headers) {
        const auto& entries = entriesByColumn[header];
        std::vector<std::string> values;
        values.reserve(entries.size());
        for (const auto& entry : entries) {
            values.push_back(entry.value);
        }

        TypeInfo typeInfo = detectColumnType(values);

        TypeClassification classification;
        classification.type = typeInfo.type;
        classification.nonMissingCount = typeInfo.nonMissingCount;
        classification.numericValueCount = typeInfo.numericCount;
        classification.numericRatio = roundTo(typeInfo.numericRatio, 4);
        classification.allValuesMissing = typeInfo.allValuesMissing;
        result.columnClassifications[header] = classification;

        if (typeInfo.type == "numeric") {
            result.numericColumns[header] = analyzeNumericColumn(header, entries, typeInfo);
        } else {
            result.categoricalColumns[header] = analyzeCategoricalColumn(header, entries, typeInfo);
        }
    }

    return result;
}

std::string formatNumber(double value) {
    if (!std::isfinite(value)) {
        return "N/A";
    }

    if (value == 0.0) {
        return "0";
    }

    double abs = std::fabs(value);
    std::ostringstream oss;
    if (abs >= 1'000'000.0 || abs < 0.0001) {
        oss << std::scientific << std::setprecision(4) << value;
    } else if (abs >= 1000.0) {
        oss << std::fixed << std::setprecision(2) << value;
    } else {
        oss << std::fixed << std::setprecision(4) << value;
    }

    std::string out = oss.str();
    auto expPos = out.find('e');
    if (expPos == std::string::npos) {
        while (!out.empty() && out.back() == '0') {
            out.pop_back();
        }
        if (!out.empty() && out.back() == '.') {
            out.pop_back();
        }
        if (out == "-0") {
            out = "0";
        }
    }
    return out;
}

std::string formatNumber(const std::optional<double>& value) {
    return value.has_value() ? formatNumber(*value) : "N/A";
}

std::string padRight(const std::string& text, std::size_t width) {
    if (text.size() >= width) {
        return text;
    }
    return text + std::string(width - text.size(), ' ');
}

void printTable(const std::string& title, const std::vector<std::string>& columns,
                const std::vector<std::vector<std::string>>& rows) {
    std::cout << "\n" << title << "\n";

    if (rows.empty()) {
        std::cout << "  (none)\n";
        return;
    }

    std::vector<std::size_t> widths(columns.size(), 0);
    for (std::size_t i = 0; i < columns.size(); ++i) {
        widths[i] = columns[i].size();
    }

    for (const auto& row : rows) {
        for (std::size_t i = 0; i < row.size(); ++i) {
            widths[i] = std::max(widths[i], row[i].size());
        }
    }

    auto renderRow = [&](const std::vector<std::string>& cells) {
        std::ostringstream line;
        for (std::size_t i = 0; i < widths.size(); ++i) {
            if (i > 0) {
                line << " | ";
            }
            std::string value = i < cells.size() ? cells[i] : "";
            line << padRight(value, widths[i]);
        }
        return line.str();
    };

    std::cout << renderRow(columns) << "\n";
    for (std::size_t i = 0; i < widths.size(); ++i) {
        if (i > 0) {
            std::cout << "-+-";
        }
        std::cout << std::string(widths[i], '-');
    }
    std::cout << "\n";

    for (const auto& row : rows) {
        std::cout << renderRow(row) << "\n";
    }
}
void summarizeForConsole(const Report& report) {
    std::cout << "CSV Statistical Analyzer\n";
    std::cout << "========================\n";
    std::cout << "Input File: " << report.metadata.inputFile << "\n";
    std::cout << "Generated : " << report.metadata.generatedAt << "\n";
    std::cout << "Columns   : " << report.metadata.totalColumns << "\n";
    std::cout << "Data Rows : " << report.metadata.dataRowCount << "\n";
    std::cout << "Malformed : " << report.metadata.malformedRowCount << "\n";

    if (!report.warnings.empty()) {
        std::cout << "\nWarnings:\n";
        for (const auto& warning : report.warnings) {
            std::cout << "- " << warning << "\n";
        }
    }

    std::vector<std::vector<std::string>> numericRows;
    for (const auto& item : report.numericColumns) {
        const auto& col = item.second;
        numericRows.push_back({
            col.column,
            std::to_string(col.nonMissingValueCount),
            formatNumber(col.mean),
            formatNumber(col.median),
            formatNumber(col.standardDeviation),
            formatNumber(col.variance),
            formatNumber(col.minimum),
            formatNumber(col.p25),
            formatNumber(col.p50),
            formatNumber(col.p75),
            formatNumber(col.maximum),
            std::to_string(col.outliers.size())
        });
    }

    printTable("Numeric Columns",
               {"Column", "Count", "Mean", "Median", "StdDev", "Variance", "Min", "P25", "P50", "P75", "Max", "Outliers"},
               numericRows);

    std::vector<std::vector<std::string>> categoricalRows;
    for (const auto& item : report.categoricalColumns) {
        const auto& col = item.second;
        std::vector<std::string> preview;
        for (std::size_t i = 0; i < std::min<std::size_t>(3, col.topFrequencies.size()); ++i) {
            preview.push_back(col.topFrequencies[i].value + " (" + std::to_string(col.topFrequencies[i].count) + ")");
        }

        std::string previewText;
        for (std::size_t i = 0; i < preview.size(); ++i) {
            if (i > 0) {
                previewText += ", ";
            }
            previewText += preview[i];
        }

        categoricalRows.push_back({
            col.column,
            std::to_string(col.nonMissingValueCount),
            std::to_string(col.uniqueCount),
            col.mostFrequentValue.has_value() ? *col.mostFrequentValue : "N/A",
            previewText.empty() ? "N/A" : previewText
        });
    }

    printTable("Categorical Columns",
               {"Column", "NonMissing", "Unique", "Most Frequent", "Top Values (up to 3 shown)"},
               categoricalRows);

    std::cout << "\nOutlier Details (IQR Method)\n";
    if (report.numericColumns.empty()) {
        std::cout << "  No numeric columns were detected.\n";
        return;
    }

    for (const auto& item : report.numericColumns) {
        const auto& col = item.second;
        if (col.outliers.empty()) {
            std::cout << "- " << col.column << ": none\n";
            continue;
        }

        std::ostringstream details;
        for (std::size_t i = 0; i < col.outliers.size(); ++i) {
            if (i > 0) {
                details << ", ";
            }
            details << "row " << col.outliers[i].rowNumber << " -> " << formatNumber(col.outliers[i].value);
        }
        std::cout << "- " << col.column << ": " << details.str() << "\n";
    }
}

std::string jsonEscape(const std::string& input) {
    std::ostringstream escaped;
    for (char ch : input) {
        switch (ch) {
            case '\\': escaped << "\\\\"; break;
            case '"': escaped << "\\\""; break;
            case '\b': escaped << "\\b"; break;
            case '\f': escaped << "\\f"; break;
            case '\n': escaped << "\\n"; break;
            case '\r': escaped << "\\r"; break;
            case '\t': escaped << "\\t"; break;
            default:
                if (static_cast<unsigned char>(ch) < 0x20) {
                    escaped << "\\u" << std::hex << std::setw(4) << std::setfill('0')
                            << static_cast<int>(static_cast<unsigned char>(ch)) << std::dec;
                } else {
                    escaped << ch;
                }
        }
    }
    return escaped.str();
}

std::string jsonNumber(double value) {
    if (!std::isfinite(value)) {
        return "null";
    }
    std::ostringstream oss;
    oss << std::fixed << std::setprecision(15) << value;
    std::string out = oss.str();
    while (!out.empty() && out.back() == '0') {
        out.pop_back();
    }
    if (!out.empty() && out.back() == '.') {
        out.pop_back();
    }
    if (out.empty() || out == "-0") {
        out = "0";
    }
    return out;
}

std::string jsonMaybeNumber(const std::optional<double>& value) {
    return value.has_value() ? jsonNumber(*value) : "null";
}

std::string jsonMaybeString(const std::optional<std::string>& value) {
    if (!value.has_value()) {
        return "null";
    }
    return "\"" + jsonEscape(*value) + "\"";
}

std::string nowIso8601Utc() {
    auto now = std::chrono::system_clock::now();
    std::time_t nowTime = std::chrono::system_clock::to_time_t(now);

    std::tm tmUtc{};
#ifdef _WIN32
    gmtime_s(&tmUtc, &nowTime);
#else
    gmtime_r(&nowTime, &tmUtc);
#endif

    std::ostringstream oss;
    oss << std::put_time(&tmUtc, "%Y-%m-%dT%H:%M:%SZ");
    return oss.str();
}

std::string csvEscape(const std::string& value) {
    if (value.find_first_of(",\"\n\r") != std::string::npos) {
        std::string escaped = value;
        std::size_t pos = 0;
        while ((pos = escaped.find('"', pos)) != std::string::npos) {
            escaped.insert(pos, 1, '"');
            pos += 2;
        }
        return "\"" + escaped + "\"";
    }
    return value;
}

std::string toFixed(double value, int precision) {
    std::ostringstream oss;
    oss << std::fixed << std::setprecision(precision) << value;
    return oss.str();
}

void generateSampleCsv(const std::filesystem::path& filePath, std::size_t rowCount) {
    const std::vector<std::string> headers = {
        "temperature_c", "revenue_usd", "latency_ms", "quality_score", "units_sold", "region", "segment"
    };
    const std::vector<std::string> regions = {"North", "South", "East", "West", "New York, NY"};
    const std::vector<std::string> segments = {"Consumer", "Enterprise", "SMB", "Public Sector"};

    std::mt19937 rng(42);

    auto normal = [&](double mean, double stdDev) {
        std::normal_distribution<double> dist(mean, stdDev);
        return dist(rng);
    };

    std::ofstream out(filePath);
    for (std::size_t i = 0; i < headers.size(); ++i) {
        if (i > 0) {
            out << ",";
        }
        out << headers[i];
    }
    out << "\n";

    for (std::size_t i = 0; i < rowCount; ++i) {
        double temperature = normal(20.0, 7.0);
        double revenue = normal(15000.0, 3500.0);
        double latencyVal = normal(120.0, 35.0);
        double qualityVal = std::clamp(normal(82.0, 9.0), 0.0, 100.0);
        long long unitsVal = std::max(0LL, static_cast<long long>(std::llround(normal(240.0, 80.0))));

        if (i % 57 == 0) {
            revenue *= 3.5;
        }

        std::string latency = (i % 41 == 0) ? "timeout" : toFixed(latencyVal, 3);
        std::string quality = (i % 23 == 0) ? "" : toFixed(qualityVal, 2);
        std::string units = (i % 37 == 0) ? "1,250" : std::to_string(unitsVal);

        std::vector<std::string> row = {
            toFixed(temperature, 3),
            toFixed(revenue, 2),
            latency,
            quality,
            units,
            regions[i % regions.size()],
            segments[(i * 3) % segments.size()]
        };

        for (std::size_t c = 0; c < row.size(); ++c) {
            if (c > 0) {
                out << ",";
            }
            out << csvEscape(row[c]);
        }
        out << "\n";
    }
}
std::string reportToJson(const Report& report) {
    std::ostringstream os;
    auto indent = [](int level) {
        return std::string(level * 2, ' ');
    };

    os << "{\n";

    os << indent(1) << "\"metadata\": {\n";
    os << indent(2) << "\"generatedAt\": \"" << jsonEscape(report.metadata.generatedAt) << "\",\n";
    os << indent(2) << "\"inputFile\": \"" << jsonEscape(report.metadata.inputFile) << "\",\n";
    os << indent(2) << "\"totalColumns\": " << report.metadata.totalColumns << ",\n";
    os << indent(2) << "\"sourceRowCount\": " << report.metadata.sourceRowCount << ",\n";
    os << indent(2) << "\"dataRowCount\": " << report.metadata.dataRowCount << ",\n";
    os << indent(2) << "\"malformedRowCount\": " << report.metadata.malformedRowCount << ",\n";
    os << indent(2) << "\"skippedBlankRowCount\": " << report.metadata.skippedBlankRowCount << "\n";
    os << indent(1) << "},\n";

    os << indent(1) << "\"columnClassifications\": {\n";
    for (auto it = report.columnClassifications.begin(); it != report.columnClassifications.end(); ++it) {
        const auto& name = it->first;
        const auto& info = it->second;
        os << indent(2) << "\"" << jsonEscape(name) << "\": {\n";
        os << indent(3) << "\"type\": \"" << jsonEscape(info.type) << "\",\n";
        os << indent(3) << "\"nonMissingCount\": " << info.nonMissingCount << ",\n";
        os << indent(3) << "\"numericValueCount\": " << info.numericValueCount << ",\n";
        os << indent(3) << "\"numericRatio\": " << jsonNumber(info.numericRatio) << ",\n";
        os << indent(3) << "\"allValuesMissing\": " << (info.allValuesMissing ? "true" : "false") << "\n";
        os << indent(2) << "}";
        if (std::next(it) != report.columnClassifications.end()) {
            os << ",";
        }
        os << "\n";
    }
    os << indent(1) << "},\n";

    os << indent(1) << "\"numericColumns\": {\n";
    for (auto it = report.numericColumns.begin(); it != report.numericColumns.end(); ++it) {
        const auto& name = it->first;
        const auto& col = it->second;
        os << indent(2) << "\"" << jsonEscape(name) << "\": {\n";
        os << indent(3) << "\"column\": \"" << jsonEscape(col.column) << "\",\n";
        os << indent(3) << "\"type\": \"numeric\",\n";
        os << indent(3) << "\"typeConfidence\": " << jsonNumber(col.typeConfidence) << ",\n";
        os << indent(3) << "\"nonMissingValueCount\": " << col.nonMissingValueCount << ",\n";
        os << indent(3) << "\"missingValueCount\": " << col.missingValueCount << ",\n";
        os << indent(3) << "\"invalidNumericValueCount\": " << col.invalidNumericValueCount << ",\n";
        os << indent(3) << "\"mean\": " << jsonMaybeNumber(col.mean) << ",\n";
        os << indent(3) << "\"median\": " << jsonMaybeNumber(col.median) << ",\n";
        os << indent(3) << "\"standardDeviation\": " << jsonMaybeNumber(col.standardDeviation) << ",\n";
        os << indent(3) << "\"variance\": " << jsonMaybeNumber(col.variance) << ",\n";
        os << indent(3) << "\"min\": " << jsonMaybeNumber(col.minimum) << ",\n";
        os << indent(3) << "\"max\": " << jsonMaybeNumber(col.maximum) << ",\n";
        os << indent(3) << "\"percentiles\": {\n";
        os << indent(4) << "\"p25\": " << jsonMaybeNumber(col.p25) << ",\n";
        os << indent(4) << "\"p50\": " << jsonMaybeNumber(col.p50) << ",\n";
        os << indent(4) << "\"p75\": " << jsonMaybeNumber(col.p75) << "\n";
        os << indent(3) << "},\n";
        os << indent(3) << "\"outliers\": {\n";
        os << indent(4) << "\"lowerBound\": " << jsonMaybeNumber(col.lowerBound) << ",\n";
        os << indent(4) << "\"upperBound\": " << jsonMaybeNumber(col.upperBound) << ",\n";
        os << indent(4) << "\"iqr\": " << jsonMaybeNumber(col.iqr) << ",\n";
        os << indent(4) << "\"count\": " << col.outliers.size() << ",\n";
        os << indent(4) << "\"values\": [\n";
        for (std::size_t i = 0; i < col.outliers.size(); ++i) {
            os << indent(5) << "{\"rowNumber\": " << col.outliers[i].rowNumber
               << ", \"value\": " << jsonNumber(col.outliers[i].value) << "}";
            if (i + 1 < col.outliers.size()) {
                os << ",";
            }
            os << "\n";
        }
        os << indent(4) << "]\n";
        os << indent(3) << "}\n";
        os << indent(2) << "}";
        if (std::next(it) != report.numericColumns.end()) {
            os << ",";
        }
        os << "\n";
    }
    os << indent(1) << "},\n";

    os << indent(1) << "\"categoricalColumns\": {\n";
    for (auto it = report.categoricalColumns.begin(); it != report.categoricalColumns.end(); ++it) {
        const auto& name = it->first;
        const auto& col = it->second;
        os << indent(2) << "\"" << jsonEscape(name) << "\": {\n";
        os << indent(3) << "\"column\": \"" << jsonEscape(col.column) << "\",\n";
        os << indent(3) << "\"type\": \"categorical\",\n";
        os << indent(3) << "\"typeConfidence\": " << jsonNumber(col.typeConfidence) << ",\n";
        os << indent(3) << "\"nonMissingValueCount\": " << col.nonMissingValueCount << ",\n";
        os << indent(3) << "\"missingValueCount\": " << col.missingValueCount << ",\n";
        os << indent(3) << "\"uniqueCount\": " << col.uniqueCount << ",\n";
        os << indent(3) << "\"mostFrequentValue\": " << jsonMaybeString(col.mostFrequentValue) << ",\n";
        os << indent(3) << "\"topFrequencies\": [\n";
        for (std::size_t i = 0; i < col.topFrequencies.size(); ++i) {
            os << indent(4) << "{\"value\": \"" << jsonEscape(col.topFrequencies[i].value)
               << "\", \"count\": " << col.topFrequencies[i].count << "}";
            if (i + 1 < col.topFrequencies.size()) {
                os << ",";
            }
            os << "\n";
        }
        os << indent(3) << "]\n";
        os << indent(2) << "}";
        if (std::next(it) != report.categoricalColumns.end()) {
            os << ",";
        }
        os << "\n";
    }
    os << indent(1) << "},\n";

    os << indent(1) << "\"malformedRows\": [\n";
    for (std::size_t i = 0; i < report.malformedRows.size(); ++i) {
        const auto& row = report.malformedRows[i];
        os << indent(2) << "{\"rowNumber\": " << row.rowNumber
           << ", \"expectedColumns\": " << row.expectedColumns
           << ", \"actualColumns\": " << row.actualColumns
           << ", \"action\": \"" << jsonEscape(row.action) << "\"}";
        if (i + 1 < report.malformedRows.size()) {
            os << ",";
        }
        os << "\n";
    }
    os << indent(1) << "],\n";

    os << indent(1) << "\"warnings\": [\n";
    for (std::size_t i = 0; i < report.warnings.size(); ++i) {
        os << indent(2) << "\"" << jsonEscape(report.warnings[i]) << "\"";
        if (i + 1 < report.warnings.size()) {
            os << ",";
        }
        os << "\n";
    }
    os << indent(1) << "]\n";

    os << "}";
    return os.str();
}

std::string readFile(const std::filesystem::path& path) {
    std::ifstream input(path, std::ios::binary);
    if (!input) {
        throw std::runtime_error("Failed to read input file");
    }
    std::ostringstream buffer;
    buffer << input.rdbuf();
    return buffer.str();
}

void writeFile(const std::filesystem::path& path, const std::string& content) {
    std::ofstream output(path, std::ios::binary);
    if (!output) {
        throw std::runtime_error("Failed to write output file");
    }
    output << content;
}

bool isWhitespaceOnly(const std::string& text) {
    for (char ch : text) {
        if (!std::isspace(static_cast<unsigned char>(ch))) {
            return false;
        }
    }
    return true;
}

} // namespace

int main(int argc, char** argv) {
    try {
        std::filesystem::path inputPath;
        if (argc < 2) {
            inputPath = std::filesystem::absolute("sample.csv");
            generateSampleCsv(inputPath, kDefaultSampleRows);
            std::cout << "No input file provided. Generated sample dataset at " << inputPath.string() << "\n";
        } else {
            inputPath = std::filesystem::absolute(argv[1]);
        }

        std::string content;
        try {
            content = readFile(inputPath);
        } catch (const std::exception&) {
            std::cerr << "Failed to read input file: " << inputPath.string() << "\n";
            return 1;
        }

        Report report;
        report.metadata.generatedAt = nowIso8601Utc();
        report.metadata.inputFile = inputPath.string();

        if (isWhitespaceOnly(content)) {
            report.warnings.push_back("The input file is empty. No rows were available for analysis.");
        } else {
            auto parsedRows = parseCsv(content);
            auto normalized = normalizeRows(parsedRows);

            report.metadata.totalColumns = normalized.headers.size();
            report.metadata.sourceRowCount = normalized.sourceRowCount;
            report.metadata.dataRowCount = normalized.dataRows.size();
            report.metadata.malformedRowCount = normalized.malformedRows.size();
            report.metadata.skippedBlankRowCount = normalized.skippedBlankRows;
            report.malformedRows = normalized.malformedRows;

            if (normalized.headers.empty()) {
                report.warnings.push_back("No header row was detected.");
            } else if (normalized.dataRows.empty()) {
                report.warnings.push_back("Header detected, but no data rows were found.");
            }

            auto analyzed = analyzeData(normalized.headers, normalized.dataRows);
            report.columnClassifications = std::move(analyzed.columnClassifications);
            report.numericColumns = std::move(analyzed.numericColumns);
            report.categoricalColumns = std::move(analyzed.categoricalColumns);
        }

        summarizeForConsole(report);

        std::filesystem::path outputPath = std::filesystem::absolute("report.json");
        writeFile(outputPath, reportToJson(report) + "\n");
        std::cout << "\nSaved JSON report to " << outputPath.string() << "\n";

        return 0;
    } catch (const std::exception& ex) {
        std::cerr << "Unexpected failure while analyzing CSV data.\n";
        std::cerr << ex.what() << "\n";
        return 1;
    }
}