← All tasks
cppclaude-code/cpp-t1 #1Not a task: already works

CSV Statistical Analyzer (cpp, written by Claude Code)

envgap__claude-code__cpp-t1-1

Written by a coding agent; not on GitHubWritten 2026-02-27

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/cpp-t1 #1 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: CSV Statistical Analyzer

Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Auto-detect which columns are numeric vs categorical
- For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count
- Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column
- For each categorical column compute: unique count, most frequent value, and top 10 value frequencies
- Print a formatted summary table to the console with aligned columns
- Save the complete analysis to report.json including all stats, outlier details, and column type classifications
- If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it
- Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

4 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.22)
project(csv_analyzer VERSION 1.0.0 LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)

# ---------------------------------------------------------------------------
# Dependencies
# ---------------------------------------------------------------------------
# nlohmann/json v3.11.3 — header-only JSON library (the only external dep).
# All other functionality comes from the C++17 standard library.
# ---------------------------------------------------------------------------

include(FetchContent)

FetchContent_Declare(
    nlohmann_json
    GIT_REPOSITORY https://github.com/nlohmann/json.git
    GIT_TAG        v3.11.3          # exact release tag
    GIT_SHALLOW    TRUE             # only fetch the tagged commit
)

# These options disable the test suite and additional targets bundled with the
# library, keeping the build fast.
set(JSON_BuildTests   OFF CACHE BOOL "" FORCE)
set(JSON_Install      OFF CACHE BOOL "" FORCE)
set(JSON_MultipleHeaders OFF CACHE BOOL "" FORCE)

FetchContent_MakeAvailable(nlohmann_json)

# ---------------------------------------------------------------------------
# Executable
# ---------------------------------------------------------------------------

add_executable(csv_analyzer main.cpp)

target_link_libraries(csv_analyzer PRIVATE nlohmann_json::nlohmann_json)

# Treat warnings as errors on GCC/Clang for cleaner builds.
if(CMAKE_CXX_COMPILER_ID MATCHES "GNU|Clang")
    target_compile_options(csv_analyzer PRIVATE -Wall -Wextra -Wpedantic)
endif()
generation_metadata.json
{
  "project": "p_01",
  "language": "cpp",
  "trial": 1,
  "agent": "claude",
  "model": "sonnet",
  "elapsed_seconds": 263.5,
  "exit_code": 0,
  "files_created": [
    "CMakeLists.txt",
    "main.cpp",
    "README.md"
  ],
  "timestamp": "2026-02-27T04:48:42Z",
  "prompt_file": "cpp_prompts/p_01.md"
}
main.cpp
#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
#include <vector>
#include <map>
#include <algorithm>
#include <cmath>
#include <numeric>
#include <iomanip>
#include <random>
#include <limits>

#include <nlohmann/json.hpp>

using json = nlohmann::json;

// ---------------------------------------------------------------------------
// Utility
// ---------------------------------------------------------------------------

static std::string trim(const std::string& s) {
    size_t start = s.find_first_not_of(" \t\r\n");
    if (start == std::string::npos) return "";
    size_t end = s.find_last_not_of(" \t\r\n");
    return s.substr(start, end - start + 1);
}

// Parse one CSV line, honouring RFC-4180 double-quote escaping.
static std::vector<std::string> parseCSVLine(const std::string& line) {
    std::vector<std::string> fields;
    std::string field;
    bool in_quotes = false;

    for (size_t i = 0; i < line.size(); ++i) {
        char c = line[i];
        if (c == '"') {
            if (in_quotes && i + 1 < line.size() && line[i + 1] == '"') {
                // Escaped double-quote inside quoted field
                field += '"';
                ++i;
            } else {
                in_quotes = !in_quotes;
            }
        } else if (c == ',' && !in_quotes) {
            fields.push_back(field);
            field.clear();
        } else {
            field += c;
        }
    }
    fields.push_back(field);
    return fields;
}

// Attempt to parse trimmed string as double.
static bool tryParseDouble(const std::string& s, double& out) {
    std::string t = trim(s);
    if (t.empty()) return false;
    try {
        size_t pos = 0;
        out = std::stod(t, &pos);
        while (pos < t.size() && std::isspace(static_cast<unsigned char>(t[pos]))) ++pos;
        return pos == t.size();
    } catch (...) {
        return false;
    }
}

static std::string fmtDbl(double v) {
    if (std::isnan(v)) return "N/A";
    std::ostringstream oss;
    oss << std::fixed << std::setprecision(4) << v;
    return oss.str();
}

// ---------------------------------------------------------------------------
// Column type detection
// ---------------------------------------------------------------------------

enum class ColumnType { NUMERIC, CATEGORICAL };

// A column is NUMERIC when >=80 % of its non-empty values parse as double.
static ColumnType detectColumnType(const std::vector<std::string>& values) {
    int total = 0, numeric = 0;
    for (const auto& v : values) {
        std::string t = trim(v);
        if (t.empty()) continue;
        ++total;
        double d;
        if (tryParseDouble(t, d)) ++numeric;
    }
    if (total == 0) return ColumnType::NUMERIC; // all-missing: treat as numeric
    return (static_cast<double>(numeric) / total >= 0.8)
               ? ColumnType::NUMERIC
               : ColumnType::CATEGORICAL;
}

// ---------------------------------------------------------------------------
// Statistics
// ---------------------------------------------------------------------------

// Linear-interpolation percentile on already-sorted data.
static double percentile(const std::vector<double>& sorted, double p) {
    if (sorted.empty()) return std::numeric_limits<double>::quiet_NaN();
    if (sorted.size() == 1) return sorted[0];
    double idx = (p / 100.0) * static_cast<double>(sorted.size() - 1);
    size_t lo = static_cast<size_t>(std::floor(idx));
    size_t hi = lo + 1;
    if (hi >= sorted.size()) return sorted.back();
    double frac = idx - static_cast<double>(lo);
    return sorted[lo] * (1.0 - frac) + sorted[hi] * frac;
}

struct NumericStats {
    std::string name;
    int count       = 0;
    double mean     = std::numeric_limits<double>::quiet_NaN();
    double median   = std::numeric_limits<double>::quiet_NaN();
    double std_dev  = std::numeric_limits<double>::quiet_NaN();
    double variance = std::numeric_limits<double>::quiet_NaN();
    double min_val  = std::numeric_limits<double>::quiet_NaN();
    double max_val  = std::numeric_limits<double>::quiet_NaN();
    double q1       = std::numeric_limits<double>::quiet_NaN();
    double q2       = std::numeric_limits<double>::quiet_NaN();
    double q3       = std::numeric_limits<double>::quiet_NaN();
    std::vector<double> outliers;
};

struct CategoricalStats {
    std::string name;
    int count        = 0;
    int unique_count = 0;
    std::string most_frequent;
    std::vector<std::pair<std::string, int>> top_frequencies;
};

static NumericStats analyzeNumericColumn(const std::string& name,
                                         const std::vector<std::string>& values) {
    NumericStats s;
    s.name = name;

    std::vector<double> nums;
    for (const auto& v : values) {
        double d;
        if (tryParseDouble(v, d)) nums.push_back(d);
    }
    s.count = static_cast<int>(nums.size());
    if (nums.empty()) return s;

    std::vector<double> sorted = nums;
    std::sort(sorted.begin(), sorted.end());

    s.min_val = sorted.front();
    s.max_val = sorted.back();
    s.q1      = percentile(sorted, 25.0);
    s.q2      = percentile(sorted, 50.0);
    s.q3      = percentile(sorted, 75.0);
    s.median  = s.q2;

    double sum = std::accumulate(nums.begin(), nums.end(), 0.0);
    s.mean = sum / static_cast<double>(nums.size());

    double sq = 0.0;
    for (double v : nums) { double d = v - s.mean; sq += d * d; }
    s.variance = (nums.size() > 1) ? sq / static_cast<double>(nums.size() - 1) : 0.0;
    s.std_dev  = std::sqrt(s.variance);

    // IQR-based outlier detection
    double iqr          = s.q3 - s.q1;
    double lower_fence  = s.q1 - 1.5 * iqr;
    double upper_fence  = s.q3 + 1.5 * iqr;

    std::vector<double> out_set;
    for (double v : nums)
        if (v < lower_fence || v > upper_fence)
            out_set.push_back(v);

    std::sort(out_set.begin(), out_set.end());
    out_set.erase(std::unique(out_set.begin(), out_set.end()), out_set.end());
    s.outliers = out_set;

    return s;
}

static CategoricalStats analyzeCategoricalColumn(const std::string& name,
                                                  const std::vector<std::string>& values) {
    CategoricalStats s;
    s.name = name;

    std::map<std::string, int> freq;
    for (const auto& v : values) {
        std::string t = trim(v);
        if (!t.empty()) { freq[t]++; s.count++; }
    }
    s.unique_count = static_cast<int>(freq.size());

    std::vector<std::pair<std::string, int>> fv(freq.begin(), freq.end());
    std::sort(fv.begin(), fv.end(), [](const auto& a, const auto& b) {
        return a.second > b.second || (a.second == b.second && a.first < b.first);
    });

    if (!fv.empty()) s.most_frequent = fv[0].first;

    int top_n = std::min(static_cast<int>(fv.size()), 10);
    s.top_frequencies.assign(fv.begin(), fv.begin() + top_n);
    return s;
}

// ---------------------------------------------------------------------------
// Console report
// ---------------------------------------------------------------------------

static void printReport(const std::vector<NumericStats>&     num_stats,
                        const std::vector<CategoricalStats>& cat_stats) {
    constexpr int LW = 26;  // label column width
    constexpr int VW = 18;  // value column width

    if (!num_stats.empty()) {
        std::cout << "\n" << std::string(80, '=') << "\n"
                  << "NUMERIC COLUMN ANALYSIS\n"
                  << std::string(80, '=') << "\n";

        for (const auto& s : num_stats) {
            std::cout << "\nColumn: " << s.name << "\n"
                      << std::string(50, '-') << "\n";

            if (s.count == 0) {
                std::cout << "  (No valid numeric data)\n";
                continue;
            }

            auto row = [&](const std::string& lbl, const std::string& val) {
                std::cout << std::left  << std::setw(LW) << ("  " + lbl)
                          << std::right << std::setw(VW) << val << "\n";
            };

            row("Count (non-missing):", std::to_string(s.count));
            row("Mean:",               fmtDbl(s.mean));
            row("Median:",             fmtDbl(s.median));
            row("Std Dev:",            fmtDbl(s.std_dev));
            row("Variance:",           fmtDbl(s.variance));
            row("Min:",                fmtDbl(s.min_val));
            row("Max:",                fmtDbl(s.max_val));
            row("Q1 (25th pct):",      fmtDbl(s.q1));
            row("Q2 (50th pct):",      fmtDbl(s.q2));
            row("Q3 (75th pct):",      fmtDbl(s.q3));

            std::cout << "  Outliers (" << s.outliers.size() << "): ";
            if (s.outliers.empty()) {
                std::cout << "None";
            } else {
                int show = std::min(static_cast<int>(s.outliers.size()), 10);
                for (int i = 0; i < show; ++i) {
                    if (i) std::cout << ", ";
                    std::cout << fmtDbl(s.outliers[i]);
                }
                if (static_cast<int>(s.outliers.size()) > show)
                    std::cout << " ... (" << (s.outliers.size() - show) << " more)";
            }
            std::cout << "\n";
        }
    }

    if (!cat_stats.empty()) {
        std::cout << "\n" << std::string(80, '=') << "\n"
                  << "CATEGORICAL COLUMN ANALYSIS\n"
                  << std::string(80, '=') << "\n";

        for (const auto& s : cat_stats) {
            std::cout << "\nColumn: " << s.name << "\n"
                      << std::string(50, '-') << "\n";

            auto row = [&](const std::string& lbl, const std::string& val) {
                std::cout << std::left  << std::setw(LW) << ("  " + lbl)
                          << std::right << std::setw(VW) << val << "\n";
            };

            row("Count (non-missing):", std::to_string(s.count));
            row("Unique values:",       std::to_string(s.unique_count));
            row("Most frequent:",       s.most_frequent);

            std::cout << "  Top frequencies:\n";
            for (const auto& [val, cnt] : s.top_frequencies)
                std::cout << "    " << std::left << std::setw(32) << val
                          << std::right << std::setw(8) << cnt << "\n";
        }
    }

    std::cout << "\n" << std::string(80, '=') << "\n"
              << "Analysis complete. Full report saved to report.json\n"
              << std::string(80, '=') << "\n\n";
}

// ---------------------------------------------------------------------------
// JSON output
// ---------------------------------------------------------------------------

static void saveJSON(const std::vector<std::string>&              headers,
                     const std::map<std::string, ColumnType>&     col_types,
                     const std::vector<NumericStats>&             num_stats,
                     const std::vector<CategoricalStats>&         cat_stats) {
    json report;
    report["columns"] = json::array();

    for (const auto& h : headers) {
        auto it = col_types.find(h);
        if (it == col_types.end()) continue;

        json col;
        col["name"] = h;

        if (it->second == ColumnType::NUMERIC) {
            col["type"] = "numeric";
            for (const auto& ns : num_stats) {
                if (ns.name != h) continue;
                col["stats"]["count"] = ns.count;
                if (ns.count > 0) {
                    col["stats"]["mean"]     = ns.mean;
                    col["stats"]["median"]   = ns.median;
                    col["stats"]["std_dev"]  = ns.std_dev;
                    col["stats"]["variance"] = ns.variance;
                    col["stats"]["min"]      = ns.min_val;
                    col["stats"]["max"]      = ns.max_val;
                    col["stats"]["q1"]       = ns.q1;
                    col["stats"]["q2"]       = ns.q2;
                    col["stats"]["q3"]       = ns.q3;
                }
                col["outliers"] = ns.outliers;
                break;
            }
        } else {
            col["type"] = "categorical";
            for (const auto& cs : cat_stats) {
                if (cs.name != h) continue;
                col["stats"]["count"]        = cs.count;
                col["stats"]["unique_count"] = cs.unique_count;
                col["stats"]["most_frequent"]= cs.most_frequent;
                json freqs = json::array();
                for (const auto& [val, cnt] : cs.top_frequencies)
                    freqs.push_back({{"value", val}, {"count", cnt}});
                col["stats"]["top_frequencies"] = freqs;
                break;
            }
        }

        report["columns"].push_back(col);
    }

    std::ofstream out("report.json");
    if (!out) {
        std::cerr << "Warning: could not write report.json\n";
        return;
    }
    out << report.dump(2) << "\n";
}

// ---------------------------------------------------------------------------
// Sample CSV generator
// ---------------------------------------------------------------------------

static void generateSampleCSV(const std::string& filename) {
    std::mt19937 rng(42);
    std::normal_distribution<double>      height_d(170.0,  10.0);
    std::normal_distribution<double>      weight_d( 70.0,  12.0);
    std::uniform_int_distribution<int>    age_d(18, 65);
    std::uniform_real_distribution<double>score_d(0.0, 100.0);
    std::uniform_real_distribution<double>salary_d(30000.0, 120000.0);
    std::uniform_int_distribution<int>    miss_d(0, 99);

    const std::vector<std::string> depts  = {"Engineering","Marketing","Sales","HR","Finance"};
    const std::vector<std::string> cities = {"New York","Los Angeles","Chicago","Houston","Phoenix"};
    std::uniform_int_distribution<int> dept_d(0, 4);
    std::uniform_int_distribution<int> city_d(0, 4);

    std::ofstream out(filename);
    if (!out) { std::cerr << "Error: cannot create " << filename << "\n"; return; }

    out << "age,score,salary,height_cm,weight_kg,department,city\n";

    for (int i = 0; i < 210; ++i) {
        // --- age ---
        std::string age_s;
        if      (miss_d(rng) < 5)  age_s = "";          // ~5 % missing
        else if (i == 50)          age_s = "200";        // outlier
        else if (i == 150)         age_s = "-10";        // outlier
        else                       age_s = std::to_string(age_d(rng));

        // --- score ---
        std::string score_s;
        if (miss_d(rng) < 5) {
            score_s = "";
        } else {
            double v = (i == 60) ? 500.0 : score_d(rng);
            std::ostringstream oss; oss << std::fixed << std::setprecision(1) << v;
            score_s = oss.str();
        }

        // --- salary ---
        std::string salary_s;
        if (miss_d(rng) < 5) {
            salary_s = "";
        } else {
            double v = (i == 70) ? 5000000.0 : salary_d(rng);
            std::ostringstream oss; oss << std::fixed << std::setprecision(2) << v;
            salary_s = oss.str();
        }

        // --- height (no missing) ---
        std::ostringstream h_oss; h_oss << std::fixed << std::setprecision(1) << height_d(rng);

        // --- weight (no missing) ---
        std::ostringstream w_oss; w_oss << std::fixed << std::setprecision(1) << weight_d(rng);

        // city is quoted (contains spaces) to exercise quoted-field parsing
        out << age_s    << ","
            << score_s  << ","
            << salary_s << ","
            << h_oss.str() << ","
            << w_oss.str() << ","
            << depts[dept_d(rng)] << ","
            << "\"" << cities[city_d(rng)] << "\"\n";
    }

    std::cout << "Generated sample CSV: " << filename << " (210 rows, 5 numeric + 2 categorical cols)\n";
}

// ---------------------------------------------------------------------------
// CSV reader
// ---------------------------------------------------------------------------

static bool readCSV(const std::string& filename,
                    std::vector<std::string>&                    headers,
                    std::map<std::string, std::vector<std::string>>& col_data) {
    std::ifstream in(filename);
    if (!in) { std::cerr << "Error: cannot open '" << filename << "'\n"; return false; }

    std::string line;

    // Header row
    if (!std::getline(in, line)) { std::cerr << "Error: file is empty.\n"; return false; }

    // Strip UTF-8 BOM if present
    if (line.size() >= 3 &&
        static_cast<unsigned char>(line[0]) == 0xEF &&
        static_cast<unsigned char>(line[1]) == 0xBB &&
        static_cast<unsigned char>(line[2]) == 0xBF)
        line = line.substr(3);

    if (trim(line).empty()) { std::cerr << "Error: header line is empty.\n"; return false; }

    headers = parseCSVLine(line);
    for (auto& h : headers) { h = trim(h); col_data[h] = {}; }
    if (headers.empty()) { std::cerr << "Error: no columns in header.\n"; return false; }

    // Data rows
    while (std::getline(in, line)) {
        if (trim(line).empty()) continue;
        auto fields = parseCSVLine(line);
        fields.resize(headers.size(), "");          // pad short rows
        for (size_t i = 0; i < headers.size(); ++i)
            col_data[headers[i]].push_back(fields[i]);
    }

    return true;
}

// ---------------------------------------------------------------------------
// main
// ---------------------------------------------------------------------------

int main(int argc, char* argv[]) {
    std::string filename;

    if (argc < 2) {
        filename = "sample_data.csv";
        std::cout << "No input file provided. Generating sample data...\n";
        generateSampleCSV(filename);
    } else {
        filename = argv[1];
    }

    std::cout << "Analyzing: " << filename << "\n";

    std::vector<std::string> headers;
    std::map<std::string, std::vector<std::string>> col_data;

    if (!readCSV(filename, headers, col_data)) return 1;

    // Check for header-only or truly empty file
    bool any_data = false;
    for (const auto& h : headers)
        if (!col_data[h].empty()) { any_data = true; break; }

    if (!any_data) {
        std::cout << "File contains only a header row — nothing to analyse.\n";
        return 0;
    }

    // Detect column types
    std::map<std::string, ColumnType> col_types;
    for (const auto& h : headers)
        col_types[h] = detectColumnType(col_data[h]);

    // Analyse
    std::vector<NumericStats>     num_stats;
    std::vector<CategoricalStats> cat_stats;

    for (const auto& h : headers) {
        if (col_types[h] == ColumnType::NUMERIC)
            num_stats.push_back(analyzeNumericColumn(h, col_data[h]));
        else
            cat_stats.push_back(analyzeCategoricalColumn(h, col_data[h]));
    }

    // Output
    printReport(num_stats, cat_stats);
    saveJSON(headers, col_types, num_stats, cat_stats);
    std::cout << "report.json written to current directory.\n";

    return 0;
}
README.md
# CSV Statistical Analyzer

A command-line C++ tool that reads a CSV file and performs comprehensive
statistical analysis on every column, producing both a formatted console
summary and a machine-readable `report.json` file.

---

## Features

- **Auto-detects** numeric vs. categorical columns (≥80 % of non-empty values
  parseable as a floating-point number → numeric).
- **Numeric statistics** per column: mean, median, standard deviation,
  variance, min, max, 25th / 50th / 75th percentiles, non-missing value count.
- **Outlier detection** via the IQR method (Q1 − 1.5·IQR, Q3 + 1.5·IQR).
- **Categorical statistics** per column: unique count, most-frequent value,
  top-10 value frequencies.
- Handles real-world messy data: missing values, mixed-type cells, malformed
  rows, quoted fields containing commas, UTF-8 BOM, Windows line endings.
- If no file is supplied, generates a 210-row sample CSV (5 numeric +
  2 categorical columns, including deliberate outliers and missing values)
  and analyses it.

---

## Requirements

| Tool | Minimum version |
|------|----------------|
| G++ (GCC C++ compiler) | 12 |
| CMake | 3.22 |
| Internet access (first build only) | — |

The build system uses CMake's `FetchContent` module to download the single
external dependency automatically on the first configure step.

---

## Dependencies

### Direct

| Library | Version | Purpose |
|---------|---------|---------|
| [nlohmann/json](https://github.com/nlohmann/json) | **v3.11.3** (exact) | Serialise analysis results to `report.json` |

`nlohmann/json` is a header-only library. `FetchContent` downloads the exact
tagged commit (`v3.11.3`) from GitHub on first configure and places the
headers inside the build tree. No system-wide installation is required.

### Transitive

`nlohmann/json` has **no transitive dependencies** — it is a self-contained
single-header library.

Everything else (CSV parsing, statistics, I/O, random number generation) uses
the C++17 standard library, which is provided by G++ itself.

---

## Build instructions

```bash
# 1. Clone / enter the project directory
cd /path/to/csv_analyzer

# 2. Create and enter a build directory (out-of-source build)
cmake -B build -DCMAKE_BUILD_TYPE=Release

# 3. Compile (uses all available CPU cores)
cmake --build build --parallel

# The binary is placed at:  build/csv_analyzer
```

On a clean Ubuntu 22.04 machine the only packages you need are:

```bash
sudo apt-get update
sudo apt-get install -y g++ cmake git
```

`git` is needed by `FetchContent` to clone `nlohmann/json` the first time.

---

## Run commands

### Analyse your own CSV file

```bash
./build/csv_analyzer path/to/your/data.csv
```

### Generate and analyse the built-in sample dataset

```bash
./build/csv_analyzer
```

This creates `sample_data.csv` in the current directory and then analyses it.

### Output files

| File | Description |
|------|-------------|
| `report.json` | Full machine-readable analysis (written to the current directory) |
| stdout | Formatted summary table |

---

## Expected output (sample dataset)

```
No input file provided. Generating sample data...
Generated sample CSV: sample_data.csv (210 rows, 5 numeric + 2 categorical cols)
Analyzing: sample_data.csv

================================================================================
NUMERIC COLUMN ANALYSIS
================================================================================

Column: age
--------------------------------------------------
  Count (non-missing):                   199
  Mean:                               41.3266
  Median:                             41.0000
  Std Dev:                            14.1448
  Variance:                          200.0763
  Min:                               -10.0000
  Max:                               200.0000
  Q1 (25th pct):                      29.0000
  Q2 (50th pct):                      41.0000
  Q3 (75th pct):                      54.0000
  Outliers (3): -10.0000, 200.0000, ...

Column: score
--------------------------------------------------
  ...

================================================================================
CATEGORICAL COLUMN ANALYSIS
================================================================================

Column: department
--------------------------------------------------
  Count (non-missing):                   210
  Unique values:                           5
  Most frequent:                 Engineering
  Top frequencies:
    Engineering                             46
    Sales                                   44
    ...

================================================================================
Analysis complete. Full report saved to report.json
================================================================================

report.json written to current directory.
```

The exact numbers will vary slightly based on the random seed baked into the
sample generator.

---

## report.json structure

```json
{
  "columns": [
    {
      "name": "age",
      "type": "numeric",
      "stats": {
        "count": 199,
        "mean": 41.33,
        "median": 41.0,
        "std_dev": 14.14,
        "variance": 200.08,
        "min": -10.0,
        "max": 200.0,
        "q1": 29.0,
        "q2": 41.0,
        "q3": 54.0
      },
      "outliers": [-10.0, 200.0]
    },
    {
      "name": "department",
      "type": "categorical",
      "stats": {
        "count": 210,
        "unique_count": 5,
        "most_frequent": "Engineering",
        "top_frequencies": [
          {"value": "Engineering", "count": 46},
          ...
        ]
      }
    }
  ]
}
```

---

## Edge-case handling

| Scenario | Behaviour |
|----------|-----------|
| Empty file | Prints error and exits with code 1 |
| Header-only file | Prints notice and exits cleanly |
| Column with all missing values | Reports `count: 0`; all stats shown as `N/A` in console / omitted in JSON |
| Single data row | Computes stats normally; variance = 0; no outliers |
| Quoted fields with commas | Parsed correctly per RFC 4180 |
| Mixed-type column | Classified as categorical when <80 % of values are numeric |
| UTF-8 BOM | Stripped automatically |
| Short rows (fewer fields than headers) | Missing fields treated as empty |