CSV Statistical Analyzer (cpp, written by Claude Code)
envgap__claude-code__cpp-t1-1
Written by a coding agent; not on GitHubWritten 2026-02-27
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
claude-code/cpp-t1 #1 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
4 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.22)
project(csv_analyzer VERSION 1.0.0 LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)
# ---------------------------------------------------------------------------
# Dependencies
# ---------------------------------------------------------------------------
# nlohmann/json v3.11.3 — header-only JSON library (the only external dep).
# All other functionality comes from the C++17 standard library.
# ---------------------------------------------------------------------------
include(FetchContent)
FetchContent_Declare(
nlohmann_json
GIT_REPOSITORY https://github.com/nlohmann/json.git
GIT_TAG v3.11.3 # exact release tag
GIT_SHALLOW TRUE # only fetch the tagged commit
)
# These options disable the test suite and additional targets bundled with the
# library, keeping the build fast.
set(JSON_BuildTests OFF CACHE BOOL "" FORCE)
set(JSON_Install OFF CACHE BOOL "" FORCE)
set(JSON_MultipleHeaders OFF CACHE BOOL "" FORCE)
FetchContent_MakeAvailable(nlohmann_json)
# ---------------------------------------------------------------------------
# Executable
# ---------------------------------------------------------------------------
add_executable(csv_analyzer main.cpp)
target_link_libraries(csv_analyzer PRIVATE nlohmann_json::nlohmann_json)
# Treat warnings as errors on GCC/Clang for cleaner builds.
if(CMAKE_CXX_COMPILER_ID MATCHES "GNU|Clang")
target_compile_options(csv_analyzer PRIVATE -Wall -Wextra -Wpedantic)
endif()
generation_metadata.json
{
"project": "p_01",
"language": "cpp",
"trial": 1,
"agent": "claude",
"model": "sonnet",
"elapsed_seconds": 263.5,
"exit_code": 0,
"files_created": [
"CMakeLists.txt",
"main.cpp",
"README.md"
],
"timestamp": "2026-02-27T04:48:42Z",
"prompt_file": "cpp_prompts/p_01.md"
}main.cpp
#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
#include <vector>
#include <map>
#include <algorithm>
#include <cmath>
#include <numeric>
#include <iomanip>
#include <random>
#include <limits>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
// ---------------------------------------------------------------------------
// Utility
// ---------------------------------------------------------------------------
static std::string trim(const std::string& s) {
size_t start = s.find_first_not_of(" \t\r\n");
if (start == std::string::npos) return "";
size_t end = s.find_last_not_of(" \t\r\n");
return s.substr(start, end - start + 1);
}
// Parse one CSV line, honouring RFC-4180 double-quote escaping.
static std::vector<std::string> parseCSVLine(const std::string& line) {
std::vector<std::string> fields;
std::string field;
bool in_quotes = false;
for (size_t i = 0; i < line.size(); ++i) {
char c = line[i];
if (c == '"') {
if (in_quotes && i + 1 < line.size() && line[i + 1] == '"') {
// Escaped double-quote inside quoted field
field += '"';
++i;
} else {
in_quotes = !in_quotes;
}
} else if (c == ',' && !in_quotes) {
fields.push_back(field);
field.clear();
} else {
field += c;
}
}
fields.push_back(field);
return fields;
}
// Attempt to parse trimmed string as double.
static bool tryParseDouble(const std::string& s, double& out) {
std::string t = trim(s);
if (t.empty()) return false;
try {
size_t pos = 0;
out = std::stod(t, &pos);
while (pos < t.size() && std::isspace(static_cast<unsigned char>(t[pos]))) ++pos;
return pos == t.size();
} catch (...) {
return false;
}
}
static std::string fmtDbl(double v) {
if (std::isnan(v)) return "N/A";
std::ostringstream oss;
oss << std::fixed << std::setprecision(4) << v;
return oss.str();
}
// ---------------------------------------------------------------------------
// Column type detection
// ---------------------------------------------------------------------------
enum class ColumnType { NUMERIC, CATEGORICAL };
// A column is NUMERIC when >=80 % of its non-empty values parse as double.
static ColumnType detectColumnType(const std::vector<std::string>& values) {
int total = 0, numeric = 0;
for (const auto& v : values) {
std::string t = trim(v);
if (t.empty()) continue;
++total;
double d;
if (tryParseDouble(t, d)) ++numeric;
}
if (total == 0) return ColumnType::NUMERIC; // all-missing: treat as numeric
return (static_cast<double>(numeric) / total >= 0.8)
? ColumnType::NUMERIC
: ColumnType::CATEGORICAL;
}
// ---------------------------------------------------------------------------
// Statistics
// ---------------------------------------------------------------------------
// Linear-interpolation percentile on already-sorted data.
static double percentile(const std::vector<double>& sorted, double p) {
if (sorted.empty()) return std::numeric_limits<double>::quiet_NaN();
if (sorted.size() == 1) return sorted[0];
double idx = (p / 100.0) * static_cast<double>(sorted.size() - 1);
size_t lo = static_cast<size_t>(std::floor(idx));
size_t hi = lo + 1;
if (hi >= sorted.size()) return sorted.back();
double frac = idx - static_cast<double>(lo);
return sorted[lo] * (1.0 - frac) + sorted[hi] * frac;
}
struct NumericStats {
std::string name;
int count = 0;
double mean = std::numeric_limits<double>::quiet_NaN();
double median = std::numeric_limits<double>::quiet_NaN();
double std_dev = std::numeric_limits<double>::quiet_NaN();
double variance = std::numeric_limits<double>::quiet_NaN();
double min_val = std::numeric_limits<double>::quiet_NaN();
double max_val = std::numeric_limits<double>::quiet_NaN();
double q1 = std::numeric_limits<double>::quiet_NaN();
double q2 = std::numeric_limits<double>::quiet_NaN();
double q3 = std::numeric_limits<double>::quiet_NaN();
std::vector<double> outliers;
};
struct CategoricalStats {
std::string name;
int count = 0;
int unique_count = 0;
std::string most_frequent;
std::vector<std::pair<std::string, int>> top_frequencies;
};
static NumericStats analyzeNumericColumn(const std::string& name,
const std::vector<std::string>& values) {
NumericStats s;
s.name = name;
std::vector<double> nums;
for (const auto& v : values) {
double d;
if (tryParseDouble(v, d)) nums.push_back(d);
}
s.count = static_cast<int>(nums.size());
if (nums.empty()) return s;
std::vector<double> sorted = nums;
std::sort(sorted.begin(), sorted.end());
s.min_val = sorted.front();
s.max_val = sorted.back();
s.q1 = percentile(sorted, 25.0);
s.q2 = percentile(sorted, 50.0);
s.q3 = percentile(sorted, 75.0);
s.median = s.q2;
double sum = std::accumulate(nums.begin(), nums.end(), 0.0);
s.mean = sum / static_cast<double>(nums.size());
double sq = 0.0;
for (double v : nums) { double d = v - s.mean; sq += d * d; }
s.variance = (nums.size() > 1) ? sq / static_cast<double>(nums.size() - 1) : 0.0;
s.std_dev = std::sqrt(s.variance);
// IQR-based outlier detection
double iqr = s.q3 - s.q1;
double lower_fence = s.q1 - 1.5 * iqr;
double upper_fence = s.q3 + 1.5 * iqr;
std::vector<double> out_set;
for (double v : nums)
if (v < lower_fence || v > upper_fence)
out_set.push_back(v);
std::sort(out_set.begin(), out_set.end());
out_set.erase(std::unique(out_set.begin(), out_set.end()), out_set.end());
s.outliers = out_set;
return s;
}
static CategoricalStats analyzeCategoricalColumn(const std::string& name,
const std::vector<std::string>& values) {
CategoricalStats s;
s.name = name;
std::map<std::string, int> freq;
for (const auto& v : values) {
std::string t = trim(v);
if (!t.empty()) { freq[t]++; s.count++; }
}
s.unique_count = static_cast<int>(freq.size());
std::vector<std::pair<std::string, int>> fv(freq.begin(), freq.end());
std::sort(fv.begin(), fv.end(), [](const auto& a, const auto& b) {
return a.second > b.second || (a.second == b.second && a.first < b.first);
});
if (!fv.empty()) s.most_frequent = fv[0].first;
int top_n = std::min(static_cast<int>(fv.size()), 10);
s.top_frequencies.assign(fv.begin(), fv.begin() + top_n);
return s;
}
// ---------------------------------------------------------------------------
// Console report
// ---------------------------------------------------------------------------
static void printReport(const std::vector<NumericStats>& num_stats,
const std::vector<CategoricalStats>& cat_stats) {
constexpr int LW = 26; // label column width
constexpr int VW = 18; // value column width
if (!num_stats.empty()) {
std::cout << "\n" << std::string(80, '=') << "\n"
<< "NUMERIC COLUMN ANALYSIS\n"
<< std::string(80, '=') << "\n";
for (const auto& s : num_stats) {
std::cout << "\nColumn: " << s.name << "\n"
<< std::string(50, '-') << "\n";
if (s.count == 0) {
std::cout << " (No valid numeric data)\n";
continue;
}
auto row = [&](const std::string& lbl, const std::string& val) {
std::cout << std::left << std::setw(LW) << (" " + lbl)
<< std::right << std::setw(VW) << val << "\n";
};
row("Count (non-missing):", std::to_string(s.count));
row("Mean:", fmtDbl(s.mean));
row("Median:", fmtDbl(s.median));
row("Std Dev:", fmtDbl(s.std_dev));
row("Variance:", fmtDbl(s.variance));
row("Min:", fmtDbl(s.min_val));
row("Max:", fmtDbl(s.max_val));
row("Q1 (25th pct):", fmtDbl(s.q1));
row("Q2 (50th pct):", fmtDbl(s.q2));
row("Q3 (75th pct):", fmtDbl(s.q3));
std::cout << " Outliers (" << s.outliers.size() << "): ";
if (s.outliers.empty()) {
std::cout << "None";
} else {
int show = std::min(static_cast<int>(s.outliers.size()), 10);
for (int i = 0; i < show; ++i) {
if (i) std::cout << ", ";
std::cout << fmtDbl(s.outliers[i]);
}
if (static_cast<int>(s.outliers.size()) > show)
std::cout << " ... (" << (s.outliers.size() - show) << " more)";
}
std::cout << "\n";
}
}
if (!cat_stats.empty()) {
std::cout << "\n" << std::string(80, '=') << "\n"
<< "CATEGORICAL COLUMN ANALYSIS\n"
<< std::string(80, '=') << "\n";
for (const auto& s : cat_stats) {
std::cout << "\nColumn: " << s.name << "\n"
<< std::string(50, '-') << "\n";
auto row = [&](const std::string& lbl, const std::string& val) {
std::cout << std::left << std::setw(LW) << (" " + lbl)
<< std::right << std::setw(VW) << val << "\n";
};
row("Count (non-missing):", std::to_string(s.count));
row("Unique values:", std::to_string(s.unique_count));
row("Most frequent:", s.most_frequent);
std::cout << " Top frequencies:\n";
for (const auto& [val, cnt] : s.top_frequencies)
std::cout << " " << std::left << std::setw(32) << val
<< std::right << std::setw(8) << cnt << "\n";
}
}
std::cout << "\n" << std::string(80, '=') << "\n"
<< "Analysis complete. Full report saved to report.json\n"
<< std::string(80, '=') << "\n\n";
}
// ---------------------------------------------------------------------------
// JSON output
// ---------------------------------------------------------------------------
static void saveJSON(const std::vector<std::string>& headers,
const std::map<std::string, ColumnType>& col_types,
const std::vector<NumericStats>& num_stats,
const std::vector<CategoricalStats>& cat_stats) {
json report;
report["columns"] = json::array();
for (const auto& h : headers) {
auto it = col_types.find(h);
if (it == col_types.end()) continue;
json col;
col["name"] = h;
if (it->second == ColumnType::NUMERIC) {
col["type"] = "numeric";
for (const auto& ns : num_stats) {
if (ns.name != h) continue;
col["stats"]["count"] = ns.count;
if (ns.count > 0) {
col["stats"]["mean"] = ns.mean;
col["stats"]["median"] = ns.median;
col["stats"]["std_dev"] = ns.std_dev;
col["stats"]["variance"] = ns.variance;
col["stats"]["min"] = ns.min_val;
col["stats"]["max"] = ns.max_val;
col["stats"]["q1"] = ns.q1;
col["stats"]["q2"] = ns.q2;
col["stats"]["q3"] = ns.q3;
}
col["outliers"] = ns.outliers;
break;
}
} else {
col["type"] = "categorical";
for (const auto& cs : cat_stats) {
if (cs.name != h) continue;
col["stats"]["count"] = cs.count;
col["stats"]["unique_count"] = cs.unique_count;
col["stats"]["most_frequent"]= cs.most_frequent;
json freqs = json::array();
for (const auto& [val, cnt] : cs.top_frequencies)
freqs.push_back({{"value", val}, {"count", cnt}});
col["stats"]["top_frequencies"] = freqs;
break;
}
}
report["columns"].push_back(col);
}
std::ofstream out("report.json");
if (!out) {
std::cerr << "Warning: could not write report.json\n";
return;
}
out << report.dump(2) << "\n";
}
// ---------------------------------------------------------------------------
// Sample CSV generator
// ---------------------------------------------------------------------------
static void generateSampleCSV(const std::string& filename) {
std::mt19937 rng(42);
std::normal_distribution<double> height_d(170.0, 10.0);
std::normal_distribution<double> weight_d( 70.0, 12.0);
std::uniform_int_distribution<int> age_d(18, 65);
std::uniform_real_distribution<double>score_d(0.0, 100.0);
std::uniform_real_distribution<double>salary_d(30000.0, 120000.0);
std::uniform_int_distribution<int> miss_d(0, 99);
const std::vector<std::string> depts = {"Engineering","Marketing","Sales","HR","Finance"};
const std::vector<std::string> cities = {"New York","Los Angeles","Chicago","Houston","Phoenix"};
std::uniform_int_distribution<int> dept_d(0, 4);
std::uniform_int_distribution<int> city_d(0, 4);
std::ofstream out(filename);
if (!out) { std::cerr << "Error: cannot create " << filename << "\n"; return; }
out << "age,score,salary,height_cm,weight_kg,department,city\n";
for (int i = 0; i < 210; ++i) {
// --- age ---
std::string age_s;
if (miss_d(rng) < 5) age_s = ""; // ~5 % missing
else if (i == 50) age_s = "200"; // outlier
else if (i == 150) age_s = "-10"; // outlier
else age_s = std::to_string(age_d(rng));
// --- score ---
std::string score_s;
if (miss_d(rng) < 5) {
score_s = "";
} else {
double v = (i == 60) ? 500.0 : score_d(rng);
std::ostringstream oss; oss << std::fixed << std::setprecision(1) << v;
score_s = oss.str();
}
// --- salary ---
std::string salary_s;
if (miss_d(rng) < 5) {
salary_s = "";
} else {
double v = (i == 70) ? 5000000.0 : salary_d(rng);
std::ostringstream oss; oss << std::fixed << std::setprecision(2) << v;
salary_s = oss.str();
}
// --- height (no missing) ---
std::ostringstream h_oss; h_oss << std::fixed << std::setprecision(1) << height_d(rng);
// --- weight (no missing) ---
std::ostringstream w_oss; w_oss << std::fixed << std::setprecision(1) << weight_d(rng);
// city is quoted (contains spaces) to exercise quoted-field parsing
out << age_s << ","
<< score_s << ","
<< salary_s << ","
<< h_oss.str() << ","
<< w_oss.str() << ","
<< depts[dept_d(rng)] << ","
<< "\"" << cities[city_d(rng)] << "\"\n";
}
std::cout << "Generated sample CSV: " << filename << " (210 rows, 5 numeric + 2 categorical cols)\n";
}
// ---------------------------------------------------------------------------
// CSV reader
// ---------------------------------------------------------------------------
static bool readCSV(const std::string& filename,
std::vector<std::string>& headers,
std::map<std::string, std::vector<std::string>>& col_data) {
std::ifstream in(filename);
if (!in) { std::cerr << "Error: cannot open '" << filename << "'\n"; return false; }
std::string line;
// Header row
if (!std::getline(in, line)) { std::cerr << "Error: file is empty.\n"; return false; }
// Strip UTF-8 BOM if present
if (line.size() >= 3 &&
static_cast<unsigned char>(line[0]) == 0xEF &&
static_cast<unsigned char>(line[1]) == 0xBB &&
static_cast<unsigned char>(line[2]) == 0xBF)
line = line.substr(3);
if (trim(line).empty()) { std::cerr << "Error: header line is empty.\n"; return false; }
headers = parseCSVLine(line);
for (auto& h : headers) { h = trim(h); col_data[h] = {}; }
if (headers.empty()) { std::cerr << "Error: no columns in header.\n"; return false; }
// Data rows
while (std::getline(in, line)) {
if (trim(line).empty()) continue;
auto fields = parseCSVLine(line);
fields.resize(headers.size(), ""); // pad short rows
for (size_t i = 0; i < headers.size(); ++i)
col_data[headers[i]].push_back(fields[i]);
}
return true;
}
// ---------------------------------------------------------------------------
// main
// ---------------------------------------------------------------------------
int main(int argc, char* argv[]) {
std::string filename;
if (argc < 2) {
filename = "sample_data.csv";
std::cout << "No input file provided. Generating sample data...\n";
generateSampleCSV(filename);
} else {
filename = argv[1];
}
std::cout << "Analyzing: " << filename << "\n";
std::vector<std::string> headers;
std::map<std::string, std::vector<std::string>> col_data;
if (!readCSV(filename, headers, col_data)) return 1;
// Check for header-only or truly empty file
bool any_data = false;
for (const auto& h : headers)
if (!col_data[h].empty()) { any_data = true; break; }
if (!any_data) {
std::cout << "File contains only a header row — nothing to analyse.\n";
return 0;
}
// Detect column types
std::map<std::string, ColumnType> col_types;
for (const auto& h : headers)
col_types[h] = detectColumnType(col_data[h]);
// Analyse
std::vector<NumericStats> num_stats;
std::vector<CategoricalStats> cat_stats;
for (const auto& h : headers) {
if (col_types[h] == ColumnType::NUMERIC)
num_stats.push_back(analyzeNumericColumn(h, col_data[h]));
else
cat_stats.push_back(analyzeCategoricalColumn(h, col_data[h]));
}
// Output
printReport(num_stats, cat_stats);
saveJSON(headers, col_types, num_stats, cat_stats);
std::cout << "report.json written to current directory.\n";
return 0;
}
README.md
# CSV Statistical Analyzer
A command-line C++ tool that reads a CSV file and performs comprehensive
statistical analysis on every column, producing both a formatted console
summary and a machine-readable `report.json` file.
---
## Features
- **Auto-detects** numeric vs. categorical columns (≥80 % of non-empty values
parseable as a floating-point number → numeric).
- **Numeric statistics** per column: mean, median, standard deviation,
variance, min, max, 25th / 50th / 75th percentiles, non-missing value count.
- **Outlier detection** via the IQR method (Q1 − 1.5·IQR, Q3 + 1.5·IQR).
- **Categorical statistics** per column: unique count, most-frequent value,
top-10 value frequencies.
- Handles real-world messy data: missing values, mixed-type cells, malformed
rows, quoted fields containing commas, UTF-8 BOM, Windows line endings.
- If no file is supplied, generates a 210-row sample CSV (5 numeric +
2 categorical columns, including deliberate outliers and missing values)
and analyses it.
---
## Requirements
| Tool | Minimum version |
|------|----------------|
| G++ (GCC C++ compiler) | 12 |
| CMake | 3.22 |
| Internet access (first build only) | — |
The build system uses CMake's `FetchContent` module to download the single
external dependency automatically on the first configure step.
---
## Dependencies
### Direct
| Library | Version | Purpose |
|---------|---------|---------|
| [nlohmann/json](https://github.com/nlohmann/json) | **v3.11.3** (exact) | Serialise analysis results to `report.json` |
`nlohmann/json` is a header-only library. `FetchContent` downloads the exact
tagged commit (`v3.11.3`) from GitHub on first configure and places the
headers inside the build tree. No system-wide installation is required.
### Transitive
`nlohmann/json` has **no transitive dependencies** — it is a self-contained
single-header library.
Everything else (CSV parsing, statistics, I/O, random number generation) uses
the C++17 standard library, which is provided by G++ itself.
---
## Build instructions
```bash
# 1. Clone / enter the project directory
cd /path/to/csv_analyzer
# 2. Create and enter a build directory (out-of-source build)
cmake -B build -DCMAKE_BUILD_TYPE=Release
# 3. Compile (uses all available CPU cores)
cmake --build build --parallel
# The binary is placed at: build/csv_analyzer
```
On a clean Ubuntu 22.04 machine the only packages you need are:
```bash
sudo apt-get update
sudo apt-get install -y g++ cmake git
```
`git` is needed by `FetchContent` to clone `nlohmann/json` the first time.
---
## Run commands
### Analyse your own CSV file
```bash
./build/csv_analyzer path/to/your/data.csv
```
### Generate and analyse the built-in sample dataset
```bash
./build/csv_analyzer
```
This creates `sample_data.csv` in the current directory and then analyses it.
### Output files
| File | Description |
|------|-------------|
| `report.json` | Full machine-readable analysis (written to the current directory) |
| stdout | Formatted summary table |
---
## Expected output (sample dataset)
```
No input file provided. Generating sample data...
Generated sample CSV: sample_data.csv (210 rows, 5 numeric + 2 categorical cols)
Analyzing: sample_data.csv
================================================================================
NUMERIC COLUMN ANALYSIS
================================================================================
Column: age
--------------------------------------------------
Count (non-missing): 199
Mean: 41.3266
Median: 41.0000
Std Dev: 14.1448
Variance: 200.0763
Min: -10.0000
Max: 200.0000
Q1 (25th pct): 29.0000
Q2 (50th pct): 41.0000
Q3 (75th pct): 54.0000
Outliers (3): -10.0000, 200.0000, ...
Column: score
--------------------------------------------------
...
================================================================================
CATEGORICAL COLUMN ANALYSIS
================================================================================
Column: department
--------------------------------------------------
Count (non-missing): 210
Unique values: 5
Most frequent: Engineering
Top frequencies:
Engineering 46
Sales 44
...
================================================================================
Analysis complete. Full report saved to report.json
================================================================================
report.json written to current directory.
```
The exact numbers will vary slightly based on the random seed baked into the
sample generator.
---
## report.json structure
```json
{
"columns": [
{
"name": "age",
"type": "numeric",
"stats": {
"count": 199,
"mean": 41.33,
"median": 41.0,
"std_dev": 14.14,
"variance": 200.08,
"min": -10.0,
"max": 200.0,
"q1": 29.0,
"q2": 41.0,
"q3": 54.0
},
"outliers": [-10.0, 200.0]
},
{
"name": "department",
"type": "categorical",
"stats": {
"count": 210,
"unique_count": 5,
"most_frequent": "Engineering",
"top_frequencies": [
{"value": "Engineering", "count": 46},
...
]
}
}
]
}
```
---
## Edge-case handling
| Scenario | Behaviour |
|----------|-----------|
| Empty file | Prints error and exits with code 1 |
| Header-only file | Prints notice and exits cleanly |
| Column with all missing values | Reports `count: 0`; all stats shown as `N/A` in console / omitted in JSON |
| Single data row | Computes stats normally; variance = 0; no outliers |
| Quoted fields with commas | Parsed correctly per RFC 4180 |
| Mixed-type column | Classified as categorical when <80 % of values are numeric |
| UTF-8 BOM | Stripped automatically |
| Short rows (fewer fields than headers) | Missing fields treated as empty |