CSV Statistical Analyzer (python, written by Gemini Code Assist)
envgap__gemini__python-t1-1
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
SyntaxError: unterminated string literal at line 50
Not a benchmark task.
- Its repair changed source code, so it is not an environment task.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
gemini/python-t1 #1 · read the task the agent was given
Gemini Code Assist wrote this python project from the task below. It does not run on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: CSV Statistical Analyzer Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument - Auto-detect which columns are numeric vs categorical - For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count - Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column - For each categorical column compute: unique count, most frequent value, and top 10 value frequencies - Print a formatted summary table to the console with aligned columns - Save the complete analysis to report.json including all stats, outlier details, and column type classifications - If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it - Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
csv_analyzer.py
import pandas as pd
import numpy as np
import json
import sys
import os
from typing import Dict, Any
def generate_sample_csv(file_path: str):
print(f"Generating sample CSV: {file_path}")
np.random.seed(42)
rows = 200
data = {
'id': range(1, rows + 1),
'age': np.random.randint(18, 80, size=rows).astype(float),
'salary': np.random.normal(50000, 15000, size=rows),
'score': np.random.uniform(0, 100, size=rows),
'height': np.random.normal(170, 10, size=rows),
'department': np.random.choice(['HR', 'Engineering', 'Sales', 'Marketing', 'Legal'], size=rows),
'city': np.random.choice(['New York', 'London', 'Paris', 'Tokyo', 'Berlin'], size=rows)
}
# Introduce some "messiness"
data['age'][5] = np.nan
data['salary'][10] = 1000000.0 # Outlier
data['salary'][15] = -50000.0 # Outlier
df = pd.DataFrame(data)
df.to_csv(file_path, index=False)
def analyze_csv(file_path: str):
try:
df = pd.read_csv(file_path)
except Exception as e:
print(f"Error reading CSV: {e}")
return
if df.empty:
print("CSV is empty.")
return
report = {
"metadata": {
"file": file_path,
"rows": len(df),
"columns": list(df.columns)
},
"statistics": {}
}
print("
" + "="*80)
print(f"{'COLUMN':<20} | {'TYPE':<10} | {'MEAN/MODE':<15} | {'STD/FREQ':<15}")
print("-"*80)
for col in df.columns:
col_data = df[col]
# Check if numeric
is_numeric = pd.api.types.is_numeric_dtype(col_data) and not col_data.isnull().all()
if is_numeric:
stats = col_data.describe()
q1 = col_data.quantile(0.25)
q3 = col_data.quantile(0.75)
iqr = q3 - q1
outliers = col_data[(col_data < (q1 - 1.5 * iqr)) | (col_data > (q3 + 1.5 * iqr))].dropna().tolist()
report["statistics"][col] = {
"type": "numeric",
"mean": float(stats['mean']),
"median": float(col_data.median()),
"std": float(stats['std']),
"var": float(col_data.var()),
"min": float(stats['min']),
"max": float(stats['max']),
"q1": float(q1),
"q3": float(q3),
"count": int(stats['count']),
"outliers": outliers
}
print(f"{col:<20} | {'Numeric':<10} | {stats['mean']:<15.2f} | {stats['std']:<15.2f}")
else:
counts = col_data.value_counts()
mode = str(counts.idxmax()) if not counts.empty else "N/A"
freq = int(counts.max()) if not counts.empty else 0
report["statistics"][col] = {
"type": "categorical",
"unique_count": int(col_data.nunique()),
"mode": mode,
"top_10": counts.head(10).to_dict()
}
print(f"{col:<20} | {'Categorical':<10} | {mode:<15} | {freq:<15}")
with open('report.json', 'w') as f:
json.dump(report, f, indent=4)
print("="*80)
print("Report saved to report.json")
if __name__ == "__main__":
input_file = sys.argv[1] if len(sys.argv) > 1 else "sample.csv"
if not os.path.exists(input_file):
generate_sample_csv(input_file)
analyze_csv(input_file)
README.md
# CSV Statistical Analyzer (Python) A comprehensive tool for statistical analysis of CSV data. ## Setup Instructions 1. Ensure Python 3.10+ is installed. 2. Create a virtual environment: ```bash python3 -m venv venv source venv/bin/activate ``` 3. Install dependencies: ```bash pip install -r requirements.txt ``` ## Run Commands - Run with a specific CSV: ```bash python csv_analyzer.py data.csv ``` - Run with generated sample data: ```bash python csv_analyzer.py ``` ## Output - Console: Formatted summary table. - File: `report.json` containing detailed statistics and outliers.
requirements.txt
numpy==1.26.4 pandas==2.2.1 python-dateutil==2.8.2 pytz==2024.1 six==1.16.0