← All tasks
pythongemini/python-t1 #1Not a task: repair changed code

CSV Statistical Analyzer (python, written by Gemini Code Assist)

envgap__gemini__python-t1-1

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

SyntaxError: unterminated string literal at line 50
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

gemini/python-t1 #1 · read the task the agent was given
Gemini Code Assist wrote this python project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: CSV Statistical Analyzer

Write a program that reads a CSV file and performs comprehensive statistical analysis on every numeric column. It should handle real-world messy data — missing values, mixed types, malformed rows — and produce both a human-readable console report and a machine-readable JSON output.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument
- Auto-detect which columns are numeric vs categorical
- For each numeric column compute: mean, median, standard deviation, variance, min, max, 25th/50th/75th percentiles, and non-missing value count
- Detect outliers using the IQR method (values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR) and list them per column
- For each categorical column compute: unique count, most frequent value, and top 10 value frequencies
- Print a formatted summary table to the console with aligned columns
- Save the complete analysis to report.json including all stats, outlier details, and column type classifications
- If no input file is given, generate a sample CSV with at least 200 rows across 5 numeric and 2 categorical columns, then analyze it
- Handle gracefully: empty files, header-only files, columns with all missing values, single-row files, quoted fields containing commas

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

csv_analyzer.py
import pandas as pd
import numpy as np
import json
import sys
import os
from typing import Dict, Any

def generate_sample_csv(file_path: str):
    print(f"Generating sample CSV: {file_path}")
    np.random.seed(42)
    rows = 200
    data = {
        'id': range(1, rows + 1),
        'age': np.random.randint(18, 80, size=rows).astype(float),
        'salary': np.random.normal(50000, 15000, size=rows),
        'score': np.random.uniform(0, 100, size=rows),
        'height': np.random.normal(170, 10, size=rows),
        'department': np.random.choice(['HR', 'Engineering', 'Sales', 'Marketing', 'Legal'], size=rows),
        'city': np.random.choice(['New York', 'London', 'Paris', 'Tokyo', 'Berlin'], size=rows)
    }
    
    # Introduce some "messiness"
    data['age'][5] = np.nan
    data['salary'][10] = 1000000.0  # Outlier
    data['salary'][15] = -50000.0    # Outlier
    
    df = pd.DataFrame(data)
    df.to_csv(file_path, index=False)

def analyze_csv(file_path: str):
    try:
        df = pd.read_csv(file_path)
    except Exception as e:
        print(f"Error reading CSV: {e}")
        return

    if df.empty:
        print("CSV is empty.")
        return

    report = {
        "metadata": {
            "file": file_path,
            "rows": len(df),
            "columns": list(df.columns)
        },
        "statistics": {}
    }

    print("
" + "="*80)
    print(f"{'COLUMN':<20} | {'TYPE':<10} | {'MEAN/MODE':<15} | {'STD/FREQ':<15}")
    print("-"*80)

    for col in df.columns:
        col_data = df[col]
        
        # Check if numeric
        is_numeric = pd.api.types.is_numeric_dtype(col_data) and not col_data.isnull().all()
        
        if is_numeric:
            stats = col_data.describe()
            q1 = col_data.quantile(0.25)
            q3 = col_data.quantile(0.75)
            iqr = q3 - q1
            outliers = col_data[(col_data < (q1 - 1.5 * iqr)) | (col_data > (q3 + 1.5 * iqr))].dropna().tolist()
            
            report["statistics"][col] = {
                "type": "numeric",
                "mean": float(stats['mean']),
                "median": float(col_data.median()),
                "std": float(stats['std']),
                "var": float(col_data.var()),
                "min": float(stats['min']),
                "max": float(stats['max']),
                "q1": float(q1),
                "q3": float(q3),
                "count": int(stats['count']),
                "outliers": outliers
            }
            print(f"{col:<20} | {'Numeric':<10} | {stats['mean']:<15.2f} | {stats['std']:<15.2f}")
        else:
            counts = col_data.value_counts()
            mode = str(counts.idxmax()) if not counts.empty else "N/A"
            freq = int(counts.max()) if not counts.empty else 0
            
            report["statistics"][col] = {
                "type": "categorical",
                "unique_count": int(col_data.nunique()),
                "mode": mode,
                "top_10": counts.head(10).to_dict()
            }
            print(f"{col:<20} | {'Categorical':<10} | {mode:<15} | {freq:<15}")

    with open('report.json', 'w') as f:
        json.dump(report, f, indent=4)
    print("="*80)
    print("Report saved to report.json")

if __name__ == "__main__":
    input_file = sys.argv[1] if len(sys.argv) > 1 else "sample.csv"
    if not os.path.exists(input_file):
        generate_sample_csv(input_file)
    analyze_csv(input_file)
README.md
# CSV Statistical Analyzer (Python)

A comprehensive tool for statistical analysis of CSV data.

## Setup Instructions

1. Ensure Python 3.10+ is installed.
2. Create a virtual environment:
   ```bash
   python3 -m venv venv
   source venv/bin/activate
   ```
3. Install dependencies:
   ```bash
   pip install -r requirements.txt
   ```

## Run Commands

- Run with a specific CSV:
  ```bash
  python csv_analyzer.py data.csv
  ```
- Run with generated sample data:
  ```bash
  python csv_analyzer.py
  ```

## Output
- Console: Formatted summary table.
- File: `report.json` containing detailed statistics and outliers.
requirements.txt
numpy==1.26.4
pandas==2.2.1
python-dateutil==2.8.2
pytz==2024.1
six==1.16.0