← All tasks
pythongemini/python-t1 #7Not a task: repair changed code

Data Profiling Tool (python, written by Gemini Code Assist)

envgap__gemini__python-t1-7

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

SyntaxError: unterminated string literal at line 47
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

gemini/python-t1 #7 · read the task the agent was given
Gemini Code Assist wrote this python project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Data Profiling Tool

Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV or JSON data file path as a command-line argument
- Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings)
- For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th)
- For string columns: compute min/max/average length, most common values (top 10), and unique count
- For all columns: count total values, missing/null values, missing percentage, and unique value count
- Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates
- Compute a pairwise correlation matrix for all numeric columns
- Print a formatted summary report to console showing key statistics per column
- Save the full profiling report as a JSON file with --output flag (default: data_profile.json)
- If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it
- Handle files with inconsistent delimiters or encoding issues gracefully

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

data_profiler.py
import pandas as pd
import numpy as np
from scipy.stats import skew
import json
import sys
import argparse
import os
from typing import Dict, Any

def generate_sample_data(path: str):
    print(f"Generating sample data: {path}")
    np.random.seed(42)
    rows = 1000
    data = {
        'id': range(rows),
        'age': np.random.randint(18, 90, size=rows).astype(float),
        'salary': np.random.normal(60000, 20000, size=rows),
        'joined_date': pd.date_range(start='2020-01-01', periods=rows, freq='D'),
        'department': np.random.choice(['HR', 'Tech', 'Sales', 'Admin'], size=rows),
        'is_active': np.random.choice([True, False], size=rows),
        'notes': ['Note ' + str(i) for i in range(rows)],
        'empty_col': [None] * rows,  # Quality issue: entirely null
        'constant_col': ['Fixed'] * rows  # Quality issue: single unique value
    }
    # Introduce outliers
    data['salary'][0] = 1000000.0
    # Introduce missing values
    data['age'][5:15] = np.nan
    
    df = pd.DataFrame(data)
    df.to_csv(path, index=False)

def profile_data(input_path: str, output_path: str):
    try:
        df = pd.read_csv(input_path)
    except Exception as e:
        print(f"Error loading data: {e}")
        return

    profile = {
        "metadata": {"rows": len(df), "columns": len(df.columns)},
        "columns": {},
        "quality_issues": [],
        "correlations": {}
    }

    print("
DATA PROFILING SUMMARY")
    print("=" * 60)

    numeric_cols = []

    for col in df.columns:
        col_data = df[col]
        missing_count = int(col_data.isnull().sum())
        unique_count = int(col_data.nunique())
        
        col_profile = {
            "missing_count": missing_count,
            "missing_percent": (missing_count / len(df)) * 100,
            "unique_count": unique_count
        }

        # Quality Issue Checks
        if missing_count == len(df):
            profile["quality_issues"].append(f"Column '{col}' is entirely null.")
        if unique_count == 1:
            profile["quality_issues"].append(f"Column '{col}' has only one unique value.")

        # Type Detection
        if pd.api.types.is_numeric_dtype(col_data):
            numeric_cols.append(col)
            valid_data = col_data.dropna()
            col_profile.update({
                "type": "numeric",
                "min": float(valid_data.min()) if not valid_data.empty else None,
                "max": float(valid_data.max()) if not valid_data.empty else None,
                "mean": float(valid_data.mean()) if not valid_data.empty else None,
                "median": float(valid_data.median()) if not valid_data.empty else None,
                "std": float(valid_data.std()) if not valid_data.empty else None,
                "skewness": float(skew(valid_data)) if len(valid_data) > 2 else 0
            })
            # Outlier detection (4 std devs)
            if not valid_data.empty:
                m, s = valid_data.mean(), valid_data.std()
                outliers = valid_data[(valid_data > m + 4*s) | (valid_data < m - 4*s)]
                if not outliers.empty:
                    profile["quality_issues"].append(f"Column '{col}' has {len(outliers)} extreme outliers (>4 STD).")

        elif pd.api.types.is_string_dtype(col_data):
            col_profile.update({
                "type": "string",
                "avg_length": float(col_data.str.len().mean()) if missing_count < len(df) else 0,
                "top_10": col_data.value_counts().head(10).to_dict()
            })
        
        profile["columns"][col] = col_profile
        print(f"{col:<15} | {col_profile.get('type', 'other'):<10} | Missing: {col_profile['missing_percent']:>5.1f}% | Unique: {unique_count}")

    # Correlation
    if numeric_cols:
        corr_matrix = df[numeric_cols].corr().to_dict()
        profile["correlations"] = corr_matrix

    with open(output_path, 'w') as f:
        json.dump(profile, f, indent=4)
    print("=" * 60)
    print(f"Profiling complete. Report saved to {output_path}")

if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("input", nargs="?")
    parser.add_argument("--output", default="data_profile.json")
    args = parser.parse_args()

    target = args.input or "sample_data.csv"
    if not os.path.exists(target):
        generate_sample_data(target)
    
    profile_data(target, args.output)
README.md
# Data Profiling Tool (Python)

A comprehensive tool for analyzing tabular datasets and identifying quality issues.

## Setup Instructions

1. Ensure Python 3.10+ is installed.
2. Create a virtual environment:
   ```bash
   python3 -m venv venv
   source venv/bin/activate
   ```
3. Install dependencies:
   ```bash
   pip install -r requirements.txt
   ```

## Run Commands

- Profile a specific CSV:
  ```bash
  python data_profiler.py data.csv
  ```
- Run with generated sample data:
  ```bash
  python data_profiler.py
  ```

## Output
- Console: Summary of column types and missing data.
- File: `data_profile.json` containing statistics, correlation matrix, and quality alerts.
requirements.txt
pandas==2.2.1
numpy==1.26.4
scipy==1.12.0
tabulate==0.9.0