Data Profiling Tool (python, written by Gemini Code Assist)
envgap__gemini__python-t1-7
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
SyntaxError: unterminated string literal at line 47
Not a benchmark task.
- Its repair changed source code, so it is not an environment task.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
gemini/python-t1 #7 · read the task the agent was given
Gemini Code Assist wrote this python project from the task below. It does not run on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Data Profiling Tool Write a program that profiles tabular datasets by analyzing column types, distributions, missing values, correlations, and data quality issues, producing a comprehensive data quality report. FUNCTIONAL REQUIREMENTS: - Accept a CSV or JSON data file path as a command-line argument - Auto-detect column data types: numeric (integer vs float), string, boolean, date/timestamp, and categorical (low cardinality strings) - For numeric columns: compute min, max, mean, median, standard deviation, skewness, and percentiles (25th, 50th, 75th, 95th, 99th) - For string columns: compute min/max/average length, most common values (top 10), and unique count - For all columns: count total values, missing/null values, missing percentage, and unique value count - Detect potential data quality issues: columns that are entirely null, columns with a single unique value, numeric columns with extreme outliers (beyond 4 standard deviations), and string columns that look like misclassified numbers or dates - Compute a pairwise correlation matrix for all numeric columns - Print a formatted summary report to console showing key statistics per column - Save the full profiling report as a JSON file with --output flag (default: data_profile.json) - If no input file is given, generate a sample dataset with 1000 rows across at least 8 columns of mixed types including intentional quality issues, then profile it - Handle files with inconsistent delimiters or encoding issues gracefully Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
data_profiler.py
import pandas as pd
import numpy as np
from scipy.stats import skew
import json
import sys
import argparse
import os
from typing import Dict, Any
def generate_sample_data(path: str):
print(f"Generating sample data: {path}")
np.random.seed(42)
rows = 1000
data = {
'id': range(rows),
'age': np.random.randint(18, 90, size=rows).astype(float),
'salary': np.random.normal(60000, 20000, size=rows),
'joined_date': pd.date_range(start='2020-01-01', periods=rows, freq='D'),
'department': np.random.choice(['HR', 'Tech', 'Sales', 'Admin'], size=rows),
'is_active': np.random.choice([True, False], size=rows),
'notes': ['Note ' + str(i) for i in range(rows)],
'empty_col': [None] * rows, # Quality issue: entirely null
'constant_col': ['Fixed'] * rows # Quality issue: single unique value
}
# Introduce outliers
data['salary'][0] = 1000000.0
# Introduce missing values
data['age'][5:15] = np.nan
df = pd.DataFrame(data)
df.to_csv(path, index=False)
def profile_data(input_path: str, output_path: str):
try:
df = pd.read_csv(input_path)
except Exception as e:
print(f"Error loading data: {e}")
return
profile = {
"metadata": {"rows": len(df), "columns": len(df.columns)},
"columns": {},
"quality_issues": [],
"correlations": {}
}
print("
DATA PROFILING SUMMARY")
print("=" * 60)
numeric_cols = []
for col in df.columns:
col_data = df[col]
missing_count = int(col_data.isnull().sum())
unique_count = int(col_data.nunique())
col_profile = {
"missing_count": missing_count,
"missing_percent": (missing_count / len(df)) * 100,
"unique_count": unique_count
}
# Quality Issue Checks
if missing_count == len(df):
profile["quality_issues"].append(f"Column '{col}' is entirely null.")
if unique_count == 1:
profile["quality_issues"].append(f"Column '{col}' has only one unique value.")
# Type Detection
if pd.api.types.is_numeric_dtype(col_data):
numeric_cols.append(col)
valid_data = col_data.dropna()
col_profile.update({
"type": "numeric",
"min": float(valid_data.min()) if not valid_data.empty else None,
"max": float(valid_data.max()) if not valid_data.empty else None,
"mean": float(valid_data.mean()) if not valid_data.empty else None,
"median": float(valid_data.median()) if not valid_data.empty else None,
"std": float(valid_data.std()) if not valid_data.empty else None,
"skewness": float(skew(valid_data)) if len(valid_data) > 2 else 0
})
# Outlier detection (4 std devs)
if not valid_data.empty:
m, s = valid_data.mean(), valid_data.std()
outliers = valid_data[(valid_data > m + 4*s) | (valid_data < m - 4*s)]
if not outliers.empty:
profile["quality_issues"].append(f"Column '{col}' has {len(outliers)} extreme outliers (>4 STD).")
elif pd.api.types.is_string_dtype(col_data):
col_profile.update({
"type": "string",
"avg_length": float(col_data.str.len().mean()) if missing_count < len(df) else 0,
"top_10": col_data.value_counts().head(10).to_dict()
})
profile["columns"][col] = col_profile
print(f"{col:<15} | {col_profile.get('type', 'other'):<10} | Missing: {col_profile['missing_percent']:>5.1f}% | Unique: {unique_count}")
# Correlation
if numeric_cols:
corr_matrix = df[numeric_cols].corr().to_dict()
profile["correlations"] = corr_matrix
with open(output_path, 'w') as f:
json.dump(profile, f, indent=4)
print("=" * 60)
print(f"Profiling complete. Report saved to {output_path}")
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("input", nargs="?")
parser.add_argument("--output", default="data_profile.json")
args = parser.parse_args()
target = args.input or "sample_data.csv"
if not os.path.exists(target):
generate_sample_data(target)
profile_data(target, args.output)
README.md
# Data Profiling Tool (Python) A comprehensive tool for analyzing tabular datasets and identifying quality issues. ## Setup Instructions 1. Ensure Python 3.10+ is installed. 2. Create a virtual environment: ```bash python3 -m venv venv source venv/bin/activate ``` 3. Install dependencies: ```bash pip install -r requirements.txt ``` ## Run Commands - Profile a specific CSV: ```bash python data_profiler.py data.csv ``` - Run with generated sample data: ```bash python data_profiler.py ``` ## Output - Console: Summary of column types and missing data. - File: `data_profile.json` containing statistics, correlation matrix, and quality alerts.
requirements.txt
pandas==2.2.1 numpy==1.26.4 scipy==1.12.0 tabulate==0.9.0