Time Series Trend Detector (python, written by Gemini Code Assist)
envgap__gemini__python-t1-9
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
SyntaxError: unterminated string literal at line 113
Not a benchmark task.
- Its repair changed source code, so it is not an environment task.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
gemini/python-t1 #9 · read the task the agent was given
Gemini Code Assist wrote this python project from the task below. It does not run on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Time Series Trend Detector Write a program that analyzes time series data to detect trends, seasonal patterns, and anomalies using statistical methods, and produces a visual summary report. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument with columns for timestamp and one or more numeric value columns - Parse timestamps in multiple formats (ISO 8601, Unix epoch, and common date formats like MM/DD/YYYY, YYYY-MM-DD HH:MM:SS) - Compute a moving average with a configurable window size via --window flag (default: 7 data points) - Detect overall trend direction (increasing, decreasing, stable) using linear regression and report the slope and R-squared value - Detect seasonality by computing autocorrelation at various lags and reporting the dominant period if one exists - Identify anomalies: data points that deviate more than a configurable number of standard deviations from the moving average (--threshold flag, default: 2.0) - Support multiple value columns: analyze each independently and report results for all - Generate a summary report with: trend direction and strength, seasonal period (if any), count and list of anomalies with their timestamps and values, basic statistics (min, max, mean, variance) - Save the report as a JSON file with --output flag (default: trend_report.json) - Export the processed data (original values, moving average, anomaly flags) as a CSV file via --export flag - If no input file is given, generate a sample time series dataset with 365 daily data points containing a linear trend, weekly seasonality, and injected anomalies, then analyze it - Handle missing timestamps and gaps in the series by interpolating or flagging them Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# Time Series Trend Detector (Python) A tool for detecting trends, seasonality, and anomalies in time series data. ## Setup Instructions 1. Ensure Python 3.10+ is installed. 2. Create a virtual environment: ```bash python3 -m venv venv source venv/bin/activate ``` 3. Install dependencies: ```bash pip install -r requirements.txt ``` ## Run Commands - Analyze a specific CSV: ```bash python trend_detector.py data.csv ``` - Configure window size and anomaly threshold: ```bash python trend_detector.py data.csv --window 14 --threshold 3.0 ``` - Run with generated sample data: ```bash python trend_detector.py ``` ## Output - Console: Summary of trend and anomaly counts. - File: `trend_report.json` with detailed statistical findings.
requirements.txt
pandas==2.2.1 numpy==1.26.4 scipy==1.12.0 statsmodels==0.14.1 python-dateutil==2.8.2
trend_detector.py
import pandas as pd
import numpy as np
import json
import sys
import argparse
import os
from scipy import stats
from statsmodels.tsa.stattools import acf
from typing import Dict, Any
def generate_sample_data(path: str):
print(f"Generating sample time series data: {path}")
np.random.seed(42)
rows = 365
date_rng = pd.date_range(start='2025-01-01', periods=rows, freq='D')
# Trend: 0.1 * index
trend = 0.1 * np.arange(rows)
# Seasonality: Weekly (7 days)
seasonality = 5 * np.sin(2 * np.pi * np.arange(rows) / 7)
# Noise
noise = np.random.normal(0, 1, size=rows)
values = trend + seasonality + noise
# Inject anomalies
values[100] += 20.0
values[200] -= 15.0
df = pd.DataFrame({'timestamp': date_rng, 'value': values})
df.to_csv(path, index=False)
def analyze_series(df: pd.DataFrame, col: str, window: int, threshold: float) -> Dict:
series = df[col]
# Moving Average
ma = series.rolling(window=window, center=True).mean()
# Trend detection (Linear Regression)
x = np.arange(len(series))
slope, intercept, r_value, p_value, std_err = stats.linregress(x, series)
trend_dir = "increasing" if slope > 0.01 else "decreasing" if slope < -0.01 else "stable"
# Seasonality (Autocorrelation)
# Check lags up to 30
acf_vals = acf(series, nlags=30, fft=True)
dominant_period = int(np.argmax(acf_vals[1:]) + 1) if acf_vals[1:].max() > 0.5 else None
# Anomalies
ma_filled = ma.fillna(method='bfill').fillna(method='ffill')
diff = series - ma_filled
std = diff.std()
anomaly_indices = np.where(np.abs(diff) > threshold * std)[0]
anomalies = df.iloc[anomaly_indices][['timestamp', col]].to_dict(orient='records')
return {
"statistics": {
"min": float(series.min()),
"max": float(series.max()),
"mean": float(series.mean()),
"std": float(series.std())
},
"trend": {
"direction": trend_dir,
"slope": float(slope),
"r_squared": float(r_value**2)
},
"seasonality": {
"dominant_period": dominant_period
},
"anomalies": {
"count": len(anomalies),
"list": anomalies
}
}
def main():
parser = argparse.ArgumentParser(description="Time Series Trend Detector")
parser.add_argument("input", nargs="?")
parser.add_argument("--window", type=int, default=7)
parser.add_argument("--threshold", type=float, default=2.0)
parser.add_argument("--output", default="trend_report.json")
parser.add_argument("--export", default="processed_data.csv")
args = parser.parse_args()
target = args.input or "sample_ts.csv"
if not os.path.exists(target):
generate_sample_data(target)
df = pd.read_csv(target)
df['timestamp'] = pd.to_datetime(df['timestamp'])
# Find numeric columns excluding timestamp
numeric_cols = df.select_dtypes(include=[np.number]).columns.tolist()
if 'timestamp' in numeric_cols: numeric_cols.remove('timestamp')
report = {"metadata": {"file": target, "rows": len(df)}, "analysis": {}}
for col in numeric_cols:
print(f"Analyzing column: {col}")
col_analysis = analyze_series(df, col, args.window, args.threshold)
report["analysis"][col] = col_analysis
# Display summary
t = col_analysis['trend']
print(f" Trend: {t['direction']} (R2={t['r_squared']:.3f})")
print(f" Anomalies detected: {col_analysis['anomalies']['count']}")
with open(args.output, 'w') as f:
json.dump(report, f, indent=4)
print(f"
Analysis complete. Report: {args.output}")
if __name__ == "__main__":
main()