← All tasks
pythonclaude-code/python-t1 #9Not a task: repair changed code

Time Series Trend Detector (python, written by Claude Code)

envgap__claude-code__python-t1-9

Written by a coding agent; not on GitHubWritten 2026-02-27

01 / FAILURE SIGNATURE

As the study recorded it

interpolate method=time requires DatetimeIndex
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/python-t1 #9 · read the task the agent was given
Claude Code wrote this python project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Time Series Trend Detector

Write a program that analyzes time series data to detect trends, seasonal patterns, and anomalies using statistical methods, and produces a visual summary report.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument with columns for timestamp and one or more numeric value columns
- Parse timestamps in multiple formats (ISO 8601, Unix epoch, and common date formats like MM/DD/YYYY, YYYY-MM-DD HH:MM:SS)
- Compute a moving average with a configurable window size via --window flag (default: 7 data points)
- Detect overall trend direction (increasing, decreasing, stable) using linear regression and report the slope and R-squared value
- Detect seasonality by computing autocorrelation at various lags and reporting the dominant period if one exists
- Identify anomalies: data points that deviate more than a configurable number of standard deviations from the moving average (--threshold flag, default: 2.0)
- Support multiple value columns: analyze each independently and report results for all
- Generate a summary report with: trend direction and strength, seasonal period (if any), count and list of anomalies with their timestamps and values, basic statistics (min, max, mean, variance)
- Save the report as a JSON file with --output flag (default: trend_report.json)
- Export the processed data (original values, moving average, anomaly flags) as a CSV file via --export flag
- If no input file is given, generate a sample time series dataset with 365 daily data points containing a linear trend, weekly seasonality, and injected anomalies, then analyze it
- Handle missing timestamps and gaps in the series by interpolating or flagging them

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

detector.py
#!/usr/bin/env python3
"""
Time Series Trend Detector - Trial 1
Uses pandas + statsmodels + matplotlib for time series analysis.
Detects trends, seasonal patterns, and anomalies in time series data.
"""

import sys
import json
import argparse
from datetime import datetime, timedelta
from pathlib import Path

import numpy as np
import pandas as pd
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
from statsmodels.tsa.seasonal import seasonal_decompose
from statsmodels.tsa.stattools import adfuller
from statsmodels.regression.linear_model import OLS
from statsmodels.tools import add_constant


def generate_sample_data(n_points=730, output_path="sample_timeseries.csv"):
    """Generate sample time series with known patterns: trend + seasonality + noise + anomalies."""
    np.random.seed(42)
    dates = pd.date_range(start="2022-01-01", periods=n_points, freq="D")

    # Linear trend
    trend = np.linspace(10, 50, n_points)

    # Seasonal component (yearly cycle)
    seasonal = 15 * np.sin(2 * np.pi * np.arange(n_points) / 365.25)

    # Weekly pattern
    weekly = 5 * np.sin(2 * np.pi * np.arange(n_points) / 7)

    # Noise
    noise = np.random.normal(0, 3, n_points)

    values = trend + seasonal + weekly + noise

    # Inject anomalies
    anomaly_indices = [100, 250, 400, 550, 680]
    for idx in anomaly_indices:
        if idx < n_points:
            values[idx] += np.random.choice([-1, 1]) * np.random.uniform(30, 50)

    # Remove a few data points to simulate missing timestamps
    drop_indices = [50, 51, 200, 201, 202, 450]
    mask = np.ones(n_points, dtype=bool)
    for idx in drop_indices:
        if idx < n_points:
            mask[idx] = False

    df = pd.DataFrame({"timestamp": dates[mask], "value": values[mask]})
    df.to_csv(output_path, index=False)
    print(f"Generated sample data with {len(df)} points -> {output_path}")
    return output_path


def load_data(filepath):
    """Load CSV with timestamp + value columns, handling various formats."""
    df = pd.read_csv(filepath)

    # Identify timestamp and value columns
    ts_col = None
    val_col = None
    for col in df.columns:
        cl = col.strip().lower()
        if cl in ("timestamp", "date", "datetime", "time", "ds"):
            ts_col = col
        elif cl in ("value", "val", "y", "count", "amount", "price", "metric"):
            val_col = col

    if ts_col is None:
        ts_col = df.columns[0]
    if val_col is None:
        val_col = df.columns[1]

    df[ts_col] = pd.to_datetime(df[ts_col], infer_datetime_format=True)
    df[val_col] = pd.to_numeric(df[val_col], errors="coerce")

    df = df[[ts_col, val_col]].rename(columns={ts_col: "timestamp", val_col: "value"})
    df = df.dropna(subset=["value"])
    df = df.sort_values("timestamp").reset_index(drop=True)

    return df


def handle_missing_timestamps(df):
    """Detect and fill missing timestamps using the median frequency."""
    diffs = df["timestamp"].diff().dropna()
    median_freq = diffs.median()

    full_range = pd.date_range(
        start=df["timestamp"].min(),
        end=df["timestamp"].max(),
        freq=median_freq,
    )

    df_full = pd.DataFrame({"timestamp": full_range})
    df_merged = pd.merge(df_full, df, on="timestamp", how="left")

    missing_count = df_merged["value"].isna().sum()
    if missing_count > 0:
        df_merged["value"] = df_merged["value"].interpolate(method="time")
        print(f"Filled {missing_count} missing timestamps via interpolation.")

    return df_merged


def compute_summary_stats(df):
    """Compute summary statistics for the time series."""
    values = df["value"]
    stats = {
        "count": int(len(values)),
        "mean": float(values.mean()),
        "std": float(values.std()),
        "min": float(values.min()),
        "max": float(values.max()),
        "median": float(values.median()),
        "q25": float(values.quantile(0.25)),
        "q75": float(values.quantile(0.75)),
        "skewness": float(values.skew()),
        "kurtosis": float(values.kurtosis()),
        "start_date": str(df["timestamp"].min()),
        "end_date": str(df["timestamp"].max()),
        "duration_days": int((df["timestamp"].max() - df["timestamp"].min()).days),
    }
    return stats


def detect_trend(df):
    """Detect linear and polynomial trends using OLS regression."""
    x = np.arange(len(df)).astype(float)
    y = df["value"].values

    # Linear trend
    X_lin = add_constant(x)
    model_lin = OLS(y, X_lin).fit()
    slope = model_lin.params[1]
    intercept = model_lin.params[0]
    r_squared_lin = model_lin.rsquared

    # Polynomial trend (degree 2)
    X_poly = np.column_stack([np.ones(len(x)), x, x ** 2])
    model_poly = OLS(y, X_poly).fit()
    r_squared_poly = model_poly.rsquared

    # Polynomial trend (degree 3)
    X_poly3 = np.column_stack([np.ones(len(x)), x, x ** 2, x ** 3])
    model_poly3 = OLS(y, X_poly3).fit()
    r_squared_poly3 = model_poly3.rsquared

    # Determine best fit
    best_degree = 1
    best_r2 = r_squared_lin
    if r_squared_poly - r_squared_lin > 0.02:
        best_degree = 2
        best_r2 = r_squared_poly
    if r_squared_poly3 - best_r2 > 0.02:
        best_degree = 3
        best_r2 = r_squared_poly3

    # Trend direction
    if abs(slope) < 1e-6:
        direction = "flat"
    elif slope > 0:
        direction = "upward"
    else:
        direction = "downward"

    # ADF test for stationarity
    adf_result = adfuller(y, maxlag=min(30, len(y) // 2 - 1))
    is_stationary = adf_result[1] < 0.05

    trend_result = {
        "linear_slope": float(slope),
        "linear_intercept": float(intercept),
        "linear_r_squared": float(r_squared_lin),
        "poly2_r_squared": float(r_squared_poly),
        "poly3_r_squared": float(r_squared_poly3),
        "best_polynomial_degree": best_degree,
        "best_r_squared": float(best_r2),
        "direction": direction,
        "slope_per_day": float(slope),
        "adf_statistic": float(adf_result[0]),
        "adf_p_value": float(adf_result[1]),
        "is_stationary": bool(is_stationary),
    }

    return trend_result, model_lin


def decompose_seasonal(df, period=None):
    """Perform seasonal decomposition."""
    if period is None:
        n = len(df)
        if n >= 730:
            period = 365
        elif n >= 60:
            period = 30
        elif n >= 14:
            period = 7
        else:
            period = max(2, n // 3)

    if period >= len(df) // 2:
        period = max(2, len(df) // 3)

    try:
        result = seasonal_decompose(
            df["value"].values,
            model="additive",
            period=period,
            extrapolate_trend="freq",
        )

        seasonal_strength = 1 - (np.var(result.resid[~np.isnan(result.resid)]) /
                                  np.var(result.seasonal[~np.isnan(result.seasonal)] +
                                         result.resid[~np.isnan(result.resid)]))
        seasonal_strength = max(0, min(1, seasonal_strength))

        decomposition = {
            "period": period,
            "seasonal_strength": float(seasonal_strength),
            "trend_component_mean": float(np.nanmean(result.trend)),
            "seasonal_amplitude": float(np.nanmax(result.seasonal) - np.nanmin(result.seasonal)),
            "residual_std": float(np.nanstd(result.resid)),
        }
        return decomposition, result

    except Exception as e:
        print(f"Seasonal decomposition failed: {e}")
        return {"period": period, "error": str(e)}, None


def detect_anomalies_zscore(df, threshold=3.0):
    """Detect anomalies using z-score method."""
    values = df["value"].values
    mean = np.mean(values)
    std = np.std(values)

    if std == 0:
        return []

    z_scores = np.abs((values - mean) / std)
    anomaly_indices = np.where(z_scores > threshold)[0]

    anomalies = []
    for idx in anomaly_indices:
        anomalies.append({
            "index": int(idx),
            "timestamp": str(df["timestamp"].iloc[idx]),
            "value": float(values[idx]),
            "z_score": float(z_scores[idx]),
            "method": "z-score",
        })
    return anomalies


def detect_anomalies_iqr(df, multiplier=1.5):
    """Detect anomalies using IQR method."""
    values = df["value"].values
    q1 = np.percentile(values, 25)
    q3 = np.percentile(values, 75)
    iqr = q3 - q1

    lower_bound = q1 - multiplier * iqr
    upper_bound = q3 + multiplier * iqr

    anomalies = []
    for idx in range(len(values)):
        if values[idx] < lower_bound or values[idx] > upper_bound:
            anomalies.append({
                "index": int(idx),
                "timestamp": str(df["timestamp"].iloc[idx]),
                "value": float(values[idx]),
                "lower_bound": float(lower_bound),
                "upper_bound": float(upper_bound),
                "method": "IQR",
            })
    return anomalies


def compute_moving_averages(df, windows=None):
    """Compute simple and exponential moving averages."""
    if windows is None:
        windows = [7, 14, 30, 90]

    windows = [w for w in windows if w < len(df)]
    results = {}

    for w in windows:
        sma = df["value"].rolling(window=w, min_periods=1).mean()
        ema = df["value"].ewm(span=w, adjust=False).mean()
        results[f"SMA_{w}"] = sma.tolist()
        results[f"EMA_{w}"] = ema.tolist()

    return results, windows


def generate_plots(df, trend_model, decomposition_result, anomalies_zscore, anomalies_iqr, ma_results, windows, output_dir):
    """Generate matplotlib plots for the analysis."""
    fig, axes = plt.subplots(4, 1, figsize=(14, 18))
    fig.suptitle("Time Series Trend Analysis Report", fontsize=16, fontweight="bold")

    # Plot 1: Original data with trend line and moving averages
    ax = axes[0]
    ax.plot(df["timestamp"], df["value"], alpha=0.5, label="Original", linewidth=0.8)
    x = np.arange(len(df))
    trend_line = trend_model.params[0] + trend_model.params[1] * x
    ax.plot(df["timestamp"], trend_line, "r--", label="Linear Trend", linewidth=2)
    if windows:
        w = windows[0]
        key = f"SMA_{w}"
        if key in ma_results:
            ax.plot(df["timestamp"], ma_results[key], label=f"SMA({w})", linewidth=1.5)
    ax.set_title("Time Series with Trend")
    ax.legend()
    ax.xaxis.set_major_formatter(mdates.DateFormatter("%Y-%m"))
    ax.tick_params(axis="x", rotation=45)

    # Plot 2: Seasonal decomposition
    ax = axes[1]
    if decomposition_result is not None:
        ax.plot(df["timestamp"], decomposition_result.seasonal, color="green", linewidth=0.8)
        ax.set_title("Seasonal Component")
    else:
        ax.text(0.5, 0.5, "Decomposition unavailable", ha="center", va="center", transform=ax.transAxes)
    ax.xaxis.set_major_formatter(mdates.DateFormatter("%Y-%m"))
    ax.tick_params(axis="x", rotation=45)

    # Plot 3: Anomalies
    ax = axes[2]
    ax.plot(df["timestamp"], df["value"], alpha=0.5, linewidth=0.8, label="Data")
    if anomalies_zscore:
        anom_idx = [a["index"] for a in anomalies_zscore]
        ax.scatter(
            df["timestamp"].iloc[anom_idx],
            df["value"].iloc[anom_idx],
            color="red", s=60, zorder=5, label="Z-score anomalies",
        )
    if anomalies_iqr:
        anom_idx = [a["index"] for a in anomalies_iqr]
        ax.scatter(
            df["timestamp"].iloc[anom_idx],
            df["value"].iloc[anom_idx],
            color="orange", marker="x", s=60, zorder=5, label="IQR anomalies",
        )
    ax.set_title("Anomaly Detection")
    ax.legend()
    ax.xaxis.set_major_formatter(mdates.DateFormatter("%Y-%m"))
    ax.tick_params(axis="x", rotation=45)

    # Plot 4: Residuals
    ax = axes[3]
    if decomposition_result is not None:
        ax.plot(df["timestamp"], decomposition_result.resid, color="purple", alpha=0.6, linewidth=0.8)
        ax.axhline(y=0, color="black", linestyle="--", linewidth=0.5)
        ax.set_title("Residuals")
    else:
        residuals = df["value"].values - trend_line
        ax.plot(df["timestamp"], residuals, color="purple", alpha=0.6, linewidth=0.8)
        ax.axhline(y=0, color="black", linestyle="--", linewidth=0.5)
        ax.set_title("Residuals (after trend removal)")
    ax.xaxis.set_major_formatter(mdates.DateFormatter("%Y-%m"))
    ax.tick_params(axis="x", rotation=45)

    plt.tight_layout()
    plot_path = Path(output_dir) / "trend_analysis.png"
    plt.savefig(plot_path, dpi=150)
    plt.close()
    print(f"Plot saved to {plot_path}")


def print_console_report(stats, trend, seasonal_info, anomalies_zscore, anomalies_iqr):
    """Print a formatted console report."""
    print("\n" + "=" * 70)
    print("       TIME SERIES TREND DETECTION REPORT")
    print("=" * 70)

    print("\n--- Summary Statistics ---")
    print(f"  Data Points:     {stats['count']}")
    print(f"  Date Range:      {stats['start_date']} to {stats['end_date']}")
    print(f"  Duration:        {stats['duration_days']} days")
    print(f"  Mean:            {stats['mean']:.4f}")
    print(f"  Std Dev:         {stats['std']:.4f}")
    print(f"  Min:             {stats['min']:.4f}")
    print(f"  Max:             {stats['max']:.4f}")
    print(f"  Median:          {stats['median']:.4f}")
    print(f"  Skewness:        {stats['skewness']:.4f}")
    print(f"  Kurtosis:        {stats['kurtosis']:.4f}")

    print("\n--- Trend Analysis ---")
    print(f"  Direction:       {trend['direction']}")
    print(f"  Linear Slope:    {trend['linear_slope']:.6f} per time step")
    print(f"  Linear R^2:      {trend['linear_r_squared']:.4f}")
    print(f"  Poly(2) R^2:     {trend['poly2_r_squared']:.4f}")
    print(f"  Poly(3) R^2:     {trend['poly3_r_squared']:.4f}")
    print(f"  Best Fit Degree: {trend['best_polynomial_degree']}")
    print(f"  Stationary:      {'Yes' if trend['is_stationary'] else 'No'} (ADF p={trend['adf_p_value']:.4f})")

    print("\n--- Seasonal Decomposition ---")
    if "error" not in seasonal_info:
        print(f"  Period:          {seasonal_info['period']}")
        print(f"  Strength:        {seasonal_info['seasonal_strength']:.4f}")
        print(f"  Amplitude:       {seasonal_info['seasonal_amplitude']:.4f}")
        print(f"  Residual Std:    {seasonal_info['residual_std']:.4f}")
    else:
        print(f"  Error: {seasonal_info['error']}")

    print("\n--- Anomaly Detection ---")
    print(f"  Z-score anomalies: {len(anomalies_zscore)}")
    for a in anomalies_zscore[:5]:
        print(f"    {a['timestamp']}  value={a['value']:.2f}  z={a['z_score']:.2f}")
    if len(anomalies_zscore) > 5:
        print(f"    ... and {len(anomalies_zscore) - 5} more")

    print(f"  IQR anomalies:     {len(anomalies_iqr)}")
    for a in anomalies_iqr[:5]:
        print(f"    {a['timestamp']}  value={a['value']:.2f}")
    if len(anomalies_iqr) > 5:
        print(f"    ... and {len(anomalies_iqr) - 5} more")

    print("\n" + "=" * 70)


def save_report(stats, trend, seasonal_info, anomalies_zscore, anomalies_iqr, output_path):
    """Save the full trend report as JSON."""
    report = {
        "generated_at": datetime.now().isoformat(),
        "summary_statistics": stats,
        "trend_analysis": trend,
        "seasonal_decomposition": seasonal_info,
        "anomalies": {
            "z_score": {
                "count": len(anomalies_zscore),
                "threshold": 3.0,
                "detections": anomalies_zscore,
            },
            "iqr": {
                "count": len(anomalies_iqr),
                "multiplier": 1.5,
                "detections": anomalies_iqr,
            },
        },
    }

    with open(output_path, "w") as f:
        json.dump(report, f, indent=2, default=str)
    print(f"Report saved to {output_path}")


def main():
    parser = argparse.ArgumentParser(description="Time Series Trend Detector")
    parser.add_argument("--input", "-i", type=str, help="Path to input CSV file")
    parser.add_argument("--output", "-o", type=str, default=".", help="Output directory")
    parser.add_argument("--period", "-p", type=int, default=None, help="Seasonal period")
    parser.add_argument("--zscore-threshold", type=float, default=3.0, help="Z-score threshold")
    parser.add_argument("--iqr-multiplier", type=float, default=1.5, help="IQR multiplier")
    args = parser.parse_args()

    output_dir = Path(args.output)
    output_dir.mkdir(parents=True, exist_ok=True)

    # Load or generate data
    if args.input:
        csv_path = args.input
    else:
        print("No input file specified. Generating sample data...")
        csv_path = str(output_dir / "sample_timeseries.csv")
        generate_sample_data(output_path=csv_path)

    print(f"Loading data from {csv_path}...")
    df = load_data(csv_path)
    print(f"Loaded {len(df)} records.")

    # Handle missing timestamps
    df = handle_missing_timestamps(df)
    print(f"After filling gaps: {len(df)} records.")

    # Summary stats
    stats = compute_summary_stats(df)

    # Trend detection
    trend_info, trend_model = detect_trend(df)

    # Seasonal decomposition
    seasonal_info, decomp_result = decompose_seasonal(df, period=args.period)

    # Anomaly detection
    anomalies_zscore = detect_anomalies_zscore(df, threshold=args.zscore_threshold)
    anomalies_iqr = detect_anomalies_iqr(df, multiplier=args.iqr_multiplier)

    # Moving averages
    ma_results, windows = compute_moving_averages(df)

    # Console report
    print_console_report(stats, trend_info, seasonal_info, anomalies_zscore, anomalies_iqr)

    # Generate plots
    generate_plots(df, trend_model, decomp_result, anomalies_zscore, anomalies_iqr, ma_results, windows, output_dir)

    # Save JSON report
    report_path = output_dir / "trend_report.json"
    save_report(stats, trend_info, seasonal_info, anomalies_zscore, anomalies_iqr, report_path)

    print("\nAnalysis complete.")


if __name__ == "__main__":
    main()
README.md
# Time Series Trend Detector - Python Trial 1

Analyzes time series data to detect trends, seasonal patterns, and anomalies using pandas, statsmodels, and matplotlib.

## Dependencies

- pandas 2.1.4
- numpy 1.26.2
- statsmodels 0.14.1
- matplotlib 3.8.2

## Usage

```bash
pip install -r requirements.txt

# Run with sample data (auto-generated)
python detector.py

# Run with your CSV file
python detector.py --input data.csv --output results/

# Custom thresholds
python detector.py --input data.csv --zscore-threshold 2.5 --iqr-multiplier 2.0 --period 365
```

## Input Format

CSV file with at least two columns: a timestamp column and a value column. The detector auto-identifies columns named `timestamp`, `date`, `datetime`, `value`, `y`, etc.

## Output

- Console report with summary statistics, trend analysis, seasonal decomposition, and anomaly detection
- `trend_report.json` - Full analysis results in JSON format
- `trend_analysis.png` - Visualization plots
requirements.txt
pandas==2.1.4
numpy==1.26.2
statsmodels==0.14.1
matplotlib==3.8.2