← All tasks
pythoncodex/python-t1 #9Not a task: already works

Time Series Trend Detector (python, written by Codex)

envgap__codex__python-t1-9

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/python-t1 #9 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Time Series Trend Detector

Write a program that analyzes time series data to detect trends, seasonal patterns, and anomalies using statistical methods, and produces a visual summary report.

FUNCTIONAL REQUIREMENTS:
- Accept a CSV file path as a command-line argument with columns for timestamp and one or more numeric value columns
- Parse timestamps in multiple formats (ISO 8601, Unix epoch, and common date formats like MM/DD/YYYY, YYYY-MM-DD HH:MM:SS)
- Compute a moving average with a configurable window size via --window flag (default: 7 data points)
- Detect overall trend direction (increasing, decreasing, stable) using linear regression and report the slope and R-squared value
- Detect seasonality by computing autocorrelation at various lags and reporting the dominant period if one exists
- Identify anomalies: data points that deviate more than a configurable number of standard deviations from the moving average (--threshold flag, default: 2.0)
- Support multiple value columns: analyze each independently and report results for all
- Generate a summary report with: trend direction and strength, seasonal period (if any), count and list of anomalies with their timestamps and values, basic statistics (min, max, mean, variance)
- Save the report as a JSON file with --output flag (default: trend_report.json)
- Export the processed data (original values, moving average, anomaly flags) as a CSV file via --export flag
- If no input file is given, generate a sample time series dataset with 365 daily data points containing a linear trend, weekly seasonality, and injected anomalies, then analyze it
- Handle missing timestamps and gaps in the series by interpolating or flagging them

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

README.md
# Time Series Trend Detector (Python)

Analyzes CSV time series data for trend, seasonality, anomalies, and summary statistics.

## Requirements

- Ubuntu 22.04
- Python 3.10+

## Dependencies

- Direct: none
- Transitive: none

`requirements.txt` is pinned and intentionally dependency-free (standard library implementation).

## Run

With input:

```bash
python src/main.py /path/to/data.csv --window 7 --threshold 2.0 --output trend_report.json --export processed.csv
```

No input (generates a 365-day sample and analyzes it):

```bash
python src/main.py
```

## Features

- Timestamp parsing: ISO 8601, Unix epoch, `MM/DD/YYYY`, `YYYY-MM-DD HH:MM:SS`
- Moving average (`--window`)
- Trend detection with linear regression slope + R-squared
- Seasonality detection via autocorrelation lag scan
- Anomaly detection against moving average (`--threshold` std dev)
- Gap interpolation and reporting
- JSON report and processed CSV export
requirements.txt
# No external dependencies required.
# Direct dependencies: none
# Transitive dependencies: none
src/main.py
#!/usr/bin/env python3
import argparse
import csv
import json
import math
from datetime import datetime, timezone, timedelta
from pathlib import Path
from typing import Any


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description="Time Series Trend Detector")
    parser.add_argument("input_file", nargs="?", help="CSV input file")
    parser.add_argument("--window", type=int, default=7, help="Moving average window size")
    parser.add_argument("--threshold", type=float, default=2.0, help="Std-dev threshold for anomalies")
    parser.add_argument("--output", default="trend_report.json", help="JSON report output path")
    parser.add_argument("--export", help="Processed CSV export path")
    return parser.parse_args()


def parse_timestamp(raw: Any) -> float | None:
    if raw is None:
        return None
    s = str(raw).strip()
    if not s:
        return None
    if s.isdigit() and len(s) == 10:
        return float(int(s))
    if s.isdigit() and len(s) == 13:
        return float(int(s) / 1000)
    for fmt in ("%m/%d/%Y", "%Y-%m-%d %H:%M:%S", "%Y-%m-%d", "%Y-%m-%dT%H:%M:%S", "%Y-%m-%dT%H:%M:%SZ"):
        try:
            dt = datetime.strptime(s, fmt).replace(tzinfo=timezone.utc)
            return dt.timestamp()
        except ValueError:
            continue
    try:
        return datetime.fromisoformat(s.replace("Z", "+00:00")).timestamp()
    except ValueError:
        return None


def mean(values: list[float]) -> float | None:
    return sum(values) / len(values) if values else None


def variance(values: list[float], m: float | None = None) -> float | None:
    if not values:
        return None
    mu = mean(values) if m is None else m
    assert mu is not None
    return sum((x - mu) ** 2 for x in values) / len(values)


def stddev(values: list[float], m: float | None = None) -> float | None:
    v = variance(values, m)
    return math.sqrt(v) if v is not None else None


def moving_average(values: list[float | None], window: int) -> list[float | None]:
    out: list[float | None] = [None] * len(values)
    for i in range(len(values)):
        start = max(0, i - window + 1)
        chunk = [x for x in values[start : i + 1] if x is not None]
        out[i] = sum(chunk) / len(chunk) if chunk else None
    return out


def linear_regression(values: list[float | None]) -> dict[str, float]:
    points = [(i, v) for i, v in enumerate(values) if v is not None]
    if len(points) < 2:
        return {"slope": 0.0, "intercept": 0.0, "r_squared": 0.0}
    xs = [float(x) for x, _ in points]
    ys = [float(y) for _, y in points]
    mx = sum(xs) / len(xs)
    my = sum(ys) / len(ys)
    cov = sum((x - mx) * (y - my) for x, y in points)
    vx = sum((x - mx) ** 2 for x in xs)
    slope = cov / vx if vx != 0 else 0.0
    intercept = my - slope * mx
    ss_res = sum((y - (slope * x + intercept)) ** 2 for x, y in points)
    ss_tot = sum((y - my) ** 2 for y in ys)
    r2 = 0.0 if ss_tot == 0 else 1.0 - ss_res / ss_tot
    return {"slope": slope, "intercept": intercept, "r_squared": r2}


def autocorrelation(values: list[float | None], lag: int) -> float | None:
    pairs = [(values[i], values[i - lag]) for i in range(lag, len(values)) if values[i] is not None and values[i - lag] is not None]
    if len(pairs) < 3:
        return None
    xs = [x for x, _ in pairs]
    ys = [y for _, y in pairs]
    mx = sum(xs) / len(xs)
    my = sum(ys) / len(ys)
    num = sum((x - mx) * (y - my) for x, y in pairs)
    dx = sum((x - mx) ** 2 for x in xs)
    dy = sum((y - my) ** 2 for y in ys)
    den = math.sqrt(dx * dy)
    return 0.0 if den == 0 else num / den


def detect_seasonality(values: list[float | None]) -> dict[str, float | int | None]:
    max_lag = min(60, len(values) // 2)
    best_lag: int | None = None
    best_corr = 0.0
    for lag in range(2, max_lag + 1):
        corr = autocorrelation(values, lag)
        if corr is None:
            continue
        if abs(corr) > abs(best_corr):
            best_corr = corr
            best_lag = lag
    if best_lag is None or abs(best_corr) < 0.3:
        return {"period": None, "autocorrelation": None}
    return {"period": best_lag, "autocorrelation": best_corr}


def infer_step(sorted_ts: list[float]) -> float:
    diffs = [sorted_ts[i] - sorted_ts[i - 1] for i in range(1, len(sorted_ts)) if sorted_ts[i] > sorted_ts[i - 1]]
    if not diffs:
        return 86400.0
    diffs.sort()
    return diffs[len(diffs) // 2]


def interpolate_gaps(rows: list[dict[str, Any]], value_cols: list[str]) -> tuple[list[dict[str, Any]], float, list[float]]:
    ordered = sorted(rows, key=lambda r: r["timestamp_ts"])
    ts_list = [r["timestamp_ts"] for r in ordered]
    step = infer_step(ts_list)
    by_ts = {r["timestamp_ts"]: r for r in ordered}
    start = ts_list[0]
    end = ts_list[-1]
    t = start
    filled: list[dict[str, Any]] = []
    gaps: list[float] = []
    while t <= end + 1e-9:
        if t in by_ts:
            item = dict(by_ts[t])
            item["gap_filled"] = False
            filled.append(item)
        else:
            item = {
                "timestamp_ts": t,
                "timestamp": datetime.fromtimestamp(t, tz=timezone.utc).isoformat().replace("+00:00", "Z"),
                "gap_filled": True,
            }
            for c in value_cols:
                item[c] = None
            filled.append(item)
            gaps.append(t)
        t += step

    for col in value_cols:
        for i in range(len(filled)):
            if filled[i][col] is not None:
                continue
            left = i - 1
            while left >= 0 and filled[left][col] is None:
                left -= 1
            right = i + 1
            while right < len(filled) and filled[right][col] is None:
                right += 1
            if left >= 0 and right < len(filled):
                ratio = (filled[i]["timestamp_ts"] - filled[left]["timestamp_ts"]) / (
                    filled[right]["timestamp_ts"] - filled[left]["timestamp_ts"]
                )
                filled[i][col] = filled[left][col] + (filled[right][col] - filled[left][col]) * ratio
    return filled, step, gaps


def analyze_column(rows: list[dict[str, Any]], col: str, window: int, threshold: float) -> dict[str, Any]:
    values = [row[col] for row in rows]
    ma = moving_average(values, window)
    valid = [v for v in values if v is not None]
    mu = mean(valid)
    varr = variance(valid, mu)
    sd = stddev(valid, mu)
    trend = linear_regression(values)
    season = detect_seasonality(values)
    slope = trend["slope"]
    direction = "stable"
    scale = sd if sd not in (None, 0) else 1.0
    if abs(slope) >= scale * 0.001:
        direction = "increasing" if slope > 0 else "decreasing"

    anomalies = []
    for i, v in enumerate(values):
        if v is None or ma[i] is None or sd in (None, 0):
            continue
        z = abs(v - ma[i]) / sd
        if z > threshold:
            anomalies.append(
                {
                    "index": i,
                    "timestamp": rows[i]["timestamp"],
                    "value": v,
                    "moving_average": ma[i],
                    "z_from_moving_average": z,
                }
            )

    return {
        "column": col,
        "stats": {
            "min": min(valid) if valid else None,
            "max": max(valid) if valid else None,
            "mean": mu,
            "variance": varr,
            "stddev": sd,
        },
        "trend": {
            "direction": direction,
            "slope": trend["slope"],
            "r_squared": trend["r_squared"],
        },
        "seasonality": season,
        "anomaly_count": len(anomalies),
        "anomalies": anomalies,
        "moving_average": ma,
    }


def parse_input_csv(path: Path) -> tuple[list[dict[str, Any]], list[str], str]:
    rows = []
    with path.open("r", encoding="utf-8-sig", newline="") as f:
        reader = csv.DictReader(f)
        if reader.fieldnames is None or len(reader.fieldnames) < 2:
            return [], [], "timestamp"
        ts_col = reader.fieldnames[0]
        value_cols = reader.fieldnames[1:]
        for rec in reader:
            ts = parse_timestamp(rec.get(ts_col))
            if ts is None:
                continue
            row = {
                "timestamp_ts": ts,
                "timestamp": datetime.fromtimestamp(ts, tz=timezone.utc).isoformat().replace("+00:00", "Z"),
            }
            for col in value_cols:
                raw = rec.get(col)
                try:
                    row[col] = float(raw) if raw not in (None, "") else None
                except (TypeError, ValueError):
                    row[col] = None
            rows.append(row)
    return rows, value_cols, ts_col


def generate_sample(path: Path) -> None:
    start = datetime(2025, 1, 1, tzinfo=timezone.utc)
    with path.open("w", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        writer.writerow(["timestamp", "metric_a", "metric_b"])
        for i in range(365):
            ts = (start + timedelta(days=i)).strftime("%Y-%m-%d")
            weekly = 10.0 * math.sin(2.0 * math.pi * i / 7.0)
            trend = i * 0.18
            a = 50 + trend + weekly + (math.sin(i) * 0.8)
            b = 30 + i * 0.05 + 5 * math.cos(2.0 * math.pi * i / 7.0)
            if i in {45, 123, 251}:
                a += 35
            if i in {200, 300}:
                b -= 20
            writer.writerow([ts, f"{a:.3f}", f"{b:.3f}"])


def export_processed(path: Path, rows: list[dict[str, Any]], analyses: list[dict[str, Any]]) -> None:
    headers = ["timestamp"]
    for a in analyses:
        c = a["column"]
        headers.extend([c, f"{c}_moving_average", f"{c}_is_anomaly"])
    with path.open("w", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        writer.writerow(headers)
        for i, row in enumerate(rows):
            out_row = [row["timestamp"]]
            for a in analyses:
                c = a["column"]
                is_anomaly = any(x["index"] == i for x in a["anomalies"])
                out_row.extend([row[c], a["moving_average"][i], 1 if is_anomaly else 0])
            writer.writerow(out_row)


def print_summary(analyses: list[dict[str, Any]], gaps: list[float], output_path: Path, export_path: Path | None) -> None:
    print("Time Series Trend Detector")
    print("==========================")
    print(f"Columns analyzed: {len(analyses)}")
    print(f"Gap points filled/interpolated: {len(gaps)}")
    print("")
    for a in analyses:
        print(f"Column: {a['column']}")
        print(
            f"  Trend      : {a['trend']['direction']} "
            f"(slope={a['trend']['slope']:.6f}, R^2={a['trend']['r_squared']:.4f})"
        )
        if a["seasonality"]["period"] is None:
            print("  Seasonality: none")
        else:
            print(
                f"  Seasonality: period={a['seasonality']['period']} "
                f"(autocorr={a['seasonality']['autocorrelation']:.4f})"
            )
        print(f"  Anomalies  : {a['anomaly_count']}")
        s = a["stats"]
        print(f"  Stats      : min={s['min']}, max={s['max']}, mean={s['mean']}, variance={s['variance']}")
        print("")
    print(f"JSON report saved : {output_path}")
    if export_path:
        print(f"Processed CSV saved: {export_path}")


def main() -> int:
    args = parse_args()
    window = max(1, args.window)
    threshold = args.threshold
    output_path = Path(args.output).resolve()
    export_path = Path(args.export).resolve() if args.export else None

    if args.input_file:
        input_path = Path(args.input_file).resolve()
        if not input_path.exists():
            print(f"Input file not found: {input_path}")
            return 1
    else:
        input_path = Path("sample_timeseries.csv").resolve()
        generate_sample(input_path)
        print(f"No input provided. Generated sample dataset: {input_path}")

    rows, value_cols, ts_col = parse_input_csv(input_path)
    if not rows or not value_cols:
        print("No valid rows or value columns found.")
        return 1

    filled, step, gaps = interpolate_gaps(rows, value_cols)
    analyses = [analyze_column(filled, col, window, threshold) for col in value_cols]

    report = {
        "metadata": {
            "input_file": str(input_path),
            "generated_at": datetime.now(tz=timezone.utc).isoformat().replace("+00:00", "Z"),
            "window": window,
            "threshold": threshold,
            "inferred_step_seconds": step,
            "interpolated_gap_points": [
                datetime.fromtimestamp(ts, tz=timezone.utc).isoformat().replace("+00:00", "Z") for ts in gaps
            ],
        },
        "timestamp_column": ts_col,
        "value_columns": value_cols,
        "analyses": [
            {
                "column": a["column"],
                "stats": a["stats"],
                "trend": a["trend"],
                "seasonality": a["seasonality"],
                "anomaly_count": a["anomaly_count"],
                "anomalies": [
                    {
                        "timestamp": x["timestamp"],
                        "value": x["value"],
                        "moving_average": x["moving_average"],
                        "z_from_moving_average": x["z_from_moving_average"],
                    }
                    for x in a["anomalies"]
                ],
            }
            for a in analyses
        ],
    }

    output_path.write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8")
    if export_path:
        export_processed(export_path, filled, analyses)
    print_summary(analyses, gaps, output_path, export_path)
    return 0


if __name__ == "__main__":
    raise SystemExit(main())