Time Series Trend Detector (python, written by Codex)
envgap__codex__python-t1-9
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/python-t1 #9 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Time Series Trend Detector Write a program that analyzes time series data to detect trends, seasonal patterns, and anomalies using statistical methods, and produces a visual summary report. FUNCTIONAL REQUIREMENTS: - Accept a CSV file path as a command-line argument with columns for timestamp and one or more numeric value columns - Parse timestamps in multiple formats (ISO 8601, Unix epoch, and common date formats like MM/DD/YYYY, YYYY-MM-DD HH:MM:SS) - Compute a moving average with a configurable window size via --window flag (default: 7 data points) - Detect overall trend direction (increasing, decreasing, stable) using linear regression and report the slope and R-squared value - Detect seasonality by computing autocorrelation at various lags and reporting the dominant period if one exists - Identify anomalies: data points that deviate more than a configurable number of standard deviations from the moving average (--threshold flag, default: 2.0) - Support multiple value columns: analyze each independently and report results for all - Generate a summary report with: trend direction and strength, seasonal period (if any), count and list of anomalies with their timestamps and values, basic statistics (min, max, mean, variance) - Save the report as a JSON file with --output flag (default: trend_report.json) - Export the processed data (original values, moving average, anomaly flags) as a CSV file via --export flag - If no input file is given, generate a sample time series dataset with 365 daily data points containing a linear trend, weekly seasonality, and injected anomalies, then analyze it - Handle missing timestamps and gaps in the series by interpolating or flagging them Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# Time Series Trend Detector (Python) Analyzes CSV time series data for trend, seasonality, anomalies, and summary statistics. ## Requirements - Ubuntu 22.04 - Python 3.10+ ## Dependencies - Direct: none - Transitive: none `requirements.txt` is pinned and intentionally dependency-free (standard library implementation). ## Run With input: ```bash python src/main.py /path/to/data.csv --window 7 --threshold 2.0 --output trend_report.json --export processed.csv ``` No input (generates a 365-day sample and analyzes it): ```bash python src/main.py ``` ## Features - Timestamp parsing: ISO 8601, Unix epoch, `MM/DD/YYYY`, `YYYY-MM-DD HH:MM:SS` - Moving average (`--window`) - Trend detection with linear regression slope + R-squared - Seasonality detection via autocorrelation lag scan - Anomaly detection against moving average (`--threshold` std dev) - Gap interpolation and reporting - JSON report and processed CSV export
requirements.txt
# No external dependencies required. # Direct dependencies: none # Transitive dependencies: none
src/main.py
#!/usr/bin/env python3
import argparse
import csv
import json
import math
from datetime import datetime, timezone, timedelta
from pathlib import Path
from typing import Any
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Time Series Trend Detector")
parser.add_argument("input_file", nargs="?", help="CSV input file")
parser.add_argument("--window", type=int, default=7, help="Moving average window size")
parser.add_argument("--threshold", type=float, default=2.0, help="Std-dev threshold for anomalies")
parser.add_argument("--output", default="trend_report.json", help="JSON report output path")
parser.add_argument("--export", help="Processed CSV export path")
return parser.parse_args()
def parse_timestamp(raw: Any) -> float | None:
if raw is None:
return None
s = str(raw).strip()
if not s:
return None
if s.isdigit() and len(s) == 10:
return float(int(s))
if s.isdigit() and len(s) == 13:
return float(int(s) / 1000)
for fmt in ("%m/%d/%Y", "%Y-%m-%d %H:%M:%S", "%Y-%m-%d", "%Y-%m-%dT%H:%M:%S", "%Y-%m-%dT%H:%M:%SZ"):
try:
dt = datetime.strptime(s, fmt).replace(tzinfo=timezone.utc)
return dt.timestamp()
except ValueError:
continue
try:
return datetime.fromisoformat(s.replace("Z", "+00:00")).timestamp()
except ValueError:
return None
def mean(values: list[float]) -> float | None:
return sum(values) / len(values) if values else None
def variance(values: list[float], m: float | None = None) -> float | None:
if not values:
return None
mu = mean(values) if m is None else m
assert mu is not None
return sum((x - mu) ** 2 for x in values) / len(values)
def stddev(values: list[float], m: float | None = None) -> float | None:
v = variance(values, m)
return math.sqrt(v) if v is not None else None
def moving_average(values: list[float | None], window: int) -> list[float | None]:
out: list[float | None] = [None] * len(values)
for i in range(len(values)):
start = max(0, i - window + 1)
chunk = [x for x in values[start : i + 1] if x is not None]
out[i] = sum(chunk) / len(chunk) if chunk else None
return out
def linear_regression(values: list[float | None]) -> dict[str, float]:
points = [(i, v) for i, v in enumerate(values) if v is not None]
if len(points) < 2:
return {"slope": 0.0, "intercept": 0.0, "r_squared": 0.0}
xs = [float(x) for x, _ in points]
ys = [float(y) for _, y in points]
mx = sum(xs) / len(xs)
my = sum(ys) / len(ys)
cov = sum((x - mx) * (y - my) for x, y in points)
vx = sum((x - mx) ** 2 for x in xs)
slope = cov / vx if vx != 0 else 0.0
intercept = my - slope * mx
ss_res = sum((y - (slope * x + intercept)) ** 2 for x, y in points)
ss_tot = sum((y - my) ** 2 for y in ys)
r2 = 0.0 if ss_tot == 0 else 1.0 - ss_res / ss_tot
return {"slope": slope, "intercept": intercept, "r_squared": r2}
def autocorrelation(values: list[float | None], lag: int) -> float | None:
pairs = [(values[i], values[i - lag]) for i in range(lag, len(values)) if values[i] is not None and values[i - lag] is not None]
if len(pairs) < 3:
return None
xs = [x for x, _ in pairs]
ys = [y for _, y in pairs]
mx = sum(xs) / len(xs)
my = sum(ys) / len(ys)
num = sum((x - mx) * (y - my) for x, y in pairs)
dx = sum((x - mx) ** 2 for x in xs)
dy = sum((y - my) ** 2 for y in ys)
den = math.sqrt(dx * dy)
return 0.0 if den == 0 else num / den
def detect_seasonality(values: list[float | None]) -> dict[str, float | int | None]:
max_lag = min(60, len(values) // 2)
best_lag: int | None = None
best_corr = 0.0
for lag in range(2, max_lag + 1):
corr = autocorrelation(values, lag)
if corr is None:
continue
if abs(corr) > abs(best_corr):
best_corr = corr
best_lag = lag
if best_lag is None or abs(best_corr) < 0.3:
return {"period": None, "autocorrelation": None}
return {"period": best_lag, "autocorrelation": best_corr}
def infer_step(sorted_ts: list[float]) -> float:
diffs = [sorted_ts[i] - sorted_ts[i - 1] for i in range(1, len(sorted_ts)) if sorted_ts[i] > sorted_ts[i - 1]]
if not diffs:
return 86400.0
diffs.sort()
return diffs[len(diffs) // 2]
def interpolate_gaps(rows: list[dict[str, Any]], value_cols: list[str]) -> tuple[list[dict[str, Any]], float, list[float]]:
ordered = sorted(rows, key=lambda r: r["timestamp_ts"])
ts_list = [r["timestamp_ts"] for r in ordered]
step = infer_step(ts_list)
by_ts = {r["timestamp_ts"]: r for r in ordered}
start = ts_list[0]
end = ts_list[-1]
t = start
filled: list[dict[str, Any]] = []
gaps: list[float] = []
while t <= end + 1e-9:
if t in by_ts:
item = dict(by_ts[t])
item["gap_filled"] = False
filled.append(item)
else:
item = {
"timestamp_ts": t,
"timestamp": datetime.fromtimestamp(t, tz=timezone.utc).isoformat().replace("+00:00", "Z"),
"gap_filled": True,
}
for c in value_cols:
item[c] = None
filled.append(item)
gaps.append(t)
t += step
for col in value_cols:
for i in range(len(filled)):
if filled[i][col] is not None:
continue
left = i - 1
while left >= 0 and filled[left][col] is None:
left -= 1
right = i + 1
while right < len(filled) and filled[right][col] is None:
right += 1
if left >= 0 and right < len(filled):
ratio = (filled[i]["timestamp_ts"] - filled[left]["timestamp_ts"]) / (
filled[right]["timestamp_ts"] - filled[left]["timestamp_ts"]
)
filled[i][col] = filled[left][col] + (filled[right][col] - filled[left][col]) * ratio
return filled, step, gaps
def analyze_column(rows: list[dict[str, Any]], col: str, window: int, threshold: float) -> dict[str, Any]:
values = [row[col] for row in rows]
ma = moving_average(values, window)
valid = [v for v in values if v is not None]
mu = mean(valid)
varr = variance(valid, mu)
sd = stddev(valid, mu)
trend = linear_regression(values)
season = detect_seasonality(values)
slope = trend["slope"]
direction = "stable"
scale = sd if sd not in (None, 0) else 1.0
if abs(slope) >= scale * 0.001:
direction = "increasing" if slope > 0 else "decreasing"
anomalies = []
for i, v in enumerate(values):
if v is None or ma[i] is None or sd in (None, 0):
continue
z = abs(v - ma[i]) / sd
if z > threshold:
anomalies.append(
{
"index": i,
"timestamp": rows[i]["timestamp"],
"value": v,
"moving_average": ma[i],
"z_from_moving_average": z,
}
)
return {
"column": col,
"stats": {
"min": min(valid) if valid else None,
"max": max(valid) if valid else None,
"mean": mu,
"variance": varr,
"stddev": sd,
},
"trend": {
"direction": direction,
"slope": trend["slope"],
"r_squared": trend["r_squared"],
},
"seasonality": season,
"anomaly_count": len(anomalies),
"anomalies": anomalies,
"moving_average": ma,
}
def parse_input_csv(path: Path) -> tuple[list[dict[str, Any]], list[str], str]:
rows = []
with path.open("r", encoding="utf-8-sig", newline="") as f:
reader = csv.DictReader(f)
if reader.fieldnames is None or len(reader.fieldnames) < 2:
return [], [], "timestamp"
ts_col = reader.fieldnames[0]
value_cols = reader.fieldnames[1:]
for rec in reader:
ts = parse_timestamp(rec.get(ts_col))
if ts is None:
continue
row = {
"timestamp_ts": ts,
"timestamp": datetime.fromtimestamp(ts, tz=timezone.utc).isoformat().replace("+00:00", "Z"),
}
for col in value_cols:
raw = rec.get(col)
try:
row[col] = float(raw) if raw not in (None, "") else None
except (TypeError, ValueError):
row[col] = None
rows.append(row)
return rows, value_cols, ts_col
def generate_sample(path: Path) -> None:
start = datetime(2025, 1, 1, tzinfo=timezone.utc)
with path.open("w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerow(["timestamp", "metric_a", "metric_b"])
for i in range(365):
ts = (start + timedelta(days=i)).strftime("%Y-%m-%d")
weekly = 10.0 * math.sin(2.0 * math.pi * i / 7.0)
trend = i * 0.18
a = 50 + trend + weekly + (math.sin(i) * 0.8)
b = 30 + i * 0.05 + 5 * math.cos(2.0 * math.pi * i / 7.0)
if i in {45, 123, 251}:
a += 35
if i in {200, 300}:
b -= 20
writer.writerow([ts, f"{a:.3f}", f"{b:.3f}"])
def export_processed(path: Path, rows: list[dict[str, Any]], analyses: list[dict[str, Any]]) -> None:
headers = ["timestamp"]
for a in analyses:
c = a["column"]
headers.extend([c, f"{c}_moving_average", f"{c}_is_anomaly"])
with path.open("w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerow(headers)
for i, row in enumerate(rows):
out_row = [row["timestamp"]]
for a in analyses:
c = a["column"]
is_anomaly = any(x["index"] == i for x in a["anomalies"])
out_row.extend([row[c], a["moving_average"][i], 1 if is_anomaly else 0])
writer.writerow(out_row)
def print_summary(analyses: list[dict[str, Any]], gaps: list[float], output_path: Path, export_path: Path | None) -> None:
print("Time Series Trend Detector")
print("==========================")
print(f"Columns analyzed: {len(analyses)}")
print(f"Gap points filled/interpolated: {len(gaps)}")
print("")
for a in analyses:
print(f"Column: {a['column']}")
print(
f" Trend : {a['trend']['direction']} "
f"(slope={a['trend']['slope']:.6f}, R^2={a['trend']['r_squared']:.4f})"
)
if a["seasonality"]["period"] is None:
print(" Seasonality: none")
else:
print(
f" Seasonality: period={a['seasonality']['period']} "
f"(autocorr={a['seasonality']['autocorrelation']:.4f})"
)
print(f" Anomalies : {a['anomaly_count']}")
s = a["stats"]
print(f" Stats : min={s['min']}, max={s['max']}, mean={s['mean']}, variance={s['variance']}")
print("")
print(f"JSON report saved : {output_path}")
if export_path:
print(f"Processed CSV saved: {export_path}")
def main() -> int:
args = parse_args()
window = max(1, args.window)
threshold = args.threshold
output_path = Path(args.output).resolve()
export_path = Path(args.export).resolve() if args.export else None
if args.input_file:
input_path = Path(args.input_file).resolve()
if not input_path.exists():
print(f"Input file not found: {input_path}")
return 1
else:
input_path = Path("sample_timeseries.csv").resolve()
generate_sample(input_path)
print(f"No input provided. Generated sample dataset: {input_path}")
rows, value_cols, ts_col = parse_input_csv(input_path)
if not rows or not value_cols:
print("No valid rows or value columns found.")
return 1
filled, step, gaps = interpolate_gaps(rows, value_cols)
analyses = [analyze_column(filled, col, window, threshold) for col in value_cols]
report = {
"metadata": {
"input_file": str(input_path),
"generated_at": datetime.now(tz=timezone.utc).isoformat().replace("+00:00", "Z"),
"window": window,
"threshold": threshold,
"inferred_step_seconds": step,
"interpolated_gap_points": [
datetime.fromtimestamp(ts, tz=timezone.utc).isoformat().replace("+00:00", "Z") for ts in gaps
],
},
"timestamp_column": ts_col,
"value_columns": value_cols,
"analyses": [
{
"column": a["column"],
"stats": a["stats"],
"trend": a["trend"],
"seasonality": a["seasonality"],
"anomaly_count": a["anomaly_count"],
"anomalies": [
{
"timestamp": x["timestamp"],
"value": x["value"],
"moving_average": x["moving_average"],
"z_from_moving_average": x["z_from_moving_average"],
}
for x in a["anomalies"]
],
}
for a in analyses
],
}
output_path.write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8")
if export_path:
export_processed(export_path, filled, analyses)
print_summary(analyses, gaps, output_path, export_path)
return 0
if __name__ == "__main__":
raise SystemExit(main())