← All tasks
pythonclaude-code/python-t1 #5Not a task: already works

Log File Pattern Analyzer (python, written by Claude Code)

envgap__claude-code__python-t1-5

Written by a coding agent; not on GitHubWritten 2026-02-27

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

claude-code/python-t1 #5 · read the task the agent was given
Claude Code wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Log File Pattern Analyzer

Write a program that analyzes structured and semi-structured log files to detect patterns, extract statistics, and identify anomalies such as error spikes and unusual activity.

FUNCTIONAL REQUIREMENTS:
- Accept a log file path as a command-line argument
- Auto-detect common log formats: Apache/Nginx access logs, syslog, and JSON-structured logs
- Parse timestamps, log levels (DEBUG, INFO, WARN, ERROR, FATAL), source identifiers, and message content
- Compute statistics: total entries, entries per log level, entries per hour/day, top 10 most frequent messages (grouped by template after removing variable parts like IPs, timestamps, and IDs)
- Detect error spikes: flag any time window where the error rate exceeds 3x the overall average error rate
- Support filtering by date range via --from and --to flags (ISO 8601 format)
- Support filtering by log level via --level flag (show that level and above)
- Print a summary report to console with counts, top patterns, and detected anomalies
- Save the full analysis as a JSON report file with --output flag (default: log_analysis.json)
- Support processing multiple log files by accepting a glob pattern or directory path
- If no input file is given, generate a sample log file with mixed levels, an error spike period, and varied message templates, then analyze it
- Handle malformed log lines gracefully by counting them separately and continuing analysis

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

analyzer.py
#!/usr/bin/env python3
"""Log File Pattern Analyzer - Trial 1 (pandas + matplotlib)

Analyzes structured/semi-structured log files to detect patterns,
extract statistics, and identify anomalies.
"""

import argparse
import json
import os
import re
import sys
import random
import math
from datetime import datetime, timedelta
from collections import Counter

import pandas as pd
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt


# ---------------------------------------------------------------------------
# Log format patterns
# ---------------------------------------------------------------------------
SYSLOG_RE = re.compile(
    r"^(?P<timestamp>\w{3}\s+\d{1,2}\s+\d{2}:\d{2}:\d{2})\s+"
    r"(?P<source>\S+)\s+(?P<process>\S+?):\s+(?P<message>.*)$"
)
APACHE_RE = re.compile(
    r'^(?P<client>\S+)\s+\S+\s+\S+\s+'
    r'\[(?P<timestamp>[^\]]+)\]\s+'
    r'"(?P<method>\S+)\s+(?P<path>\S+)\s+\S+"\s+'
    r'(?P<status>\d{3})\s+(?P<size>\S+)'
    r'(?:\s+"[^"]*"\s+"[^"]*")?'
    r'(?:\s+(?P<response_time>\d+))?'
)
JSON_LOG_KEYS = {"timestamp", "level", "message"}

SYSLOG_LEVEL_MAP = {
    "emerg": "CRITICAL", "alert": "CRITICAL", "crit": "CRITICAL",
    "err": "ERROR", "error": "ERROR",
    "warn": "WARNING", "warning": "WARNING",
    "notice": "INFO", "info": "INFO",
    "debug": "DEBUG",
}

STATUS_LEVEL = {
    "2": "INFO", "3": "INFO", "4": "WARNING", "5": "ERROR",
}


# ---------------------------------------------------------------------------
# Sample log generator
# ---------------------------------------------------------------------------
def generate_sample_log(path: str, num_lines: int = 2000) -> str:
    """Generate a mixed-format sample log file."""
    levels = ["DEBUG", "INFO", "INFO", "INFO", "WARNING", "ERROR", "CRITICAL"]
    sources = ["web-server", "auth-service", "db-worker", "scheduler", "cache"]
    methods = ["GET", "POST", "PUT", "DELETE"]
    paths = ["/api/users", "/api/orders", "/api/products", "/health", "/login", "/api/search"]
    messages = [
        "Request processed successfully",
        "Connection established",
        "Cache miss for key user_session",
        "Database query took 320ms",
        "Authentication failed for user admin",
        "Rate limit exceeded",
        "Timeout waiting for upstream",
        "Disk usage above 90%",
        "Memory allocation failed",
        "Service restarted",
    ]

    base_time = datetime(2024, 6, 1, 0, 0, 0)
    lines = []

    for i in range(num_lines):
        ts = base_time + timedelta(seconds=i * 2 + random.randint(0, 3))
        fmt_choice = random.choices(["syslog", "apache", "json"], weights=[30, 40, 30])[0]

        # Inject anomaly spike between lines 800-850
        if 800 <= i <= 850:
            level = random.choice(["ERROR", "CRITICAL"])
        else:
            level = random.choice(levels)

        if fmt_choice == "syslog":
            src = random.choice(sources)
            msg = random.choice(messages)
            syslog_ts = ts.strftime("%b %d %H:%M:%S")
            lines.append(f"{syslog_ts} {src} app[{random.randint(1000,9999)}]: [{level}] {msg}")

        elif fmt_choice == "apache":
            ip = f"192.168.{random.randint(1,10)}.{random.randint(1,254)}"
            method = random.choice(methods)
            p = random.choice(paths)
            status = {"INFO": 200, "WARNING": 404, "ERROR": 500, "DEBUG": 200, "CRITICAL": 503}[level]
            size = random.randint(200, 50000)
            rt = random.randint(5, 2000)
            apache_ts = ts.strftime("%d/%b/%Y:%H:%M:%S +0000")
            lines.append(
                f'{ip} - - [{apache_ts}] "{method} {p} HTTP/1.1" {status} {size} "-" "Mozilla/5.0" {rt}'
            )

        else:
            record = {
                "timestamp": ts.isoformat(),
                "level": level,
                "source": random.choice(sources),
                "message": random.choice(messages),
            }
            lines.append(json.dumps(record))

        # Inject occasional malformed lines
        if random.random() < 0.02:
            lines.append("<<<MALFORMED LINE -- random garbage @#$% >>>")

    with open(path, "w", encoding="utf-8") as fh:
        fh.write("\n".join(lines) + "\n")

    return path


# ---------------------------------------------------------------------------
# Parsing helpers
# ---------------------------------------------------------------------------
def _infer_syslog_level(message: str) -> str:
    msg_lower = message.lower()
    for keyword, level in SYSLOG_LEVEL_MAP.items():
        if keyword in msg_lower:
            return level
    # Check bracketed level like [ERROR]
    m = re.search(r"\[(\w+)\]", message)
    if m:
        candidate = m.group(1).upper()
        if candidate in {"DEBUG", "INFO", "WARNING", "ERROR", "CRITICAL"}:
            return candidate
    return "INFO"


def parse_line(line: str) -> dict | None:
    """Attempt to parse a single log line. Returns dict or None on failure."""
    line = line.strip()
    if not line:
        return None

    # Try JSON first
    if line.startswith("{"):
        try:
            obj = json.loads(line)
            if "timestamp" in obj:
                ts = pd.to_datetime(obj["timestamp"], errors="coerce")
                return {
                    "timestamp": ts,
                    "level": obj.get("level", "INFO").upper(),
                    "source": obj.get("source", "unknown"),
                    "message": obj.get("message", ""),
                    "response_time": obj.get("response_time"),
                    "format": "json",
                }
        except (json.JSONDecodeError, KeyError):
            pass

    # Try Apache combined
    m = APACHE_RE.match(line)
    if m:
        d = m.groupdict()
        ts = pd.to_datetime(d["timestamp"], format="%d/%b/%Y:%H:%M:%S %z", errors="coerce")
        status_class = d["status"][0]
        level = STATUS_LEVEL.get(status_class, "INFO")
        rt = int(d["response_time"]) if d.get("response_time") else None
        return {
            "timestamp": ts,
            "level": level,
            "source": d["client"],
            "message": f'{d["method"]} {d["path"]} {d["status"]}',
            "response_time": rt,
            "format": "apache",
        }

    # Try syslog
    m = SYSLOG_RE.match(line)
    if m:
        d = m.groupdict()
        try:
            ts = pd.to_datetime(d["timestamp"], format="%b %d %H:%M:%S", errors="coerce")
            if pd.notna(ts):
                ts = ts.replace(year=datetime.now().year)
        except Exception:
            ts = pd.NaT
        level = _infer_syslog_level(d["message"])
        return {
            "timestamp": ts,
            "level": level,
            "source": d["source"],
            "message": d["message"],
            "response_time": None,
            "format": "syslog",
        }

    return None


# ---------------------------------------------------------------------------
# Core analysis
# ---------------------------------------------------------------------------
def analyse_logs(filepath: str) -> dict:
    """Parse and analyse the given log file, returning a report dict."""
    records = []
    malformed = 0
    total_lines = 0

    with open(filepath, "r", encoding="utf-8", errors="replace") as fh:
        for line in fh:
            total_lines += 1
            parsed = parse_line(line)
            if parsed is None:
                malformed += 1
            else:
                records.append(parsed)

    if not records:
        return {
            "file": filepath,
            "total_lines": total_lines,
            "parsed_lines": 0,
            "malformed_lines": malformed,
            "error": "No parseable log lines found.",
        }

    df = pd.DataFrame(records)
    df["timestamp"] = pd.to_datetime(df["timestamp"], errors="coerce", utc=True)
    df.sort_values("timestamp", inplace=True)
    df.reset_index(drop=True, inplace=True)

    report: dict = {
        "file": filepath,
        "total_lines": total_lines,
        "parsed_lines": len(df),
        "malformed_lines": malformed,
    }

    # --- Level distribution ---
    level_counts = df["level"].value_counts().to_dict()
    report["level_distribution"] = {k: int(v) for k, v in level_counts.items()}

    # --- Error rate ---
    error_count = int(df["level"].isin(["ERROR", "CRITICAL"]).sum())
    report["error_count"] = error_count
    report["error_rate"] = round(error_count / len(df) * 100, 2)

    # --- Format distribution ---
    fmt_counts = df["format"].value_counts().to_dict()
    report["format_distribution"] = {k: int(v) for k, v in fmt_counts.items()}

    # --- Top sources ---
    top_sources = df["source"].value_counts().head(10).to_dict()
    report["top_sources"] = {k: int(v) for k, v in top_sources.items()}

    # --- Response time stats ---
    rt = df["response_time"].dropna()
    if len(rt) > 0:
        report["response_time"] = {
            "count": int(len(rt)),
            "mean_ms": round(float(rt.mean()), 2),
            "median_ms": round(float(rt.median()), 2),
            "p95_ms": round(float(rt.quantile(0.95)), 2),
            "p99_ms": round(float(rt.quantile(0.99)), 2),
            "max_ms": round(float(rt.max()), 2),
            "min_ms": round(float(rt.min()), 2),
        }

    # --- Time window analysis ---
    valid_ts = df.dropna(subset=["timestamp"])
    if len(valid_ts) > 1:
        valid_ts = valid_ts.set_index("timestamp")
        time_range = valid_ts.index.max() - valid_ts.index.min()
        report["time_range"] = {
            "start": str(valid_ts.index.min()),
            "end": str(valid_ts.index.max()),
            "duration_seconds": round(time_range.total_seconds(), 1),
        }

        # Determine bucket size
        total_sec = time_range.total_seconds()
        if total_sec <= 3600:
            freq = "1min"
        elif total_sec <= 86400:
            freq = "5min"
        else:
            freq = "1h"

        # Request rates
        counts_series = valid_ts.resample(freq).size()
        report["request_rate"] = {
            "bucket": freq,
            "mean_per_bucket": round(float(counts_series.mean()), 2),
            "max_per_bucket": int(counts_series.max()),
            "min_per_bucket": int(counts_series.min()),
        }

        # Error time series
        errors_ts = valid_ts[valid_ts["level"].isin(["ERROR", "CRITICAL"])]
        error_series = errors_ts.resample(freq).size()

        # Anomaly detection (simple z-score on error counts per bucket)
        anomalies = []
        if len(error_series) > 3:
            mean_err = error_series.mean()
            std_err = error_series.std()
            if std_err > 0:
                z_scores = (error_series - mean_err) / std_err
                for ts_bucket, z in z_scores.items():
                    if z > 2.0:
                        anomalies.append({
                            "window": str(ts_bucket),
                            "error_count": int(error_series[ts_bucket]),
                            "z_score": round(float(z), 2),
                            "type": "error_spike",
                        })

        # Pattern anomaly: look for repeated errors
        if len(errors_ts) > 0:
            msg_counts = errors_ts["message"].value_counts()
            for msg, cnt in msg_counts.head(5).items():
                if cnt > len(errors_ts) * 0.2:
                    anomalies.append({
                        "type": "repeated_error",
                        "message": str(msg),
                        "count": int(cnt),
                        "percentage": round(cnt / len(errors_ts) * 100, 2),
                    })

        report["anomalies"] = anomalies
        report["anomaly_count"] = len(anomalies)

        # --- Generate chart ---
        try:
            fig, axes = plt.subplots(2, 2, figsize=(14, 10))
            fig.suptitle("Log Analysis Report", fontsize=16)

            # 1. Level distribution pie chart
            ax = axes[0, 0]
            labels = list(level_counts.keys())
            values = list(level_counts.values())
            colors = {"DEBUG": "#6c757d", "INFO": "#0d6efd", "WARNING": "#ffc107",
                       "ERROR": "#dc3545", "CRITICAL": "#842029"}
            pie_colors = [colors.get(l, "#adb5bd") for l in labels]
            ax.pie(values, labels=labels, autopct="%1.1f%%", colors=pie_colors)
            ax.set_title("Log Level Distribution")

            # 2. Request rate over time
            ax = axes[0, 1]
            counts_series.plot(ax=ax, color="#0d6efd")
            ax.set_title(f"Request Rate ({freq} buckets)")
            ax.set_ylabel("Count")

            # 3. Error rate over time
            ax = axes[1, 0]
            error_series.plot(ax=ax, color="#dc3545")
            ax.set_title(f"Error Rate ({freq} buckets)")
            ax.set_ylabel("Error Count")
            # Mark anomalies
            for a in anomalies:
                if a["type"] == "error_spike":
                    anom_ts = pd.Timestamp(a["window"])
                    if anom_ts in error_series.index:
                        ax.axvline(x=anom_ts, color="orange", linestyle="--", alpha=0.7)

            # 4. Response time distribution
            ax = axes[1, 1]
            if len(rt) > 0:
                rt.hist(bins=50, ax=ax, color="#198754", edgecolor="white")
                ax.set_title("Response Time Distribution (ms)")
                ax.set_xlabel("ms")
            else:
                ax.text(0.5, 0.5, "No response time data", ha="center", va="center",
                        transform=ax.transAxes)
                ax.set_title("Response Time Distribution")

            plt.tight_layout()
            chart_path = os.path.join(os.path.dirname(filepath), "analysis_chart.png")
            fig.savefig(chart_path, dpi=100)
            plt.close(fig)
            report["chart"] = chart_path
        except Exception as e:
            report["chart_error"] = str(e)

    return report


# ---------------------------------------------------------------------------
# Console output
# ---------------------------------------------------------------------------
def print_report(report: dict) -> None:
    """Pretty-print the analysis report to the console."""
    sep = "=" * 70
    print(f"\n{sep}")
    print("  LOG FILE PATTERN ANALYZER - ANALYSIS REPORT")
    print(sep)
    print(f"  File:            {report.get('file', 'N/A')}")
    print(f"  Total lines:     {report.get('total_lines', 0)}")
    print(f"  Parsed lines:    {report.get('parsed_lines', 0)}")
    print(f"  Malformed lines: {report.get('malformed_lines', 0)}")
    print()

    if "error" in report:
        print(f"  ERROR: {report['error']}")
        print(sep)
        return

    print("  -- Level Distribution --")
    for level, count in sorted(report.get("level_distribution", {}).items()):
        pct = count / report["parsed_lines"] * 100
        bar = "#" * int(pct / 2)
        print(f"    {level:<10} {count:>6}  ({pct:5.1f}%)  {bar}")

    print()
    print(f"  Error count: {report.get('error_count', 0)}")
    print(f"  Error rate:  {report.get('error_rate', 0)}%")
    print()

    if "response_time" in report:
        rt = report["response_time"]
        print("  -- Response Time (ms) --")
        print(f"    Mean:   {rt['mean_ms']}")
        print(f"    Median: {rt['median_ms']}")
        print(f"    P95:    {rt['p95_ms']}")
        print(f"    P99:    {rt['p99_ms']}")
        print(f"    Max:    {rt['max_ms']}")
        print()

    if "time_range" in report:
        tr = report["time_range"]
        print("  -- Time Range --")
        print(f"    Start:    {tr['start']}")
        print(f"    End:      {tr['end']}")
        print(f"    Duration: {tr['duration_seconds']}s")
        print()

    if "request_rate" in report:
        rr = report["request_rate"]
        print(f"  -- Request Rate ({rr['bucket']} buckets) --")
        print(f"    Mean: {rr['mean_per_bucket']}")
        print(f"    Max:  {rr['max_per_bucket']}")
        print(f"    Min:  {rr['min_per_bucket']}")
        print()

    if "format_distribution" in report:
        print("  -- Log Format Distribution --")
        for fmt, cnt in report["format_distribution"].items():
            print(f"    {fmt:<10} {cnt:>6}")
        print()

    if "top_sources" in report:
        print("  -- Top Sources --")
        for src, cnt in report["top_sources"].items():
            print(f"    {src:<25} {cnt:>6}")
        print()

    anomalies = report.get("anomalies", [])
    print(f"  -- Anomalies Detected: {len(anomalies)} --")
    for i, a in enumerate(anomalies, 1):
        if a["type"] == "error_spike":
            print(f"    [{i}] ERROR SPIKE at {a['window']} "
                  f"(count={a['error_count']}, z={a['z_score']})")
        elif a["type"] == "repeated_error":
            print(f"    [{i}] REPEATED ERROR: \"{a['message']}\" "
                  f"(count={a['count']}, {a['percentage']}%)")
        else:
            print(f"    [{i}] {a}")

    if "chart" in report:
        print(f"\n  Chart saved to: {report['chart']}")

    print(sep)


# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main() -> None:
    parser = argparse.ArgumentParser(
        description="Log File Pattern Analyzer (pandas + matplotlib)"
    )
    parser.add_argument("logfile", nargs="?", default=None,
                        help="Path to the log file. If omitted, a sample is generated.")
    parser.add_argument("-o", "--output", default="analysis_report.json",
                        help="Output JSON report path (default: analysis_report.json)")
    args = parser.parse_args()

    if args.logfile is None:
        sample_path = os.path.join(os.getcwd(), "sample.log")
        print(f"No log file specified. Generating sample log at {sample_path} ...")
        generate_sample_log(sample_path)
        args.logfile = sample_path

    if not os.path.isfile(args.logfile):
        print(f"Error: File not found: {args.logfile}", file=sys.stderr)
        sys.exit(1)

    print(f"Analyzing {args.logfile} ...")
    report = analyse_logs(args.logfile)
    print_report(report)

    # Write JSON report
    output_path = args.output
    # Make chart path relative for JSON portability
    report_copy = dict(report)
    if "chart" in report_copy:
        report_copy["chart"] = os.path.basename(report_copy["chart"])

    with open(output_path, "w", encoding="utf-8") as fh:
        json.dump(report_copy, fh, indent=2, default=str)
    print(f"\nJSON report written to {output_path}")


if __name__ == "__main__":
    main()
README.md
# Log File Pattern Analyzer (Trial 1 - Python: pandas + matplotlib)

Analyzes structured/semi-structured log files to detect patterns, extract statistics, and identify anomalies.

## Dependencies
- pandas 2.2.2
- matplotlib 3.9.0

## Installation
```bash
pip install -r requirements.txt
```

## Usage
```bash
# Analyze a log file
python analyzer.py /path/to/logfile.log

# Generate sample log and analyze
python analyzer.py

# Custom output path
python analyzer.py /path/to/logfile.log -o my_report.json
```

## Features
- Parses syslog, Apache combined, and JSON log formats
- Extracts timestamps, log levels, sources
- Computes error rates, request rates, response time statistics
- Detects anomalies (error spikes via z-score, repeated error patterns)
- Time-window analysis with configurable bucket sizes
- Generates analysis chart (PNG) and JSON report
- Handles malformed lines gracefully
requirements.txt
pandas==2.2.2
matplotlib==3.9.0