← All tasks
pythoncodex/python-t1 #21Not a task: already works

Image Histogram Analyzer (python, written by Codex)

envgap__codex__python-t1-21

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/python-t1 #21 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Image Histogram Analyzer

Write a program that computes and analyzes color histograms of images, providing statistical analysis of color distribution, channel comparisons, and similarity scoring between images.

FUNCTIONAL REQUIREMENTS:
- Accept an image file path as a command-line argument
- Compute per-channel histograms (Red, Green, Blue) with 256 bins each, plus a luminance/grayscale histogram
- Calculate statistics for each channel: mean intensity, median, standard deviation, skewness, dominant intensity ranges, and dynamic range (difference between darkest and brightest used values)
- Detect if an image is overexposed (high mean, clipped highlights), underexposed (low mean, clipped shadows), or low contrast (narrow histogram spread)
- Support histogram comparison between two images via --compare flag: compute correlation, chi-squared distance, intersection, and Bhattacharyya distance between their histograms
- Support cumulative histogram computation for each channel via --cumulative flag
- Generate a histogram data output as a CSV file with columns (bin, red_count, green_count, blue_count, luminance_count) via --export flag
- Support analyzing specific regions of an image via --crop flag (x,y,width,height)
- Print a text-based summary to console: per-channel statistics, exposure assessment, contrast assessment, and color balance analysis
- Save the full analysis as JSON with --output flag (default: histogram_analysis.json)
- Support batch analysis of multiple images via --batch flag with a summary comparison table
- If no input is given, generate three sample images (one overexposed, one underexposed, one well-balanced), analyze each, and display comparative results
- Handle errors: unsupported image formats, corrupted files, grayscale images (single-channel analysis)

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

README.md
# Image Histogram Analyzer (Python)

Computes RGB/luminance histograms, statistical channel analysis, exposure/contrast detection, histogram similarity metrics, CSV export, JSON output, and batch analysis.

## Requirements

- Ubuntu 22.04
- Python 3.10+

## Dependency (Pinned)

- `Pillow==10.4.0`

## Setup

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

## Run

```bash
python src/main.py input.jpg
python src/main.py input.jpg --crop 50,50,400,300 --cumulative --export histogram.csv --output analysis.json
python src/main.py input.jpg --compare other.jpg
python src/main.py ./images --batch --output batch_histogram_analysis.json
python src/main.py
```

## Output

- Console summary with per-channel statistics, exposure and contrast assessment, plus color balance analysis.
- JSON output defaults to `histogram_analysis.json`.
- CSV schema: `bin,red_count,green_count,blue_count,luminance_count`.
- No-input mode generates overexposed, underexposed, and balanced sample images and compares their histograms.
requirements.txt
Pillow==10.4.0
src/main.py
#!/usr/bin/env python3
from __future__ import annotations

import argparse
import csv
import json
from pathlib import Path
from typing import Iterable

from PIL import Image

SUPPORTED = {".png", ".jpg", ".jpeg", ".bmp", ".tif", ".tiff", ".webp"}


def parse_crop(raw: str | None, width: int, height: int) -> tuple[int, int, int, int]:
    if not raw:
        return 0, 0, width, height
    try:
        x, y, w, h = [int(v.strip()) for v in raw.split(",", 3)]
    except Exception as exc:
        raise ValueError("Invalid --crop. Use x,y,width,height.") from exc
    if w <= 0 or h <= 0:
        raise ValueError("Crop width and height must be positive.")
    x = max(0, x)
    y = max(0, y)
    w = min(w, width - x)
    h = min(h, height - y)
    if w <= 0 or h <= 0:
        raise ValueError("Crop is outside image bounds.")
    return x, y, w, h


def cumulative(hist: list[int]) -> list[int]:
    out = []
    run = 0
    for v in hist:
        run += v
        out.append(run)
    return out


def stats_from_hist(hist: list[int]) -> dict:
    total = sum(hist)
    if total == 0:
        return {
            "mean": 0.0,
            "median": 0.0,
            "stddev": 0.0,
            "skewness": 0.0,
            "dominant_ranges": [],
            "dynamic_range": 0,
        }

    mean = sum(i * hist[i] for i in range(256)) / total
    variance = sum(((i - mean) ** 2) * hist[i] for i in range(256)) / total
    stddev = variance ** 0.5
    if stddev > 0:
        m3 = sum(((i - mean) ** 3) * hist[i] for i in range(256)) / total
        skewness = m3 / (stddev ** 3)
    else:
        skewness = 0.0

    half = total / 2
    run = 0
    median = 0
    for i, count in enumerate(hist):
        run += count
        if run >= half:
            median = i
            break

    min_used = next((i for i, v in enumerate(hist) if v > 0), 0)
    max_used = next((i for i in range(255, -1, -1) if hist[i] > 0), 0)
    dynamic_range = max(0, max_used - min_used)

    ranges = []
    for start in range(0, 256, 16):
        ranges.append({
            "range": f"{start}-{start + 15}",
            "count": sum(hist[start:start + 16]),
        })
    ranges.sort(key=lambda row: row["count"], reverse=True)

    return {
        "mean": mean,
        "median": float(median),
        "stddev": stddev,
        "skewness": skewness,
        "dominant_ranges": ranges[:3],
        "dynamic_range": dynamic_range,
    }


def normalize(hist: list[int]) -> list[float]:
    total = sum(hist)
    if total == 0:
        return [0.0] * len(hist)
    return [v / total for v in hist]


def compare_histograms(hist_a: list[int], hist_b: list[int]) -> dict:
    a = normalize(hist_a)
    b = normalize(hist_b)
    mean_a = sum(a) / len(a)
    mean_b = sum(b) / len(b)

    cov = 0.0
    var_a = 0.0
    var_b = 0.0
    chi2 = 0.0
    inter = 0.0
    bc = 0.0
    for i in range(256):
        da = a[i] - mean_a
        db = b[i] - mean_b
        cov += da * db
        var_a += da * da
        var_b += db * db
        chi2 += ((a[i] - b[i]) ** 2) / (a[i] + b[i] + 1e-12)
        inter += min(a[i], b[i])
        bc += (a[i] * b[i]) ** 0.5
    corr = cov / ((var_a * var_b) ** 0.5) if var_a > 0 and var_b > 0 else 0.0
    bhatta = max(0.0, 1.0 - bc) ** 0.5
    return {
        "correlation": corr,
        "chi_squared": chi2,
        "intersection": inter,
        "bhattacharyya": bhatta,
    }


def assessment(l_stats: dict, l_hist: list[int]) -> dict:
    total = max(1, sum(l_hist))
    clip_high = l_hist[255] / total
    clip_low = l_hist[0] / total
    if l_stats["mean"] > 190 and clip_high > 0.015:
        exposure = "overexposed"
    elif l_stats["mean"] < 65 and clip_low > 0.015:
        exposure = "underexposed"
    else:
        exposure = "normal"
    low_contrast = l_stats["dynamic_range"] < 80 or l_stats["stddev"] < 35
    return {
        "exposure": exposure,
        "low_contrast": low_contrast,
        "clipped_highlights_ratio": clip_high,
        "clipped_shadows_ratio": clip_low,
    }


def balance(stats: dict) -> dict:
    rm = stats["red"]["mean"]
    gm = stats["green"]["mean"]
    bm = stats["blue"]["mean"]
    diff = {"rg": abs(rm - gm), "rb": abs(rm - bm), "gb": abs(gm - bm)}
    balanced = diff["rg"] < 10 and diff["rb"] < 10 and diff["gb"] < 10
    dominant = max([("red", rm), ("green", gm), ("blue", bm)], key=lambda p: p[1])[0]
    return {"balanced": balanced, "dominant_channel": dominant, "mean_differences": diff}


def analyze_image(image_path: Path, args: argparse.Namespace) -> dict:
    if image_path.suffix.lower() not in SUPPORTED:
        raise ValueError(f"Unsupported image format: {image_path}")

    with Image.open(image_path) as original:
        crop = parse_crop(args.crop, original.width, original.height)
        x, y, w, h = crop
        image = original.crop((x, y, x + w, y + h))
        grayscale = image.mode in {"1", "L", "LA"}

        if grayscale:
            l_image = image.convert("L")
            lum = l_image.histogram()[:256]
            red = lum[:]
            green = lum[:]
            blue = lum[:]
        else:
            rgb = image.convert("RGB")
            red, green, blue = [channel.histogram()[:256] for channel in rgb.split()]
            lum = rgb.convert("L").histogram()[:256]

    stats = {
        "red": stats_from_hist(red),
        "green": stats_from_hist(green),
        "blue": stats_from_hist(blue),
        "luminance": stats_from_hist(lum),
    }
    assess = assessment(stats["luminance"], lum)
    bal = balance(stats)
    output = {
        "image": str(image_path.resolve()),
        "width": w,
        "height": h,
        "crop": {"x": x, "y": y, "width": w, "height": h},
        "grayscale": grayscale,
        "stats": stats,
        "assessment": assess,
        "color_balance": bal,
        "histograms": {"red": red, "green": green, "blue": blue, "luminance": lum},
    }
    if args.cumulative:
        output["cumulative_histograms"] = {
            "red": cumulative(red),
            "green": cumulative(green),
            "blue": cumulative(blue),
            "luminance": cumulative(lum),
        }
    return output


def print_summary(analysis: dict) -> None:
    def fmt(v: float) -> str:
        return f"{v:.3f}"

    print(f"Image: {analysis['image']}")
    print(f"Dimensions: {analysis['width']}x{analysis['height']}" + (" | Grayscale" if analysis["grayscale"] else ""))
    for channel in ("red", "green", "blue", "luminance"):
        s = analysis["stats"][channel]
        print(
            f"{channel.upper()}: mean={fmt(s['mean'])} median={fmt(s['median'])} "
            f"stddev={fmt(s['stddev'])} skew={fmt(s['skewness'])} dynamic={s['dynamic_range']}"
        )
    print(f"Exposure assessment: {analysis['assessment']['exposure']}")
    print("Contrast assessment: " + ("low contrast" if analysis["assessment"]["low_contrast"] else "normal contrast"))
    if analysis["color_balance"]["balanced"]:
        print("Color balance: balanced")
    else:
        print(f"Color balance: cast toward {analysis['color_balance']['dominant_channel']}")
    print()


def export_csv(path_value: Path, analysis: dict) -> None:
    path_value.parent.mkdir(parents=True, exist_ok=True)
    with path_value.open("w", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        writer.writerow(["bin", "red_count", "green_count", "blue_count", "luminance_count"])
        for i in range(256):
            writer.writerow([
                i,
                analysis["histograms"]["red"][i],
                analysis["histograms"]["green"][i],
                analysis["histograms"]["blue"][i],
                analysis["histograms"]["luminance"][i],
            ])


def write_json(path_value: Path, payload: dict) -> None:
    path_value.parent.mkdir(parents=True, exist_ok=True)
    path_value.write_text(json.dumps(payload, indent=2), encoding="utf-8")


def summary_table(analyses: list[dict]) -> None:
    print("Batch summary:")
    print("Image | Mean(L) | Std(L) | Exposure | Contrast")
    print("----- | ------- | ------ | -------- | --------")
    for analysis in analyses:
        lum = analysis["stats"]["luminance"]
        print(
            f"{Path(analysis['image']).name} | {lum['mean']:.2f} | {lum['stddev']:.2f} | "
            f"{analysis['assessment']['exposure']} | {'low' if analysis['assessment']['low_contrast'] else 'normal'}"
        )
    print()


def create_sample(kind: str, out_path: Path) -> None:
    width, height = 640, 420
    image = Image.new("RGB", (width, height))
    pixels = []
    for y in range(height):
        for x in range(width):
            if kind == "overexposed":
                base = 210 + ((x + y) % 45)
            elif kind == "underexposed":
                base = 15 + ((x + y) % 40)
            else:
                base = 65 + ((x * 3 + y * 2) % 130)
            noise = ((x * 17 + y * 31) % 23) - 11
            v = max(0, min(255, base + noise))
            pixels.append((
                v,
                max(0, min(255, v + (3 if kind == "balanced" else 0))),
                max(0, min(255, v - (3 if kind == "balanced" else 0))),
            ))
    image.putdata(pixels)
    out_path.parent.mkdir(parents=True, exist_ok=True)
    image.save(out_path, format="PNG")


def demo(args: argparse.Namespace, output_json: Path) -> int:
    sample_dir = Path("sample_histograms").resolve()
    over = sample_dir / "overexposed.png"
    under = sample_dir / "underexposed.png"
    balanced = sample_dir / "balanced.png"
    create_sample("overexposed", over)
    create_sample("underexposed", under)
    create_sample("balanced", balanced)
    analyses = [analyze_image(over, args), analyze_image(under, args), analyze_image(balanced, args)]
    for analysis in analyses:
        print_summary(analysis)
    summary_table(analyses)
    write_json(output_json, {"mode": "demo", "analyses": analyses})
    return 0


def iter_images(folder: Path) -> Iterable[Path]:
    for p in folder.iterdir():
        if p.is_file() and p.suffix.lower() in SUPPORTED:
            yield p.resolve()


def parser() -> argparse.ArgumentParser:
    p = argparse.ArgumentParser(description="Image Histogram Analyzer")
    p.add_argument("input", nargs="?")
    p.add_argument("--compare")
    p.add_argument("--cumulative", action="store_true")
    p.add_argument("--export")
    p.add_argument("--crop")
    p.add_argument("--output", default="histogram_analysis.json")
    p.add_argument("--batch", action="store_true")
    return p


def main() -> int:
    args = parser().parse_args()
    output_json = Path(args.output).resolve()

    if not args.input and not args.batch:
        return demo(args, output_json)
    if not args.input:
        raise ValueError("Input path required.")

    input_path = Path(args.input).resolve()
    if args.batch:
        if not input_path.is_dir():
            raise ValueError("--batch requires a directory input.")
        analyses = [analyze_image(p, args) for p in iter_images(input_path)]
        for analysis in analyses:
            print_summary(analysis)
        summary_table(analyses)
        write_json(output_json, {"mode": "batch", "analyses": analyses})
        return 0

    if not input_path.is_file():
        raise FileNotFoundError(f"Input file not found: {input_path}")
    analysis = analyze_image(input_path, args)
    print_summary(analysis)

    comparison = None
    if args.compare:
        other = analyze_image(Path(args.compare).resolve(), args)
        comparison = {
            "target_image": other["image"],
            "red": compare_histograms(analysis["histograms"]["red"], other["histograms"]["red"]),
            "green": compare_histograms(analysis["histograms"]["green"], other["histograms"]["green"]),
            "blue": compare_histograms(analysis["histograms"]["blue"], other["histograms"]["blue"]),
            "luminance": compare_histograms(analysis["histograms"]["luminance"], other["histograms"]["luminance"]),
        }
        print("Comparison metrics (luminance):")
        print(json.dumps(comparison["luminance"], indent=2))
        print()

    if args.export:
        export_csv(Path(args.export).resolve(), analysis)
    write_json(output_json, {"mode": "single", "analysis": analysis, "comparison": comparison})
    return 0


if __name__ == "__main__":
    try:
        raise SystemExit(main())
    except Exception as exc:
        print(f"Error: {exc}")
        raise SystemExit(1)