← All tasks
pythoncodex/python-t1 #2Not a task: already works

JSON Schema Validator (python, written by Codex)

envgap__codex__python-t1-2

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/python-t1 #2 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: JSON Schema Validator

Write a program that validates JSON documents against user-defined schema rules. It should support type checking, required fields, nested object validation, array constraints, and pattern matching. It should collect all violations — not stop at the first — and produce a detailed validation report.

FUNCTIONAL REQUIREMENTS:
- Accept two command-line arguments: a JSON data file and a JSON schema file
- Support validation rules: type checking (string, number, boolean, array, object, null), required fields, min/max for numbers, minLength/maxLength for strings, enum values, regex pattern matching
- Validate nested objects and arrays recursively to any depth
- Validate array items (all items must conform to a specified type or sub-schema)
- Collect ALL validation errors across the entire document, reporting the full JSON path for each violation (e.g., $.users[2].email)
- Print the validation report to the console with one line per violation showing path, expected rule, and actual value
- Save the validation report as validation_report.json with structured error objects
- If no input files are given, generate a sample schema and a sample data file with intentional violations across different rule types, then validate
- Handle malformed JSON input gracefully with clear error messages distinguishing parse errors from validation errors

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

README.md
# JSON Schema Validator (Python)

CLI tool for validating JSON data against user-defined schema rules with recursive traversal and full violation collection.

## Features

- Accepts `data.json` and `schema.json` command-line arguments.
- Supports:
  - `type` (`string`, `number`, `boolean`, `array`, `object`, `null`)
  - `required`
  - `min` / `max` (`minimum` / `maximum` also supported)
  - `minLength` / `maxLength`
  - `enum`
  - `pattern` (regex)
  - `minItems` / `maxItems`
- Validates nested objects and arrays recursively to any depth.
- Validates all array items against `items` sub-schema.
- Collects all violations with JSON paths such as `$.users[2].email`.
- Prints one line per violation to the console.
- Saves structured `validation_report.json`.
- With no input arguments, generates `sample_data.json` and `sample_schema.json` containing intentional violations, then validates.
- Distinguishes parse errors from validation errors.

## Requirements

- Ubuntu 22.04
- Python 3.10+

## Dependencies (Pinned)

No external dependencies are used.

- Direct dependencies: none
- Transitive dependencies: none

`requirements.txt` is included and explicitly documents zero dependencies.

## Setup

```bash
cd TMLR/code_generation/codex_generated/p_02/python
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

## Run

Validate your files:

```bash
python3 src/main.py ./data.json ./schema.json
```

Generate sample files and validate:

```bash
python3 src/main.py
```

## Output

- Console: one line per violation with path, expected rule, and actual value.
- File: `validation_report.json` with structured error objects.
- If no inputs: `sample_data.json` and `sample_schema.json` are generated.
requirements.txt
# No external dependencies required.
# Direct dependencies: none
# Transitive dependencies: none
src/main.py
#!/usr/bin/env python3
import json
import os
import re
import sys
from datetime import datetime, timezone


def json_type_of(value):
    if value is None:
        return "null"
    if isinstance(value, bool):
        return "boolean"
    if isinstance(value, list):
        return "array"
    if isinstance(value, dict):
        return "object"
    if isinstance(value, str):
        return "string"
    if isinstance(value, (int, float)):
        return "number"
    return type(value).__name__


def value_preview(value):
    if value is _MISSING:
        return "undefined"
    if isinstance(value, (dict, list)):
        if isinstance(value, list):
            return f"<array length={len(value)}>"
        return "<object>"
    return json.dumps(value)


def matches_type(value, expected):
    return json_type_of(value) == expected


def deep_equal(left, right):
    return left == right


def add_error(errors, path_text, rule, expected, actual):
    errors.append(
        {
            "path": path_text,
            "rule": rule,
            "expected": expected,
            "actual": value_preview(actual),
        }
    )


def validate(data, schema, path_text, errors):
    if not isinstance(schema, dict):
        return

    if "type" in schema:
        declared = schema["type"]
        allowed_types = declared if isinstance(declared, list) else [declared]
        if not any(matches_type(data, item) for item in allowed_types):
            expected = " | ".join(str(item) for item in allowed_types)
            add_error(errors, path_text, "type", expected, data)
            return

    if "enum" in schema and isinstance(schema["enum"], list):
        if not any(deep_equal(data, allowed) for allowed in schema["enum"]):
            add_error(errors, path_text, "enum", json.dumps(schema["enum"]), data)

    data_type = json_type_of(data)

    if data_type == "number":
        if isinstance(schema.get("min"), (int, float)) and data < schema["min"]:
            add_error(errors, path_text, "min", f">= {schema['min']}", data)
        if isinstance(schema.get("max"), (int, float)) and data > schema["max"]:
            add_error(errors, path_text, "max", f"<= {schema['max']}", data)
        if isinstance(schema.get("minimum"), (int, float)) and data < schema["minimum"]:
            add_error(errors, path_text, "minimum", f">= {schema['minimum']}", data)
        if isinstance(schema.get("maximum"), (int, float)) and data > schema["maximum"]:
            add_error(errors, path_text, "maximum", f"<= {schema['maximum']}", data)

    if data_type == "string":
        if isinstance(schema.get("minLength"), int) and len(data) < schema["minLength"]:
            add_error(errors, path_text, "minLength", f"length >= {schema['minLength']}", data)
        if isinstance(schema.get("maxLength"), int) and len(data) > schema["maxLength"]:
            add_error(errors, path_text, "maxLength", f"length <= {schema['maxLength']}", data)
        if isinstance(schema.get("pattern"), str):
            pattern = schema["pattern"]
            try:
                if re.search(pattern, data) is None:
                    add_error(errors, path_text, "pattern", f"match /{pattern}/", data)
            except re.error as exc:
                add_error(errors, path_text, "pattern", f'invalid regex "{pattern}": {exc}', data)

    if data_type == "object":
        required = schema.get("required")
        if isinstance(required, list):
            for key in required:
                if isinstance(key, str) and key not in data:
                    add_error(errors, f"{path_text}.{key}", "required", "field must exist", _MISSING)

        properties = schema.get("properties")
        if isinstance(properties, dict):
            for key, sub_schema in properties.items():
                if key in data:
                    validate(data[key], sub_schema, f"{path_text}.{key}", errors)

    if data_type == "array":
        if isinstance(schema.get("minItems"), int) and len(data) < schema["minItems"]:
            add_error(errors, path_text, "minItems", f"items >= {schema['minItems']}", data)
        if isinstance(schema.get("maxItems"), int) and len(data) > schema["maxItems"]:
            add_error(errors, path_text, "maxItems", f"items <= {schema['maxItems']}", data)
        item_schema = schema.get("items")
        if isinstance(item_schema, dict):
            for index, item in enumerate(data):
                validate(item, item_schema, f"{path_text}[{index}]", errors)


def generate_sample_schema():
    return {
        "type": "object",
        "required": ["users", "status", "metadata", "note"],
        "properties": {
            "users": {
                "type": "array",
                "minItems": 1,
                "items": {
                    "type": "object",
                    "required": ["name", "age", "email", "tags", "active"],
                    "properties": {
                        "name": {"type": "string", "minLength": 3, "maxLength": 20},
                        "age": {"type": "number", "min": 0, "max": 120},
                        "email": {
                            "type": "string",
                            "pattern": r"^[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}$",
                        },
                        "tags": {
                            "type": "array",
                            "minItems": 1,
                            "items": {"type": "string", "minLength": 2, "maxLength": 12},
                        },
                        "active": {"type": "boolean"},
                    },
                },
            },
            "status": {"type": "string", "enum": ["new", "active", "disabled"]},
            "metadata": {
                "type": "object",
                "required": ["version", "source"],
                "properties": {
                    "version": {"type": "string", "pattern": r"^\d+\.\d+\.\d+$"},
                    "source": {
                        "type": "object",
                        "required": ["systemId"],
                        "properties": {"systemId": {"type": "string", "minLength": 5}},
                    },
                },
            },
            "note": {"type": "null"},
        },
    }


def generate_sample_data():
    return {
        "users": [
            {
                "name": "Alice",
                "age": 29,
                "email": "alice@example.com",
                "tags": ["owner", "ops"],
                "active": True,
            },
            {
                "name": "Al",
                "age": -5,
                "email": "invalid-email",
                "tags": ["ok", 5, ""],
                "active": "true",
            },
            {
                "name": 123,
                "age": 200,
                "tags": [],
                "active": False,
            },
        ],
        "status": "archived",
        "metadata": {"version": "1.0", "source": {"systemId": 42}},
        "note": "should be null",
    }


def load_json(path_value, label):
    if not os.path.exists(path_value):
        raise RuntimeError(f"{label} file read error ({path_value}): file not found")
    try:
        with open(path_value, "r", encoding="utf-8") as handle:
            return json.load(handle)
    except json.JSONDecodeError as exc:
        raise RuntimeError(
            f"{label} JSON parse error ({path_value}): line {exc.lineno}, col {exc.colno}: {exc.msg}"
        ) from exc
    except OSError as exc:
        raise RuntimeError(f"{label} file read error ({path_value}): {exc}") from exc


def print_report(errors):
    if not errors:
        print("Validation succeeded: no violations found.")
        return
    print(f"Validation failed: {len(errors)} violation(s) found.")
    for item in errors:
        print(
            f"{item['path']} | expected: {item['rule']} -> {item['expected']} | actual: {item['actual']}"
        )


def main():
    args = sys.argv[1:]
    if len(args) not in (0, 2):
        print("Usage: python3 src/main.py <data.json> <schema.json>")
        print("   or: python3 src/main.py")
        return 1

    if len(args) == 0:
        data_path = os.path.abspath("sample_data.json")
        schema_path = os.path.abspath("sample_schema.json")
        with open(data_path, "w", encoding="utf-8") as handle:
            json.dump(generate_sample_data(), handle, indent=2)
            handle.write("\n")
        with open(schema_path, "w", encoding="utf-8") as handle:
            json.dump(generate_sample_schema(), handle, indent=2)
            handle.write("\n")
        print(f"No input files provided. Generated sample data: {data_path}")
        print(f"No input files provided. Generated sample schema: {schema_path}")
    else:
        data_path = os.path.abspath(args[0])
        schema_path = os.path.abspath(args[1])

    try:
        data = load_json(data_path, "Data")
        schema = load_json(schema_path, "Schema")
    except RuntimeError as exc:
        print(exc)
        return 1

    errors = []
    validate(data, schema, "$", errors)
    print_report(errors)

    report = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        "dataFile": data_path,
        "schemaFile": schema_path,
        "valid": len(errors) == 0,
        "errorCount": len(errors),
        "errors": errors,
    }
    report_path = os.path.abspath("validation_report.json")
    with open(report_path, "w", encoding="utf-8") as handle:
        json.dump(report, handle, indent=2)
        handle.write("\n")
    print(f"Saved structured report: {report_path}")
    return 0


_MISSING = object()

if __name__ == "__main__":
    raise SystemExit(main())