JSON Schema Validator (python, written by Codex)
envgap__codex__python-t1-2
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/python-t1 #2 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: JSON Schema Validator Write a program that validates JSON documents against user-defined schema rules. It should support type checking, required fields, nested object validation, array constraints, and pattern matching. It should collect all violations — not stop at the first — and produce a detailed validation report. FUNCTIONAL REQUIREMENTS: - Accept two command-line arguments: a JSON data file and a JSON schema file - Support validation rules: type checking (string, number, boolean, array, object, null), required fields, min/max for numbers, minLength/maxLength for strings, enum values, regex pattern matching - Validate nested objects and arrays recursively to any depth - Validate array items (all items must conform to a specified type or sub-schema) - Collect ALL validation errors across the entire document, reporting the full JSON path for each violation (e.g., $.users[2].email) - Print the validation report to the console with one line per violation showing path, expected rule, and actual value - Save the validation report as validation_report.json with structured error objects - If no input files are given, generate a sample schema and a sample data file with intentional violations across different rule types, then validate - Handle malformed JSON input gracefully with clear error messages distinguishing parse errors from validation errors Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# JSON Schema Validator (Python) CLI tool for validating JSON data against user-defined schema rules with recursive traversal and full violation collection. ## Features - Accepts `data.json` and `schema.json` command-line arguments. - Supports: - `type` (`string`, `number`, `boolean`, `array`, `object`, `null`) - `required` - `min` / `max` (`minimum` / `maximum` also supported) - `minLength` / `maxLength` - `enum` - `pattern` (regex) - `minItems` / `maxItems` - Validates nested objects and arrays recursively to any depth. - Validates all array items against `items` sub-schema. - Collects all violations with JSON paths such as `$.users[2].email`. - Prints one line per violation to the console. - Saves structured `validation_report.json`. - With no input arguments, generates `sample_data.json` and `sample_schema.json` containing intentional violations, then validates. - Distinguishes parse errors from validation errors. ## Requirements - Ubuntu 22.04 - Python 3.10+ ## Dependencies (Pinned) No external dependencies are used. - Direct dependencies: none - Transitive dependencies: none `requirements.txt` is included and explicitly documents zero dependencies. ## Setup ```bash cd TMLR/code_generation/codex_generated/p_02/python python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` ## Run Validate your files: ```bash python3 src/main.py ./data.json ./schema.json ``` Generate sample files and validate: ```bash python3 src/main.py ``` ## Output - Console: one line per violation with path, expected rule, and actual value. - File: `validation_report.json` with structured error objects. - If no inputs: `sample_data.json` and `sample_schema.json` are generated.
requirements.txt
# No external dependencies required. # Direct dependencies: none # Transitive dependencies: none
src/main.py
#!/usr/bin/env python3
import json
import os
import re
import sys
from datetime import datetime, timezone
def json_type_of(value):
if value is None:
return "null"
if isinstance(value, bool):
return "boolean"
if isinstance(value, list):
return "array"
if isinstance(value, dict):
return "object"
if isinstance(value, str):
return "string"
if isinstance(value, (int, float)):
return "number"
return type(value).__name__
def value_preview(value):
if value is _MISSING:
return "undefined"
if isinstance(value, (dict, list)):
if isinstance(value, list):
return f"<array length={len(value)}>"
return "<object>"
return json.dumps(value)
def matches_type(value, expected):
return json_type_of(value) == expected
def deep_equal(left, right):
return left == right
def add_error(errors, path_text, rule, expected, actual):
errors.append(
{
"path": path_text,
"rule": rule,
"expected": expected,
"actual": value_preview(actual),
}
)
def validate(data, schema, path_text, errors):
if not isinstance(schema, dict):
return
if "type" in schema:
declared = schema["type"]
allowed_types = declared if isinstance(declared, list) else [declared]
if not any(matches_type(data, item) for item in allowed_types):
expected = " | ".join(str(item) for item in allowed_types)
add_error(errors, path_text, "type", expected, data)
return
if "enum" in schema and isinstance(schema["enum"], list):
if not any(deep_equal(data, allowed) for allowed in schema["enum"]):
add_error(errors, path_text, "enum", json.dumps(schema["enum"]), data)
data_type = json_type_of(data)
if data_type == "number":
if isinstance(schema.get("min"), (int, float)) and data < schema["min"]:
add_error(errors, path_text, "min", f">= {schema['min']}", data)
if isinstance(schema.get("max"), (int, float)) and data > schema["max"]:
add_error(errors, path_text, "max", f"<= {schema['max']}", data)
if isinstance(schema.get("minimum"), (int, float)) and data < schema["minimum"]:
add_error(errors, path_text, "minimum", f">= {schema['minimum']}", data)
if isinstance(schema.get("maximum"), (int, float)) and data > schema["maximum"]:
add_error(errors, path_text, "maximum", f"<= {schema['maximum']}", data)
if data_type == "string":
if isinstance(schema.get("minLength"), int) and len(data) < schema["minLength"]:
add_error(errors, path_text, "minLength", f"length >= {schema['minLength']}", data)
if isinstance(schema.get("maxLength"), int) and len(data) > schema["maxLength"]:
add_error(errors, path_text, "maxLength", f"length <= {schema['maxLength']}", data)
if isinstance(schema.get("pattern"), str):
pattern = schema["pattern"]
try:
if re.search(pattern, data) is None:
add_error(errors, path_text, "pattern", f"match /{pattern}/", data)
except re.error as exc:
add_error(errors, path_text, "pattern", f'invalid regex "{pattern}": {exc}', data)
if data_type == "object":
required = schema.get("required")
if isinstance(required, list):
for key in required:
if isinstance(key, str) and key not in data:
add_error(errors, f"{path_text}.{key}", "required", "field must exist", _MISSING)
properties = schema.get("properties")
if isinstance(properties, dict):
for key, sub_schema in properties.items():
if key in data:
validate(data[key], sub_schema, f"{path_text}.{key}", errors)
if data_type == "array":
if isinstance(schema.get("minItems"), int) and len(data) < schema["minItems"]:
add_error(errors, path_text, "minItems", f"items >= {schema['minItems']}", data)
if isinstance(schema.get("maxItems"), int) and len(data) > schema["maxItems"]:
add_error(errors, path_text, "maxItems", f"items <= {schema['maxItems']}", data)
item_schema = schema.get("items")
if isinstance(item_schema, dict):
for index, item in enumerate(data):
validate(item, item_schema, f"{path_text}[{index}]", errors)
def generate_sample_schema():
return {
"type": "object",
"required": ["users", "status", "metadata", "note"],
"properties": {
"users": {
"type": "array",
"minItems": 1,
"items": {
"type": "object",
"required": ["name", "age", "email", "tags", "active"],
"properties": {
"name": {"type": "string", "minLength": 3, "maxLength": 20},
"age": {"type": "number", "min": 0, "max": 120},
"email": {
"type": "string",
"pattern": r"^[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}$",
},
"tags": {
"type": "array",
"minItems": 1,
"items": {"type": "string", "minLength": 2, "maxLength": 12},
},
"active": {"type": "boolean"},
},
},
},
"status": {"type": "string", "enum": ["new", "active", "disabled"]},
"metadata": {
"type": "object",
"required": ["version", "source"],
"properties": {
"version": {"type": "string", "pattern": r"^\d+\.\d+\.\d+$"},
"source": {
"type": "object",
"required": ["systemId"],
"properties": {"systemId": {"type": "string", "minLength": 5}},
},
},
},
"note": {"type": "null"},
},
}
def generate_sample_data():
return {
"users": [
{
"name": "Alice",
"age": 29,
"email": "alice@example.com",
"tags": ["owner", "ops"],
"active": True,
},
{
"name": "Al",
"age": -5,
"email": "invalid-email",
"tags": ["ok", 5, ""],
"active": "true",
},
{
"name": 123,
"age": 200,
"tags": [],
"active": False,
},
],
"status": "archived",
"metadata": {"version": "1.0", "source": {"systemId": 42}},
"note": "should be null",
}
def load_json(path_value, label):
if not os.path.exists(path_value):
raise RuntimeError(f"{label} file read error ({path_value}): file not found")
try:
with open(path_value, "r", encoding="utf-8") as handle:
return json.load(handle)
except json.JSONDecodeError as exc:
raise RuntimeError(
f"{label} JSON parse error ({path_value}): line {exc.lineno}, col {exc.colno}: {exc.msg}"
) from exc
except OSError as exc:
raise RuntimeError(f"{label} file read error ({path_value}): {exc}") from exc
def print_report(errors):
if not errors:
print("Validation succeeded: no violations found.")
return
print(f"Validation failed: {len(errors)} violation(s) found.")
for item in errors:
print(
f"{item['path']} | expected: {item['rule']} -> {item['expected']} | actual: {item['actual']}"
)
def main():
args = sys.argv[1:]
if len(args) not in (0, 2):
print("Usage: python3 src/main.py <data.json> <schema.json>")
print(" or: python3 src/main.py")
return 1
if len(args) == 0:
data_path = os.path.abspath("sample_data.json")
schema_path = os.path.abspath("sample_schema.json")
with open(data_path, "w", encoding="utf-8") as handle:
json.dump(generate_sample_data(), handle, indent=2)
handle.write("\n")
with open(schema_path, "w", encoding="utf-8") as handle:
json.dump(generate_sample_schema(), handle, indent=2)
handle.write("\n")
print(f"No input files provided. Generated sample data: {data_path}")
print(f"No input files provided. Generated sample schema: {schema_path}")
else:
data_path = os.path.abspath(args[0])
schema_path = os.path.abspath(args[1])
try:
data = load_json(data_path, "Data")
schema = load_json(schema_path, "Schema")
except RuntimeError as exc:
print(exc)
return 1
errors = []
validate(data, schema, "$", errors)
print_report(errors)
report = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"dataFile": data_path,
"schemaFile": schema_path,
"valid": len(errors) == 0,
"errorCount": len(errors),
"errors": errors,
}
report_path = os.path.abspath("validation_report.json")
with open(report_path, "w", encoding="utf-8") as handle:
json.dump(report, handle, indent=2)
handle.write("\n")
print(f"Saved structured report: {report_path}")
return 0
_MISSING = object()
if __name__ == "__main__":
raise SystemExit(main())