XML to JSON Converter (python, written by Gemini Code Assist)
envgap__gemini__python-t1-3
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
gemini/python-t1 #3 · read the task the agent was given
Gemini Code Assist wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: XML to JSON Converter Write a program that converts XML documents to JSON format while preserving the document structure including attributes, namespaces, CDATA sections, and mixed content. It should handle complex real-world XML with deeply nested elements. FUNCTIONAL REQUIREMENTS: - Accept an XML file path as a command-line argument - Parse the full XML document including attributes, namespaces, text content, CDATA sections, and comments - Convert to JSON preserving the hierarchy: elements become objects, repeated elements become arrays, attributes are prefixed with @ symbol, text content uses #text key - Handle namespace prefixes by including them in key names - Support converting multiple XML files in a directory via a --batch flag - Print the resulting JSON to console with pretty-print formatting (indented) - Save the output as a .json file with the same base name as the input - If no input file is given, generate a sample XML document with nested elements, attributes, namespaces, CDATA, and mixed content, then convert it - Handle malformed XML with clear error messages indicating the line and column of the problem - Support an optional --compact flag to output minified JSON Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# XML to JSON Converter (Python) A tool for converting XML documents to JSON while preserving structure, attributes, and namespaces. ## Setup Instructions 1. Ensure Python 3.10+ is installed. 2. Create a virtual environment: ```bash python3 -m venv venv source venv/bin/activate ``` 3. Install dependencies: ```bash pip install -r requirements.txt ``` ## Run Commands - Convert a single file: ```bash python xml_to_json.py data.xml ``` - Batch convert a directory: ```bash python xml_to_json.py ./xml_files --batch ``` - Compact (minified) output: ```bash python xml_to_json.py data.xml --compact ``` - Run with generated sample data: ```bash python xml_to_json.py ``` ## Output - Console: Pretty-printed JSON (unless --compact is used). - File: `.json` file with the same base name as the input.
requirements.txt
xmltodict==0.13.0
xml_to_json.py
import xmltodict
import json
import sys
import os
import argparse
from typing import Dict, Any
def generate_sample_xml(file_path: str):
print(f"Generating sample XML: {file_path}")
sample_content = """<?xml version="1.0" encoding="UTF-8"?>
<root xmlns:ns="http://example.com/ns">
<header id="101" version="1.0">
<ns:title>Sample Document</ns:title>
<date>2026-03-02</date>
</header>
<body>
<item category="tech">
<name>Laptop</name>
<description><![CDATA[A high-performance laptop with <b>Retina</b> display.]]></description>
<price currency="USD">1200</price>
</item>
<item category="science">
<name>Microscope</name>
<description>Advanced optical microscope.</description>
<price currency="EUR">850</price>
</item>
<mixed_content>
Some text
<inner>Inside tag</inner>
More text
</mixed_content>
</body>
<!-- This is a comment -->
</root>
"""
with open(file_path, "w") as f:
f.write(sample_content)
def convert_file(input_path: str, compact: bool):
try:
with open(input_path, 'r') as f:
xml_content = f.read()
# xmltodict preserves attributes (@), namespaces, CDATA, and repetition
data = xmltodict.parse(xml_content, process_namespaces=False, xml_attribs=True)
indent = None if compact else 4
json_output = json.dumps(data, indent=indent)
base_name = os.path.splitext(input_path)[0]
output_path = f"{base_name}.json"
with open(output_path, 'w') as f:
f.write(json_output)
print(f"Successfully converted {input_path} to {output_path}")
if not compact:
print(json_output)
except Exception as e:
print(f"Error converting {input_path}: {e}")
def main():
parser = argparse.ArgumentParser(description="XML to JSON Converter")
parser.add_argument("input", nargs="?", help="Input XML file or directory")
parser.add_argument("--batch", action="store_true", help="Batch process directory")
parser.add_argument("--compact", action="store_true", help="Minified JSON output")
args = parser.parse_args()
if not args.input:
sample_file = "sample.xml"
if not os.path.exists(sample_file):
generate_sample_xml(sample_file)
args.input = sample_file
if args.batch:
if not os.path.isdir(args.input):
print(f"Error: {args.input} is not a directory")
return
for file in os.listdir(args.input):
if file.endswith(".xml"):
convert_file(os.path.join(args.input, file), args.compact)
else:
convert_file(args.input, args.compact)
if __name__ == "__main__":
main()