YAML Config Merger (python, written by Codex)
envgap__codex__python-t1-4
Written by a coding agent; not on GitHubWritten 2026-03-02
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/python-t1 #4 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.
Task given to the agent:
TASK: YAML Config Merger
Write a program that merges multiple YAML configuration files into a single unified configuration, supporting deep merging of nested structures, array handling strategies, and conflict resolution.
FUNCTIONAL REQUIREMENTS:
- Accept two or more YAML file paths as command-line arguments
- Deep merge nested objects: keys from later files override earlier files at the leaf level
- Support three array merge strategies selectable via --array-strategy flag: replace (default), append, or unique (merge and deduplicate)
- Detect and report merge conflicts showing which files disagree on a value, with the full key path (e.g., database.connection.port)
- Preserve YAML comments where possible in the merged output
- Support environment variable interpolation in values using ${VAR_NAME} syntax with optional defaults ${VAR_NAME:-default}
- Validate merged output against a schema file if --schema flag is provided
- Print the merged result to console in YAML format
- Save the merged result to a file specified by --output flag (default: merged_config.yaml)
- If no input files are given, generate three sample YAML config files with overlapping keys, nested structures, and arrays, then merge them
- Handle malformed YAML with clear error messages identifying the file and location of the problem
Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# YAML Config Merger (Python)
Merges multiple YAML config files into one with deep merge, array strategies, conflict reporting, env interpolation, optional schema validation, and YAML output.
## Features
- Accepts 2+ YAML file paths.
- Deep merge nested objects; later files override leaf values.
- `--array-strategy`: `replace` (default), `append`, `unique`.
- Conflict reporting with full key paths and disagreeing files.
- Preserves YAML comments where possible.
- Environment interpolation:
- `${VAR_NAME}`
- `${VAR_NAME:-default}`
- Optional schema validation via `--schema`.
- Prints merged YAML.
- Writes output file via `--output` (default `merged_config.yaml`).
- No args: generates three sample YAML files, then merges.
- Malformed YAML errors include file and line.
## Requirements
- Ubuntu 22.04
- Python 3.10+
## Dependencies (Pinned)
No external dependencies are used.
- Direct dependencies: none
- Transitive dependencies: none
## Run
```bash
cd TMLR/code_generation/codex_generated/p_04/python
python3 src/main.py config1.yaml config2.yaml config3.yaml
```
```bash
python3 src/main.py --array-strategy unique --output merged.yaml config1.yaml config2.yaml
```
```bash
python3 src/main.py --schema schema.yaml config1.yaml config2.yaml
```
```bash
python3 src/main.py
```
requirements.txt
# No external dependencies required. # Direct dependencies: none # Transitive dependencies: none
src/main.py
#!/usr/bin/env python3
import json
import os
import re
import sys
class YamlParseError(Exception):
def __init__(self, message, line):
super().__init__(f"{message} at line {line}")
self.line = line
def deep_equal(a, b):
return a == b
def deep_clone(value):
return json.loads(json.dumps(value))
def scalar_from_text(text):
if text in ("null", "~"):
return None
if text == "true":
return True
if text == "false":
return False
if re.fullmatch(r"-?\d+(\.\d+)?", text):
return float(text) if "." in text else int(text)
if (text.startswith('"') and text.endswith('"')) or (text.startswith("'") and text.endswith("'")):
return text[1:-1]
return text
def yaml_scalar(value):
if value is None:
return "null"
if isinstance(value, bool):
return "true" if value else "false"
if isinstance(value, (int, float)):
return str(value)
if isinstance(value, str):
if value == "":
return '""'
if re.search(r'[:#\-\[\]\{\},]|\A\s|\s\Z|\n', value):
return json.dumps(value)
return value
return json.dumps(value)
def parse_yaml(text):
lines = text.replace("\r\n", "\n").replace("\r", "\n").split("\n")
comments = {}
def line_info(index):
raw = lines[index] if index < len(lines) else ""
indent = len(raw) - len(raw.lstrip(" "))
trimmed = raw.strip()
return raw, indent, trimmed, index + 1
def parse_block(start, indent, path_prefix):
index = start
pending_comments = []
while index < len(lines):
_, cur_indent, trimmed, line_no = line_info(index)
if trimmed == "":
index += 1
continue
if cur_indent < indent:
break
if cur_indent > indent:
raise YamlParseError("Unexpected indentation", line_no)
if trimmed.startswith("#"):
pending_comments.append(trimmed[1:].strip())
index += 1
continue
break
if index >= len(lines):
return None, index
_, first_indent, first_trim, _ = line_info(index)
if first_indent != indent:
return None, index
is_seq = first_trim.startswith("- ")
if is_seq:
arr = []
while index < len(lines):
_, cur_indent, trimmed, line_no = line_info(index)
if trimmed == "" or trimmed.startswith("#"):
index += 1
continue
if cur_indent < indent:
break
if cur_indent > indent:
raise YamlParseError("Unexpected indentation", line_no)
if not trimmed.startswith("- "):
break
rest = trimmed[2:].strip()
if rest == "":
nested, next_index = parse_block(index + 1, indent + 2, f"{path_prefix}[{len(arr)}]")
arr.append(nested)
index = next_index
elif re.match(r"^[^:#]+:\s*", rest):
split = rest.find(":")
key = rest[:split].strip()
value_text = rest[split + 1 :].strip()
item = {}
if value_text == "":
nested, next_index = parse_block(
index + 1, indent + 4, f"{path_prefix}[{len(arr)}].{key}"
)
item[key] = nested
arr.append(item)
index = next_index
else:
item[key] = scalar_from_text(value_text)
arr.append(item)
index += 1
else:
arr.append(scalar_from_text(rest))
index += 1
return arr, index
obj = {}
local_comments = list(pending_comments)
while index < len(lines):
_, cur_indent, trimmed, line_no = line_info(index)
if trimmed == "":
index += 1
continue
if cur_indent < indent:
break
if cur_indent > indent:
raise YamlParseError("Unexpected indentation", line_no)
if trimmed.startswith("#"):
local_comments.append(trimmed[1:].strip())
index += 1
continue
split = trimmed.find(":")
if split < 1:
raise YamlParseError("Invalid mapping entry", line_no)
key = trimmed[:split].strip()
rest = trimmed[split + 1 :].strip()
full_path = f"{path_prefix}.{key}" if path_prefix else key
if local_comments:
comments[full_path] = list(local_comments)
local_comments = []
if rest == "":
nested, next_index = parse_block(index + 1, indent + 2, full_path)
obj[key] = nested
index = next_index
else:
inline = re.split(r"\s+#", rest, maxsplit=1)[0].strip()
obj[key] = scalar_from_text(inline)
index += 1
return obj, index
value, _ = parse_block(0, 0, "")
return (value if value is not None else {}), comments
def interpolate_env(value):
if isinstance(value, str):
pattern = re.compile(r"\$\{([A-Za-z_][A-Za-z0-9_]*)(:-([^}]*))?\}")
def repl(match):
key = match.group(1)
default = match.group(3)
if key in os.environ:
return os.environ[key]
return default if default is not None else ""
return pattern.sub(repl, value)
if isinstance(value, list):
return [interpolate_env(v) for v in value]
if isinstance(value, dict):
return {k: interpolate_env(v) for k, v in value.items()}
return value
def merge_values(base, incoming, path_text, strategy, conflicts, source_map, current_file):
if isinstance(base, dict) and isinstance(incoming, dict):
merged = dict(base)
for key, val in incoming.items():
child_path = f"{path_text}.{key}" if path_text else key
if key in merged:
merged[key] = merge_values(
merged[key], val, child_path, strategy, conflicts, source_map, current_file
)
else:
merged[key] = deep_clone(val)
source_map[child_path] = current_file
return merged
if isinstance(base, list) and isinstance(incoming, list):
if strategy == "append":
merged = [deep_clone(v) for v in base] + [deep_clone(v) for v in incoming]
elif strategy == "unique":
merged = [deep_clone(v) for v in base]
for item in incoming:
if not any(deep_equal(x, item) for x in merged):
merged.append(deep_clone(item))
else:
merged = [deep_clone(v) for v in incoming]
if not deep_equal(base, incoming):
conflicts.append(
{
"path": path_text,
"previousFile": source_map.get(path_text, "unknown"),
"currentFile": current_file,
"previousValue": base,
"currentValue": incoming,
}
)
source_map[path_text] = current_file
return merged
if not deep_equal(base, incoming):
conflicts.append(
{
"path": path_text,
"previousFile": source_map.get(path_text, "unknown"),
"currentFile": current_file,
"previousValue": base,
"currentValue": incoming,
}
)
source_map[path_text] = current_file
return deep_clone(incoming)
def validate_schema(data, schema, path_text, errors):
if not isinstance(schema, dict):
return
declared_type = schema.get("type")
actual_type = (
"null"
if data is None
else "array"
if isinstance(data, list)
else "object"
if isinstance(data, dict)
else "boolean"
if isinstance(data, bool)
else "number"
if isinstance(data, (int, float))
else "string"
if isinstance(data, str)
else "unknown"
)
if isinstance(declared_type, str) and declared_type != actual_type:
errors.append(f"{path_text or '$'}: expected type {declared_type}, got {actual_type}")
return
enum_values = schema.get("enum")
if isinstance(enum_values, list) and not any(deep_equal(v, data) for v in enum_values):
errors.append(f"{path_text or '$'}: value not in enum")
if actual_type == "object":
req = schema.get("required")
if isinstance(req, list):
for key in req:
if key not in data:
errors.append(f"{path_text or '$'}: missing required key {key}")
props = schema.get("properties")
if isinstance(props, dict):
for key, child_schema in props.items():
if key in data:
child_path = f"{path_text}.{key}" if path_text else key
validate_schema(data[key], child_schema, child_path, errors)
elif actual_type == "array":
item_schema = schema.get("items")
if item_schema is not None:
for idx, item in enumerate(data):
validate_schema(item, item_schema, f"{path_text}[{idx}]", errors)
def to_yaml(value, comments_map, path_text="", indent=0):
pad = " " * indent
lines = []
if isinstance(value, list):
for idx, item in enumerate(value):
if isinstance(item, dict) or isinstance(item, list):
lines.append(f"{pad}-")
lines.append(to_yaml(item, comments_map, f"{path_text}[{idx}]", indent + 2))
else:
lines.append(f"{pad}- {yaml_scalar(item)}")
return "\n".join(line for line in lines if line != "")
if isinstance(value, dict):
for key, val in value.items():
child_path = f"{path_text}.{key}" if path_text else key
if child_path in comments_map:
for comment in comments_map[child_path]:
lines.append(f"{pad}# {comment}")
if isinstance(val, (dict, list)):
lines.append(f"{pad}{key}:")
lines.append(to_yaml(val, comments_map, child_path, indent + 2))
else:
lines.append(f"{pad}{key}: {yaml_scalar(val)}")
return "\n".join(line for line in lines if line != "")
return f"{pad}{yaml_scalar(value)}"
def load_yaml_file(file_path):
try:
with open(file_path, "r", encoding="utf-8") as handle:
text = handle.read()
except OSError as exc:
raise RuntimeError(f"Failed to read {file_path}: {exc}") from exc
try:
data, comments = parse_yaml(text)
return {"data": data, "comments": comments}
except YamlParseError as exc:
raise RuntimeError(f"Malformed YAML in {file_path}: {exc}") from exc
def sample_files():
return {
"config_1.yaml": """# Base app config
app:
name: merger-demo
env: ${APP_ENV:-dev}
database:
host: localhost
port: 5432
tags:
- core
- primary
features:
enabled:
- auth
- api
""",
"config_2.yaml": """# Override database and features
database:
port: 5433
user: ${DB_USER:-admin}
tags:
- analytics
features:
enabled:
- api
- billing
""",
"config_3.yaml": """# Production tuning
app:
env: prod
database:
host: db.internal
retries: 5
features:
enabled:
- auth
- billing
""",
}
def main():
args = sys.argv[1:]
files = []
array_strategy = "replace"
schema_path = None
output_path = "merged_config.yaml"
i = 0
while i < len(args):
arg = args[i]
if arg == "--array-strategy":
i += 1
array_strategy = args[i] if i < len(args) else ""
elif arg == "--schema":
i += 1
schema_path = args[i] if i < len(args) else None
elif arg == "--output":
i += 1
output_path = args[i] if i < len(args) else output_path
else:
files.append(arg)
i += 1
if array_strategy not in ("replace", "append", "unique"):
print("Invalid --array-strategy. Use replace, append, or unique.")
return 1
if not files:
generated = sample_files()
files = []
for name, content in generated.items():
abs_path = os.path.abspath(name)
with open(abs_path, "w", encoding="utf-8") as handle:
handle.write(content)
files.append(abs_path)
print(f"No input files provided. Generated sample files: {', '.join(files)}")
if len(files) < 2:
print("Provide at least two YAML files.")
return 1
parsed_files = []
for file_name in files:
abs_file = os.path.abspath(file_name)
try:
parsed = load_yaml_file(abs_file)
except RuntimeError as exc:
print(exc)
return 1
parsed_files.append({"file": abs_file, **parsed})
merged = deep_clone(parsed_files[0]["data"])
source_map = {}
def collect_leaf_paths(value, path_prefix):
if isinstance(value, dict):
for key, child in value.items():
child_path = f"{path_prefix}.{key}" if path_prefix else key
collect_leaf_paths(child, child_path)
else:
source_map[path_prefix] = parsed_files[0]["file"]
collect_leaf_paths(merged, "")
merged_comments = dict(parsed_files[0]["comments"])
conflicts = []
for parsed in parsed_files[1:]:
merged = merge_values(
merged,
parsed["data"],
"",
array_strategy,
conflicts,
source_map,
parsed["file"],
)
for key, comments in parsed["comments"].items():
if key not in merged_comments:
merged_comments[key] = comments
merged = interpolate_env(merged)
if schema_path:
try:
schema = load_yaml_file(os.path.abspath(schema_path))["data"]
except RuntimeError as exc:
print(exc)
return 1
schema_errors = []
validate_schema(merged, schema, "", schema_errors)
if schema_errors:
print("Schema validation failed:")
for err in schema_errors:
print(f"- {err}")
return 1
if conflicts:
print("Merge conflicts detected:")
for conflict in conflicts:
path_value = conflict["path"] if conflict["path"] else "<root>"
print(
f"- {path_value}: {conflict['previousFile']} disagrees with {conflict['currentFile']}"
)
else:
print("No merge conflicts detected.")
yaml_output = to_yaml(merged, merged_comments)
print("\nMerged YAML:\n")
print(yaml_output)
out_abs = os.path.abspath(output_path)
with open(out_abs, "w", encoding="utf-8") as handle:
handle.write(yaml_output + "\n")
print(f"\nSaved merged config: {out_abs}")
return 0
if __name__ == "__main__":
raise SystemExit(main())