← All tasks
pythoncodex/python-t1 #4Not a task: already works

YAML Config Merger (python, written by Codex)

envgap__codex__python-t1-4

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/python-t1 #4 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: YAML Config Merger

Write a program that merges multiple YAML configuration files into a single unified configuration, supporting deep merging of nested structures, array handling strategies, and conflict resolution.

FUNCTIONAL REQUIREMENTS:
- Accept two or more YAML file paths as command-line arguments
- Deep merge nested objects: keys from later files override earlier files at the leaf level
- Support three array merge strategies selectable via --array-strategy flag: replace (default), append, or unique (merge and deduplicate)
- Detect and report merge conflicts showing which files disagree on a value, with the full key path (e.g., database.connection.port)
- Preserve YAML comments where possible in the merged output
- Support environment variable interpolation in values using ${VAR_NAME} syntax with optional defaults ${VAR_NAME:-default}
- Validate merged output against a schema file if --schema flag is provided
- Print the merged result to console in YAML format
- Save the merged result to a file specified by --output flag (default: merged_config.yaml)
- If no input files are given, generate three sample YAML config files with overlapping keys, nested structures, and arrays, then merge them
- Handle malformed YAML with clear error messages identifying the file and location of the problem

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

README.md
# YAML Config Merger (Python)

Merges multiple YAML config files into one with deep merge, array strategies, conflict reporting, env interpolation, optional schema validation, and YAML output.

## Features

- Accepts 2+ YAML file paths.
- Deep merge nested objects; later files override leaf values.
- `--array-strategy`: `replace` (default), `append`, `unique`.
- Conflict reporting with full key paths and disagreeing files.
- Preserves YAML comments where possible.
- Environment interpolation:
  - `${VAR_NAME}`
  - `${VAR_NAME:-default}`
- Optional schema validation via `--schema`.
- Prints merged YAML.
- Writes output file via `--output` (default `merged_config.yaml`).
- No args: generates three sample YAML files, then merges.
- Malformed YAML errors include file and line.

## Requirements

- Ubuntu 22.04
- Python 3.10+

## Dependencies (Pinned)

No external dependencies are used.

- Direct dependencies: none
- Transitive dependencies: none

## Run

```bash
cd TMLR/code_generation/codex_generated/p_04/python
python3 src/main.py config1.yaml config2.yaml config3.yaml
```

```bash
python3 src/main.py --array-strategy unique --output merged.yaml config1.yaml config2.yaml
```

```bash
python3 src/main.py --schema schema.yaml config1.yaml config2.yaml
```

```bash
python3 src/main.py
```
requirements.txt
# No external dependencies required.
# Direct dependencies: none
# Transitive dependencies: none
src/main.py
#!/usr/bin/env python3
import json
import os
import re
import sys


class YamlParseError(Exception):
    def __init__(self, message, line):
        super().__init__(f"{message} at line {line}")
        self.line = line


def deep_equal(a, b):
    return a == b


def deep_clone(value):
    return json.loads(json.dumps(value))


def scalar_from_text(text):
    if text in ("null", "~"):
        return None
    if text == "true":
        return True
    if text == "false":
        return False
    if re.fullmatch(r"-?\d+(\.\d+)?", text):
        return float(text) if "." in text else int(text)
    if (text.startswith('"') and text.endswith('"')) or (text.startswith("'") and text.endswith("'")):
        return text[1:-1]
    return text


def yaml_scalar(value):
    if value is None:
        return "null"
    if isinstance(value, bool):
        return "true" if value else "false"
    if isinstance(value, (int, float)):
        return str(value)
    if isinstance(value, str):
        if value == "":
            return '""'
        if re.search(r'[:#\-\[\]\{\},]|\A\s|\s\Z|\n', value):
            return json.dumps(value)
        return value
    return json.dumps(value)


def parse_yaml(text):
    lines = text.replace("\r\n", "\n").replace("\r", "\n").split("\n")
    comments = {}

    def line_info(index):
        raw = lines[index] if index < len(lines) else ""
        indent = len(raw) - len(raw.lstrip(" "))
        trimmed = raw.strip()
        return raw, indent, trimmed, index + 1

    def parse_block(start, indent, path_prefix):
        index = start
        pending_comments = []

        while index < len(lines):
            _, cur_indent, trimmed, line_no = line_info(index)
            if trimmed == "":
                index += 1
                continue
            if cur_indent < indent:
                break
            if cur_indent > indent:
                raise YamlParseError("Unexpected indentation", line_no)
            if trimmed.startswith("#"):
                pending_comments.append(trimmed[1:].strip())
                index += 1
                continue
            break

        if index >= len(lines):
            return None, index

        _, first_indent, first_trim, _ = line_info(index)
        if first_indent != indent:
            return None, index
        is_seq = first_trim.startswith("- ")

        if is_seq:
            arr = []
            while index < len(lines):
                _, cur_indent, trimmed, line_no = line_info(index)
                if trimmed == "" or trimmed.startswith("#"):
                    index += 1
                    continue
                if cur_indent < indent:
                    break
                if cur_indent > indent:
                    raise YamlParseError("Unexpected indentation", line_no)
                if not trimmed.startswith("- "):
                    break

                rest = trimmed[2:].strip()
                if rest == "":
                    nested, next_index = parse_block(index + 1, indent + 2, f"{path_prefix}[{len(arr)}]")
                    arr.append(nested)
                    index = next_index
                elif re.match(r"^[^:#]+:\s*", rest):
                    split = rest.find(":")
                    key = rest[:split].strip()
                    value_text = rest[split + 1 :].strip()
                    item = {}
                    if value_text == "":
                        nested, next_index = parse_block(
                            index + 1, indent + 4, f"{path_prefix}[{len(arr)}].{key}"
                        )
                        item[key] = nested
                        arr.append(item)
                        index = next_index
                    else:
                        item[key] = scalar_from_text(value_text)
                        arr.append(item)
                        index += 1
                else:
                    arr.append(scalar_from_text(rest))
                    index += 1
            return arr, index

        obj = {}
        local_comments = list(pending_comments)
        while index < len(lines):
            _, cur_indent, trimmed, line_no = line_info(index)
            if trimmed == "":
                index += 1
                continue
            if cur_indent < indent:
                break
            if cur_indent > indent:
                raise YamlParseError("Unexpected indentation", line_no)
            if trimmed.startswith("#"):
                local_comments.append(trimmed[1:].strip())
                index += 1
                continue

            split = trimmed.find(":")
            if split < 1:
                raise YamlParseError("Invalid mapping entry", line_no)
            key = trimmed[:split].strip()
            rest = trimmed[split + 1 :].strip()
            full_path = f"{path_prefix}.{key}" if path_prefix else key
            if local_comments:
                comments[full_path] = list(local_comments)
                local_comments = []

            if rest == "":
                nested, next_index = parse_block(index + 1, indent + 2, full_path)
                obj[key] = nested
                index = next_index
            else:
                inline = re.split(r"\s+#", rest, maxsplit=1)[0].strip()
                obj[key] = scalar_from_text(inline)
                index += 1
        return obj, index

    value, _ = parse_block(0, 0, "")
    return (value if value is not None else {}), comments


def interpolate_env(value):
    if isinstance(value, str):
        pattern = re.compile(r"\$\{([A-Za-z_][A-Za-z0-9_]*)(:-([^}]*))?\}")

        def repl(match):
            key = match.group(1)
            default = match.group(3)
            if key in os.environ:
                return os.environ[key]
            return default if default is not None else ""

        return pattern.sub(repl, value)
    if isinstance(value, list):
        return [interpolate_env(v) for v in value]
    if isinstance(value, dict):
        return {k: interpolate_env(v) for k, v in value.items()}
    return value


def merge_values(base, incoming, path_text, strategy, conflicts, source_map, current_file):
    if isinstance(base, dict) and isinstance(incoming, dict):
        merged = dict(base)
        for key, val in incoming.items():
            child_path = f"{path_text}.{key}" if path_text else key
            if key in merged:
                merged[key] = merge_values(
                    merged[key], val, child_path, strategy, conflicts, source_map, current_file
                )
            else:
                merged[key] = deep_clone(val)
                source_map[child_path] = current_file
        return merged

    if isinstance(base, list) and isinstance(incoming, list):
        if strategy == "append":
            merged = [deep_clone(v) for v in base] + [deep_clone(v) for v in incoming]
        elif strategy == "unique":
            merged = [deep_clone(v) for v in base]
            for item in incoming:
                if not any(deep_equal(x, item) for x in merged):
                    merged.append(deep_clone(item))
        else:
            merged = [deep_clone(v) for v in incoming]
        if not deep_equal(base, incoming):
            conflicts.append(
                {
                    "path": path_text,
                    "previousFile": source_map.get(path_text, "unknown"),
                    "currentFile": current_file,
                    "previousValue": base,
                    "currentValue": incoming,
                }
            )
        source_map[path_text] = current_file
        return merged

    if not deep_equal(base, incoming):
        conflicts.append(
            {
                "path": path_text,
                "previousFile": source_map.get(path_text, "unknown"),
                "currentFile": current_file,
                "previousValue": base,
                "currentValue": incoming,
            }
        )
    source_map[path_text] = current_file
    return deep_clone(incoming)


def validate_schema(data, schema, path_text, errors):
    if not isinstance(schema, dict):
        return
    declared_type = schema.get("type")
    actual_type = (
        "null"
        if data is None
        else "array"
        if isinstance(data, list)
        else "object"
        if isinstance(data, dict)
        else "boolean"
        if isinstance(data, bool)
        else "number"
        if isinstance(data, (int, float))
        else "string"
        if isinstance(data, str)
        else "unknown"
    )
    if isinstance(declared_type, str) and declared_type != actual_type:
        errors.append(f"{path_text or '$'}: expected type {declared_type}, got {actual_type}")
        return

    enum_values = schema.get("enum")
    if isinstance(enum_values, list) and not any(deep_equal(v, data) for v in enum_values):
        errors.append(f"{path_text or '$'}: value not in enum")

    if actual_type == "object":
        req = schema.get("required")
        if isinstance(req, list):
            for key in req:
                if key not in data:
                    errors.append(f"{path_text or '$'}: missing required key {key}")
        props = schema.get("properties")
        if isinstance(props, dict):
            for key, child_schema in props.items():
                if key in data:
                    child_path = f"{path_text}.{key}" if path_text else key
                    validate_schema(data[key], child_schema, child_path, errors)
    elif actual_type == "array":
        item_schema = schema.get("items")
        if item_schema is not None:
            for idx, item in enumerate(data):
                validate_schema(item, item_schema, f"{path_text}[{idx}]", errors)


def to_yaml(value, comments_map, path_text="", indent=0):
    pad = " " * indent
    lines = []

    if isinstance(value, list):
        for idx, item in enumerate(value):
            if isinstance(item, dict) or isinstance(item, list):
                lines.append(f"{pad}-")
                lines.append(to_yaml(item, comments_map, f"{path_text}[{idx}]", indent + 2))
            else:
                lines.append(f"{pad}- {yaml_scalar(item)}")
        return "\n".join(line for line in lines if line != "")

    if isinstance(value, dict):
        for key, val in value.items():
            child_path = f"{path_text}.{key}" if path_text else key
            if child_path in comments_map:
                for comment in comments_map[child_path]:
                    lines.append(f"{pad}# {comment}")
            if isinstance(val, (dict, list)):
                lines.append(f"{pad}{key}:")
                lines.append(to_yaml(val, comments_map, child_path, indent + 2))
            else:
                lines.append(f"{pad}{key}: {yaml_scalar(val)}")
        return "\n".join(line for line in lines if line != "")

    return f"{pad}{yaml_scalar(value)}"


def load_yaml_file(file_path):
    try:
        with open(file_path, "r", encoding="utf-8") as handle:
            text = handle.read()
    except OSError as exc:
        raise RuntimeError(f"Failed to read {file_path}: {exc}") from exc
    try:
        data, comments = parse_yaml(text)
        return {"data": data, "comments": comments}
    except YamlParseError as exc:
        raise RuntimeError(f"Malformed YAML in {file_path}: {exc}") from exc


def sample_files():
    return {
        "config_1.yaml": """# Base app config
app:
  name: merger-demo
  env: ${APP_ENV:-dev}
database:
  host: localhost
  port: 5432
  tags:
    - core
    - primary
features:
  enabled:
    - auth
    - api
""",
        "config_2.yaml": """# Override database and features
database:
  port: 5433
  user: ${DB_USER:-admin}
  tags:
    - analytics
features:
  enabled:
    - api
    - billing
""",
        "config_3.yaml": """# Production tuning
app:
  env: prod
database:
  host: db.internal
  retries: 5
features:
  enabled:
    - auth
    - billing
""",
    }


def main():
    args = sys.argv[1:]
    files = []
    array_strategy = "replace"
    schema_path = None
    output_path = "merged_config.yaml"

    i = 0
    while i < len(args):
        arg = args[i]
        if arg == "--array-strategy":
            i += 1
            array_strategy = args[i] if i < len(args) else ""
        elif arg == "--schema":
            i += 1
            schema_path = args[i] if i < len(args) else None
        elif arg == "--output":
            i += 1
            output_path = args[i] if i < len(args) else output_path
        else:
            files.append(arg)
        i += 1

    if array_strategy not in ("replace", "append", "unique"):
        print("Invalid --array-strategy. Use replace, append, or unique.")
        return 1

    if not files:
        generated = sample_files()
        files = []
        for name, content in generated.items():
            abs_path = os.path.abspath(name)
            with open(abs_path, "w", encoding="utf-8") as handle:
                handle.write(content)
            files.append(abs_path)
        print(f"No input files provided. Generated sample files: {', '.join(files)}")

    if len(files) < 2:
        print("Provide at least two YAML files.")
        return 1

    parsed_files = []
    for file_name in files:
        abs_file = os.path.abspath(file_name)
        try:
            parsed = load_yaml_file(abs_file)
        except RuntimeError as exc:
            print(exc)
            return 1
        parsed_files.append({"file": abs_file, **parsed})

    merged = deep_clone(parsed_files[0]["data"])
    source_map = {}

    def collect_leaf_paths(value, path_prefix):
        if isinstance(value, dict):
            for key, child in value.items():
                child_path = f"{path_prefix}.{key}" if path_prefix else key
                collect_leaf_paths(child, child_path)
        else:
            source_map[path_prefix] = parsed_files[0]["file"]

    collect_leaf_paths(merged, "")

    merged_comments = dict(parsed_files[0]["comments"])
    conflicts = []

    for parsed in parsed_files[1:]:
        merged = merge_values(
            merged,
            parsed["data"],
            "",
            array_strategy,
            conflicts,
            source_map,
            parsed["file"],
        )
        for key, comments in parsed["comments"].items():
            if key not in merged_comments:
                merged_comments[key] = comments

    merged = interpolate_env(merged)

    if schema_path:
        try:
            schema = load_yaml_file(os.path.abspath(schema_path))["data"]
        except RuntimeError as exc:
            print(exc)
            return 1
        schema_errors = []
        validate_schema(merged, schema, "", schema_errors)
        if schema_errors:
            print("Schema validation failed:")
            for err in schema_errors:
                print(f"- {err}")
            return 1

    if conflicts:
        print("Merge conflicts detected:")
        for conflict in conflicts:
            path_value = conflict["path"] if conflict["path"] else "<root>"
            print(
                f"- {path_value}: {conflict['previousFile']} disagrees with {conflict['currentFile']}"
            )
    else:
        print("No merge conflicts detected.")

    yaml_output = to_yaml(merged, merged_comments)
    print("\nMerged YAML:\n")
    print(yaml_output)

    out_abs = os.path.abspath(output_path)
    with open(out_abs, "w", encoding="utf-8") as handle:
        handle.write(yaml_output + "\n")
    print(f"\nSaved merged config: {out_abs}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())