← All tasks
pythoncodex/python-t1 #30Not a task: already works

Markdown to HTML Converter (python, written by Codex)

envgap__codex__python-t1-30

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
requirements.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/python-t1 #30 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: Markdown to HTML Converter

Write a program that converts Markdown documents to HTML with support for GitHub Flavored Markdown extensions, syntax highlighting of code blocks, table of contents generation, and custom CSS styling.

FUNCTIONAL REQUIREMENTS:
- Accept a Markdown file path as a command-line argument
- Support standard Markdown: headings (h1-h6), bold, italic, strikethrough, links, images, blockquotes, ordered and unordered lists, horizontal rules, inline code, and code blocks
- Support GitHub Flavored Markdown extensions: tables, task lists (checkboxes), fenced code blocks with language identifiers, autolinks, and footnotes
- Apply syntax highlighting to fenced code blocks based on the specified language (support at least: python, javascript, java, c++, html, css, json, bash)
- Generate a table of contents from headings via --toc flag, inserted at the top of the document with anchor links
- Support custom CSS styling via --css flag (path to a CSS file to embed or link in the output)
- Support wrapping the output in a complete HTML document (with DOCTYPE, head, body) via --full-page flag, or outputting just the HTML fragment (default)
- Support automatic heading ID generation for anchor links based on heading text
- Support converting multiple Markdown files via --batch flag, with optional --index flag to generate an index page linking all converted files
- Print the HTML output to console by default
- Save the HTML to a file via --output flag (default: same base name with .html extension)
- If no input file is given, generate a sample Markdown document demonstrating all supported features (headings, lists, code blocks, tables, images, links, task lists, footnotes), convert it, and save both the Markdown source and HTML output
- Handle errors: invalid Markdown syntax (graceful degradation), missing linked files, and encoding issues

Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include:
- Source code
- requirements.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

README.md
# Markdown to HTML Converter (Python)

Converts Markdown to HTML with GFM-style features, TOC generation, basic syntax highlighting, and batch conversion.

## Requirements
- Ubuntu 22.04
- Python 3.10+

## Dependencies
- No external dependencies (standard library only)

## Setup
```bash
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

## Run
Single file:
```bash
python3 src/main.py README.md --toc --full-page --output README.html
```

With CSS:
```bash
python3 src/main.py notes.md --css styles.css --full-page
```

Batch conversion with index:
```bash
python3 src/main.py --batch docs/a.md docs/b.md docs/c.md --output out --index --full-page
```

No input file:
```bash
python3 src/main.py
```
Generates sample Markdown and converted HTML.
requirements.txt
# No external dependencies required.
# Uses Python 3.10+ standard library only.
src/main.py
#!/usr/bin/env python3
"""Markdown to HTML converter with GFM-style extensions and batch support."""

from __future__ import annotations

import argparse
import html
import os
import re
from pathlib import Path


SUPPORTED_LANGS = {"python", "javascript", "java", "c++", "cpp", "html", "css", "json", "bash", "sh"}


def slugify(text: str, used: set[str]) -> str:
    base = re.sub(r"[^a-z0-9\s-]", "", text.lower()).strip()
    base = re.sub(r"\s+", "-", base) or "section"
    candidate = base
    idx = 2
    while candidate in used:
        candidate = f"{base}-{idx}"
        idx += 1
    used.add(candidate)
    return candidate


def highlight_code(code: str, lang: str) -> str:
    safe = html.escape(code)
    lang = (lang or "").lower()
    if lang not in SUPPORTED_LANGS:
        return safe
    keyword_map = {
        "python": ["def", "class", "import", "from", "if", "elif", "else", "for", "while", "return"],
        "javascript": ["function", "const", "let", "var", "if", "else", "for", "while", "return", "class"],
        "java": ["public", "private", "class", "static", "void", "if", "else", "for", "while", "return", "new"],
        "c++": ["int", "double", "class", "public", "private", "if", "else", "for", "while", "return", "auto"],
        "cpp": ["int", "double", "class", "public", "private", "if", "else", "for", "while", "return", "auto"],
        "html": ["html", "head", "body", "div", "span", "script", "style"],
        "css": ["color", "display", "position", "margin", "padding", "font", "background"],
        "json": ["true", "false", "null"],
        "bash": ["if", "then", "fi", "for", "do", "done", "echo", "export"],
        "sh": ["if", "then", "fi", "for", "do", "done", "echo", "export"],
    }
    out = safe
    for kw in keyword_map.get(lang, []):
        out = re.sub(rf"\b{re.escape(kw)}\b", f'<span class="kw">{kw}</span>', out)
    out = re.sub(r'("[^"]*")', r'<span class="str">\1</span>', out)
    out = re.sub(r"\b(\d+(\.\d+)?)\b", r'<span class="num">\1</span>', out)
    return out


def inline_md(text: str, footnote_refs: list[str]) -> str:
    out = html.escape(text)
    out = re.sub(r"`([^`]+)`", r"<code>\1</code>", out)
    out = re.sub(r"!\[([^\]]*)\]\(([^)]+)\)", r'<img alt="\1" src="\2" />', out)
    out = re.sub(r"\[([^\]]+)\]\(([^)]+)\)", r'<a href="\2">\1</a>', out)
    out = re.sub(r"~~([^~]+)~~", r"<del>\1</del>", out)
    out = re.sub(r"\*\*([^*]+)\*\*", r"<strong>\1</strong>", out)
    out = re.sub(r"\*([^*]+)\*", r"<em>\1</em>", out)
    out = re.sub(r"\b(https?://[^\s<]+)\b", r'<a href="\1">\1</a>', out)

    def _footnote_ref(m: re.Match[str]) -> str:
        fid = m.group(1)
        if fid not in footnote_refs:
            footnote_refs.append(fid)
        idx = footnote_refs.index(fid) + 1
        return f'<sup id="fnref-{fid}"><a href="#fn-{fid}">[{idx}]</a></sup>'

    out = re.sub(r"\[\^([^\]]+)]", _footnote_ref, out)
    return out


def parse_markdown(md: str, toc: bool) -> tuple[str, list[dict]]:
    lines = md.replace("\r\n", "\n").split("\n")
    footnote_defs: dict[str, str] = {}
    filtered: list[str] = []
    for line in lines:
        m = re.match(r"^\[\^([^\]]+)]:\s*(.+)$", line)
        if m:
            footnote_defs[m.group(1)] = m.group(2)
        else:
            filtered.append(line)

    used_ids: set[str] = set()
    headings: list[dict] = []
    refs: list[str] = []
    html_out: list[str] = []
    i = 0
    paragraph: list[str] = []

    def flush_paragraph() -> None:
        if paragraph:
            html_out.append(f"<p>{inline_md(' '.join(paragraph), refs)}</p>")
            paragraph.clear()

    while i < len(filtered):
        line = filtered[i]
        if not line.strip():
            flush_paragraph()
            i += 1
            continue

        m_heading = re.match(r"^(#{1,6})\s+(.+)$", line)
        if m_heading:
            flush_paragraph()
            level = len(m_heading.group(1))
            txt = m_heading.group(2).strip()
            hid = slugify(txt, used_ids)
            headings.append({"level": level, "text": txt, "id": hid})
            html_out.append(f'<h{level} id="{hid}">{inline_md(txt, refs)}</h{level}>')
            i += 1
            continue

        if re.match(r"^(-{3,}|\*{3,}|_{3,})$", line.strip()):
            flush_paragraph()
            html_out.append("<hr />")
            i += 1
            continue

        m_fence = re.match(r"^```([a-zA-Z0-9+_-]*)\s*$", line)
        if m_fence:
            flush_paragraph()
            lang = m_fence.group(1) or ""
            i += 1
            code_lines: list[str] = []
            while i < len(filtered) and not filtered[i].startswith("```"):
                code_lines.append(filtered[i])
                i += 1
            if i < len(filtered):
                i += 1
            block = highlight_code("\n".join(code_lines), lang)
            html_out.append(f'<pre><code class="language-{html.escape(lang)}">{block}</code></pre>')
            continue

        if line.startswith(">"):
            flush_paragraph()
            block: list[str] = []
            while i < len(filtered) and filtered[i].startswith(">"):
                block.append(re.sub(r"^>\s?", "", filtered[i]))
                i += 1
            html_out.append(f"<blockquote>{inline_md(' '.join(block), refs)}</blockquote>")
            continue

        if re.match(r"^(\*|-|\+|\d+\.)\s+", line):
            flush_paragraph()
            ordered = bool(re.match(r"^\d+\.\s+", line))
            tag = "ol" if ordered else "ul"
            items: list[str] = []
            while i < len(filtered) and re.match(r"^(\*|-|\+|\d+\.)\s+", filtered[i]):
                raw = re.sub(r"^(\*|-|\+|\d+\.)\s+", "", filtered[i])
                m_task = re.match(r"^\[(x|X| )]\s+(.+)$", raw)
                if m_task:
                    checked = " checked" if m_task.group(1).lower() == "x" else ""
                    items.append(
                        f'<li><input type="checkbox" disabled{checked} /> {inline_md(m_task.group(2), refs)}</li>'
                    )
                else:
                    items.append(f"<li>{inline_md(raw, refs)}</li>")
                i += 1
            html_out.append(f"<{tag}>{''.join(items)}</{tag}>")
            continue

        if "|" in line and i + 1 < len(filtered) and re.match(r"^[:\-\|\s]+$", filtered[i + 1]):
            flush_paragraph()
            headers = [c.strip() for c in line.split("|") if c.strip()]
            i += 2
            rows: list[list[str]] = []
            while i < len(filtered) and "|" in filtered[i]:
                rows.append([c.strip() for c in filtered[i].split("|") if c.strip()])
                i += 1
            table = "<table><thead><tr>" + "".join(f"<th>{inline_md(h, refs)}</th>" for h in headers) + "</tr></thead><tbody>"
            for row in rows:
                table += "<tr>" + "".join(f"<td>{inline_md(c, refs)}</td>" for c in row) + "</tr>"
            table += "</tbody></table>"
            html_out.append(table)
            continue

        paragraph.append(line)
        i += 1

    flush_paragraph()

    if refs:
        foot = ['<section class="footnotes"><hr /><ol>']
        for fid in refs:
            txt = footnote_defs.get(fid, "(missing footnote)")
            foot.append(f'<li id="fn-{fid}">{inline_md(txt, refs)} <a href="#fnref-{fid}">&#8617;</a></li>')
        foot.append("</ol></section>")
        html_out.append("".join(foot))

    toc_html = ""
    if toc and headings:
        toc_html = '<nav class="toc"><h2>Table of Contents</h2><ul>'
        for h in headings:
            toc_html += f'<li class="toc-level-{h["level"]}"><a href="#{h["id"]}">{html.escape(h["text"])}</a></li>'
        toc_html += "</ul></nav>"

    return (toc_html + "\n".join(html_out), headings)


def default_css() -> str:
    return """
body { font-family: Arial, sans-serif; margin: 2rem; line-height: 1.6; }
pre { background: #f4f4f4; padding: 1rem; overflow-x: auto; }
code { font-family: Consolas, monospace; }
table { border-collapse: collapse; width: 100%; margin: 1rem 0; }
th, td { border: 1px solid #ddd; padding: 0.5rem; text-align: left; }
blockquote { border-left: 4px solid #ddd; margin: 1rem 0; padding-left: 1rem; color: #555; }
.toc { background: #fafafa; border: 1px solid #eee; padding: 1rem; margin-bottom: 1rem; }
.kw { color: #0a4; font-weight: bold; }
.str { color: #b03; }
.num { color: #06c; }
""".strip()


def wrap_full_page(fragment: str, title: str, css_path: str | None) -> str:
    if css_path:
        css_file = Path(css_path)
        if css_file.exists():
            css_block = f"<style>{css_file.read_text(encoding='utf-8')}</style>"
        else:
            css_block = f'<link rel="stylesheet" href="{html.escape(css_path)}" />'
    else:
        css_block = f"<style>{default_css()}</style>"
    return f"""<!doctype html>
<html>
  <head>
    <meta charset="utf-8" />
    <meta name="viewport" content="width=device-width, initial-scale=1" />
    <title>{html.escape(title)}</title>
    {css_block}
  </head>
  <body>
{fragment}
  </body>
</html>
"""


def sample_markdown() -> str:
    return """# Markdown Demo

## Features
- [x] Task done
- [ ] Task pending

| Language | Status |
|---|---|
| Python | Great |
| JavaScript | Great |

Inline code: `print("hello")`

> This is a blockquote with a [link](https://example.com).

```python
def add(a, b):
    return a + b
```

Footnote reference[^note].

[^note]: This is a sample footnote.
"""


def convert_file(input_path: Path, cfg: argparse.Namespace) -> str:
    md = input_path.read_text(encoding="utf-8", errors="replace")
    fragment, _ = parse_markdown(md, toc=cfg.toc)
    output = wrap_full_page(fragment, input_path.name, cfg.css) if cfg.full_page else fragment
    return output


def write_single(input_path: Path, output_html: str, output_override: str | None) -> Path:
    out_path = Path(output_override) if output_override else input_path.with_suffix(".html")
    out_path.write_text(output_html, encoding="utf-8")
    return out_path


def write_index(paths: list[Path], output_dir: Path) -> Path:
    links = "\n".join(f'<li><a href="{html.escape(p.name)}">{html.escape(p.name)}</a></li>' for p in paths)
    body = f"<h1>Converted Markdown Index</h1><ul>{links}</ul>"
    content = wrap_full_page(body, "Markdown Index", None)
    index_path = output_dir / "index.html"
    index_path.write_text(content, encoding="utf-8")
    return index_path


def main() -> int:
    parser = argparse.ArgumentParser(description="Markdown to HTML converter")
    parser.add_argument("inputs", nargs="*", help="Markdown input file(s)")
    parser.add_argument("--toc", action="store_true", help="Generate TOC")
    parser.add_argument("--css", help="Custom CSS file path to embed/link")
    parser.add_argument("--full-page", action="store_true", help="Output full HTML document")
    parser.add_argument("--batch", action="store_true", help="Batch convert multiple files")
    parser.add_argument("--index", action="store_true", help="Generate index page in batch mode")
    parser.add_argument("--output", help="Output file path (single) or output dir (batch)")
    cfg = parser.parse_args()

    if not cfg.inputs:
        sample_md_path = Path("sample_markdown.md")
        sample_html_path = Path("sample_markdown.html")
        sample_md_path.write_text(sample_markdown(), encoding="utf-8")
        html_out = convert_file(sample_md_path, argparse.Namespace(toc=True, css=cfg.css, full_page=True))
        sample_html_path.write_text(html_out, encoding="utf-8")
        print(html_out)
        print(f"Generated sample files: {sample_md_path} and {sample_html_path}")
        return 0

    if not cfg.batch:
        in_path = Path(cfg.inputs[0])
        if not in_path.exists():
            raise FileNotFoundError(f"Input file not found: {in_path}")
        html_out = convert_file(in_path, cfg)
        out_path = write_single(in_path, html_out, cfg.output)
        print(html_out)
        print(f"Saved HTML: {out_path}")
        return 0

    out_dir = Path(cfg.output) if cfg.output else Path.cwd()
    out_dir.mkdir(parents=True, exist_ok=True)
    converted_paths: list[Path] = []
    for input_str in cfg.inputs:
        in_path = Path(input_str)
        if not in_path.exists():
            print(f"Warning: missing file {in_path}")
            continue
        html_out = convert_file(in_path, cfg)
        out_path = out_dir / f"{in_path.stem}.html"
        out_path.write_text(html_out, encoding="utf-8")
        print(html_out)
        converted_paths.append(out_path)
    if cfg.index:
        index_path = write_index(converted_paths, out_dir)
        print(f"Generated index page: {index_path}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())