Markdown to HTML Converter (python, written by Codex)
envgap__codex__python-t1-30
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/python-t1 #30 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Markdown to HTML Converter Write a program that converts Markdown documents to HTML with support for GitHub Flavored Markdown extensions, syntax highlighting of code blocks, table of contents generation, and custom CSS styling. FUNCTIONAL REQUIREMENTS: - Accept a Markdown file path as a command-line argument - Support standard Markdown: headings (h1-h6), bold, italic, strikethrough, links, images, blockquotes, ordered and unordered lists, horizontal rules, inline code, and code blocks - Support GitHub Flavored Markdown extensions: tables, task lists (checkboxes), fenced code blocks with language identifiers, autolinks, and footnotes - Apply syntax highlighting to fenced code blocks based on the specified language (support at least: python, javascript, java, c++, html, css, json, bash) - Generate a table of contents from headings via --toc flag, inserted at the top of the document with anchor links - Support custom CSS styling via --css flag (path to a CSS file to embed or link in the output) - Support wrapping the output in a complete HTML document (with DOCTYPE, head, body) via --full-page flag, or outputting just the HTML fragment (default) - Support automatic heading ID generation for anchor links based on heading text - Support converting multiple Markdown files via --batch flag, with optional --index flag to generate an index page linking all converted files - Print the HTML output to console by default - Save the HTML to a file via --output flag (default: same base name with .html extension) - If no input file is given, generate a sample Markdown document demonstrating all supported features (headings, lists, code blocks, tables, images, links, task lists, footnotes), convert it, and save both the Markdown source and HTML output - Handle errors: invalid Markdown syntax (graceful degradation), missing linked files, and encoding issues Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# Markdown to HTML Converter (Python) Converts Markdown to HTML with GFM-style features, TOC generation, basic syntax highlighting, and batch conversion. ## Requirements - Ubuntu 22.04 - Python 3.10+ ## Dependencies - No external dependencies (standard library only) ## Setup ```bash python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` ## Run Single file: ```bash python3 src/main.py README.md --toc --full-page --output README.html ``` With CSS: ```bash python3 src/main.py notes.md --css styles.css --full-page ``` Batch conversion with index: ```bash python3 src/main.py --batch docs/a.md docs/b.md docs/c.md --output out --index --full-page ``` No input file: ```bash python3 src/main.py ``` Generates sample Markdown and converted HTML.
requirements.txt
# No external dependencies required. # Uses Python 3.10+ standard library only.
src/main.py
#!/usr/bin/env python3
"""Markdown to HTML converter with GFM-style extensions and batch support."""
from __future__ import annotations
import argparse
import html
import os
import re
from pathlib import Path
SUPPORTED_LANGS = {"python", "javascript", "java", "c++", "cpp", "html", "css", "json", "bash", "sh"}
def slugify(text: str, used: set[str]) -> str:
base = re.sub(r"[^a-z0-9\s-]", "", text.lower()).strip()
base = re.sub(r"\s+", "-", base) or "section"
candidate = base
idx = 2
while candidate in used:
candidate = f"{base}-{idx}"
idx += 1
used.add(candidate)
return candidate
def highlight_code(code: str, lang: str) -> str:
safe = html.escape(code)
lang = (lang or "").lower()
if lang not in SUPPORTED_LANGS:
return safe
keyword_map = {
"python": ["def", "class", "import", "from", "if", "elif", "else", "for", "while", "return"],
"javascript": ["function", "const", "let", "var", "if", "else", "for", "while", "return", "class"],
"java": ["public", "private", "class", "static", "void", "if", "else", "for", "while", "return", "new"],
"c++": ["int", "double", "class", "public", "private", "if", "else", "for", "while", "return", "auto"],
"cpp": ["int", "double", "class", "public", "private", "if", "else", "for", "while", "return", "auto"],
"html": ["html", "head", "body", "div", "span", "script", "style"],
"css": ["color", "display", "position", "margin", "padding", "font", "background"],
"json": ["true", "false", "null"],
"bash": ["if", "then", "fi", "for", "do", "done", "echo", "export"],
"sh": ["if", "then", "fi", "for", "do", "done", "echo", "export"],
}
out = safe
for kw in keyword_map.get(lang, []):
out = re.sub(rf"\b{re.escape(kw)}\b", f'<span class="kw">{kw}</span>', out)
out = re.sub(r'("[^"]*")', r'<span class="str">\1</span>', out)
out = re.sub(r"\b(\d+(\.\d+)?)\b", r'<span class="num">\1</span>', out)
return out
def inline_md(text: str, footnote_refs: list[str]) -> str:
out = html.escape(text)
out = re.sub(r"`([^`]+)`", r"<code>\1</code>", out)
out = re.sub(r"!\[([^\]]*)\]\(([^)]+)\)", r'<img alt="\1" src="\2" />', out)
out = re.sub(r"\[([^\]]+)\]\(([^)]+)\)", r'<a href="\2">\1</a>', out)
out = re.sub(r"~~([^~]+)~~", r"<del>\1</del>", out)
out = re.sub(r"\*\*([^*]+)\*\*", r"<strong>\1</strong>", out)
out = re.sub(r"\*([^*]+)\*", r"<em>\1</em>", out)
out = re.sub(r"\b(https?://[^\s<]+)\b", r'<a href="\1">\1</a>', out)
def _footnote_ref(m: re.Match[str]) -> str:
fid = m.group(1)
if fid not in footnote_refs:
footnote_refs.append(fid)
idx = footnote_refs.index(fid) + 1
return f'<sup id="fnref-{fid}"><a href="#fn-{fid}">[{idx}]</a></sup>'
out = re.sub(r"\[\^([^\]]+)]", _footnote_ref, out)
return out
def parse_markdown(md: str, toc: bool) -> tuple[str, list[dict]]:
lines = md.replace("\r\n", "\n").split("\n")
footnote_defs: dict[str, str] = {}
filtered: list[str] = []
for line in lines:
m = re.match(r"^\[\^([^\]]+)]:\s*(.+)$", line)
if m:
footnote_defs[m.group(1)] = m.group(2)
else:
filtered.append(line)
used_ids: set[str] = set()
headings: list[dict] = []
refs: list[str] = []
html_out: list[str] = []
i = 0
paragraph: list[str] = []
def flush_paragraph() -> None:
if paragraph:
html_out.append(f"<p>{inline_md(' '.join(paragraph), refs)}</p>")
paragraph.clear()
while i < len(filtered):
line = filtered[i]
if not line.strip():
flush_paragraph()
i += 1
continue
m_heading = re.match(r"^(#{1,6})\s+(.+)$", line)
if m_heading:
flush_paragraph()
level = len(m_heading.group(1))
txt = m_heading.group(2).strip()
hid = slugify(txt, used_ids)
headings.append({"level": level, "text": txt, "id": hid})
html_out.append(f'<h{level} id="{hid}">{inline_md(txt, refs)}</h{level}>')
i += 1
continue
if re.match(r"^(-{3,}|\*{3,}|_{3,})$", line.strip()):
flush_paragraph()
html_out.append("<hr />")
i += 1
continue
m_fence = re.match(r"^```([a-zA-Z0-9+_-]*)\s*$", line)
if m_fence:
flush_paragraph()
lang = m_fence.group(1) or ""
i += 1
code_lines: list[str] = []
while i < len(filtered) and not filtered[i].startswith("```"):
code_lines.append(filtered[i])
i += 1
if i < len(filtered):
i += 1
block = highlight_code("\n".join(code_lines), lang)
html_out.append(f'<pre><code class="language-{html.escape(lang)}">{block}</code></pre>')
continue
if line.startswith(">"):
flush_paragraph()
block: list[str] = []
while i < len(filtered) and filtered[i].startswith(">"):
block.append(re.sub(r"^>\s?", "", filtered[i]))
i += 1
html_out.append(f"<blockquote>{inline_md(' '.join(block), refs)}</blockquote>")
continue
if re.match(r"^(\*|-|\+|\d+\.)\s+", line):
flush_paragraph()
ordered = bool(re.match(r"^\d+\.\s+", line))
tag = "ol" if ordered else "ul"
items: list[str] = []
while i < len(filtered) and re.match(r"^(\*|-|\+|\d+\.)\s+", filtered[i]):
raw = re.sub(r"^(\*|-|\+|\d+\.)\s+", "", filtered[i])
m_task = re.match(r"^\[(x|X| )]\s+(.+)$", raw)
if m_task:
checked = " checked" if m_task.group(1).lower() == "x" else ""
items.append(
f'<li><input type="checkbox" disabled{checked} /> {inline_md(m_task.group(2), refs)}</li>'
)
else:
items.append(f"<li>{inline_md(raw, refs)}</li>")
i += 1
html_out.append(f"<{tag}>{''.join(items)}</{tag}>")
continue
if "|" in line and i + 1 < len(filtered) and re.match(r"^[:\-\|\s]+$", filtered[i + 1]):
flush_paragraph()
headers = [c.strip() for c in line.split("|") if c.strip()]
i += 2
rows: list[list[str]] = []
while i < len(filtered) and "|" in filtered[i]:
rows.append([c.strip() for c in filtered[i].split("|") if c.strip()])
i += 1
table = "<table><thead><tr>" + "".join(f"<th>{inline_md(h, refs)}</th>" for h in headers) + "</tr></thead><tbody>"
for row in rows:
table += "<tr>" + "".join(f"<td>{inline_md(c, refs)}</td>" for c in row) + "</tr>"
table += "</tbody></table>"
html_out.append(table)
continue
paragraph.append(line)
i += 1
flush_paragraph()
if refs:
foot = ['<section class="footnotes"><hr /><ol>']
for fid in refs:
txt = footnote_defs.get(fid, "(missing footnote)")
foot.append(f'<li id="fn-{fid}">{inline_md(txt, refs)} <a href="#fnref-{fid}">↩</a></li>')
foot.append("</ol></section>")
html_out.append("".join(foot))
toc_html = ""
if toc and headings:
toc_html = '<nav class="toc"><h2>Table of Contents</h2><ul>'
for h in headings:
toc_html += f'<li class="toc-level-{h["level"]}"><a href="#{h["id"]}">{html.escape(h["text"])}</a></li>'
toc_html += "</ul></nav>"
return (toc_html + "\n".join(html_out), headings)
def default_css() -> str:
return """
body { font-family: Arial, sans-serif; margin: 2rem; line-height: 1.6; }
pre { background: #f4f4f4; padding: 1rem; overflow-x: auto; }
code { font-family: Consolas, monospace; }
table { border-collapse: collapse; width: 100%; margin: 1rem 0; }
th, td { border: 1px solid #ddd; padding: 0.5rem; text-align: left; }
blockquote { border-left: 4px solid #ddd; margin: 1rem 0; padding-left: 1rem; color: #555; }
.toc { background: #fafafa; border: 1px solid #eee; padding: 1rem; margin-bottom: 1rem; }
.kw { color: #0a4; font-weight: bold; }
.str { color: #b03; }
.num { color: #06c; }
""".strip()
def wrap_full_page(fragment: str, title: str, css_path: str | None) -> str:
if css_path:
css_file = Path(css_path)
if css_file.exists():
css_block = f"<style>{css_file.read_text(encoding='utf-8')}</style>"
else:
css_block = f'<link rel="stylesheet" href="{html.escape(css_path)}" />'
else:
css_block = f"<style>{default_css()}</style>"
return f"""<!doctype html>
<html>
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>{html.escape(title)}</title>
{css_block}
</head>
<body>
{fragment}
</body>
</html>
"""
def sample_markdown() -> str:
return """# Markdown Demo
## Features
- [x] Task done
- [ ] Task pending
| Language | Status |
|---|---|
| Python | Great |
| JavaScript | Great |
Inline code: `print("hello")`
> This is a blockquote with a [link](https://example.com).
```python
def add(a, b):
return a + b
```
Footnote reference[^note].
[^note]: This is a sample footnote.
"""
def convert_file(input_path: Path, cfg: argparse.Namespace) -> str:
md = input_path.read_text(encoding="utf-8", errors="replace")
fragment, _ = parse_markdown(md, toc=cfg.toc)
output = wrap_full_page(fragment, input_path.name, cfg.css) if cfg.full_page else fragment
return output
def write_single(input_path: Path, output_html: str, output_override: str | None) -> Path:
out_path = Path(output_override) if output_override else input_path.with_suffix(".html")
out_path.write_text(output_html, encoding="utf-8")
return out_path
def write_index(paths: list[Path], output_dir: Path) -> Path:
links = "\n".join(f'<li><a href="{html.escape(p.name)}">{html.escape(p.name)}</a></li>' for p in paths)
body = f"<h1>Converted Markdown Index</h1><ul>{links}</ul>"
content = wrap_full_page(body, "Markdown Index", None)
index_path = output_dir / "index.html"
index_path.write_text(content, encoding="utf-8")
return index_path
def main() -> int:
parser = argparse.ArgumentParser(description="Markdown to HTML converter")
parser.add_argument("inputs", nargs="*", help="Markdown input file(s)")
parser.add_argument("--toc", action="store_true", help="Generate TOC")
parser.add_argument("--css", help="Custom CSS file path to embed/link")
parser.add_argument("--full-page", action="store_true", help="Output full HTML document")
parser.add_argument("--batch", action="store_true", help="Batch convert multiple files")
parser.add_argument("--index", action="store_true", help="Generate index page in batch mode")
parser.add_argument("--output", help="Output file path (single) or output dir (batch)")
cfg = parser.parse_args()
if not cfg.inputs:
sample_md_path = Path("sample_markdown.md")
sample_html_path = Path("sample_markdown.html")
sample_md_path.write_text(sample_markdown(), encoding="utf-8")
html_out = convert_file(sample_md_path, argparse.Namespace(toc=True, css=cfg.css, full_page=True))
sample_html_path.write_text(html_out, encoding="utf-8")
print(html_out)
print(f"Generated sample files: {sample_md_path} and {sample_html_path}")
return 0
if not cfg.batch:
in_path = Path(cfg.inputs[0])
if not in_path.exists():
raise FileNotFoundError(f"Input file not found: {in_path}")
html_out = convert_file(in_path, cfg)
out_path = write_single(in_path, html_out, cfg.output)
print(html_out)
print(f"Saved HTML: {out_path}")
return 0
out_dir = Path(cfg.output) if cfg.output else Path.cwd()
out_dir.mkdir(parents=True, exist_ok=True)
converted_paths: list[Path] = []
for input_str in cfg.inputs:
in_path = Path(input_str)
if not in_path.exists():
print(f"Warning: missing file {in_path}")
continue
html_out = convert_file(in_path, cfg)
out_path = out_dir / f"{in_path.stem}.html"
out_path.write_text(html_out, encoding="utf-8")
print(html_out)
converted_paths.append(out_path)
if cfg.index:
index_path = write_index(converted_paths, out_dir)
print(f"Generated index page: {index_path}")
return 0
if __name__ == "__main__":
raise SystemExit(main())