HTML to Plain Text Extractor (java, written by Codex)
envgap__codex__java-t1-34
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/java-t1 #34 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: HTML to Plain Text Extractor Write a program that converts HTML documents to clean plain text, intelligently handling formatting, tables, lists, and links while removing all markup and scripts. FUNCTIONAL REQUIREMENTS: - Accept an HTML file path as a command-line argument - Strip all HTML tags, CSS styles, JavaScript, and comments while preserving readable text content - Convert HTML formatting to plain text equivalents: headings become UPPERCASE with underlines, bold text is wrapped in *asterisks*, lists become indented with bullets (- ) or numbers (1.), horizontal rules become dashed lines - Convert HTML tables to aligned plain text tables with column padding and separator rows - Convert hyperlinks to "text [URL]" format, or optionally strip URLs via --no-urls flag - Preserve paragraph spacing: consecutive block elements get blank line separators - Handle HTML entities: decode & < > — etc. to their text equivalents - Support extracting text from only specific HTML elements via --selector flag (CSS selector syntax, e.g., --selector "article" or --selector ".content") - Support extracting and listing all URLs found in the document via --extract-urls flag - Set maximum line width via --width flag (default: 80 characters) with word wrapping - Support batch conversion of multiple HTML files via --batch flag - Print the plain text output to console by default - Save to a file via --output flag (default: same base name with .txt extension) - If no input is given, generate a sample HTML page with headings, paragraphs, links, tables, lists, images, inline styles, scripts, and HTML entities, then convert it and display both the original HTML and the extracted text - Handle errors: malformed HTML (parse gracefully), encoding detection, and binary file detection Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>org.tmlr</groupId>
<artifactId>html-to-text-extractor</artifactId>
<version>1.0.0</version>
<properties>
<maven.compiler.source>17</maven.compiler.source>
<maven.compiler.target>17</maven.compiler.target>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
</properties>
<dependencies>
<dependency>
<groupId>org.jsoup</groupId>
<artifactId>jsoup</artifactId>
<version>1.17.2</version>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-assembly-plugin</artifactId>
<version>3.7.1</version>
<configuration>
<archive>
<manifest>
<mainClass>HtmlToTextExtractor</mainClass>
</manifest>
</archive>
<descriptorRefs>
<descriptorRef>jar-with-dependencies</descriptorRef>
</descriptorRefs>
</configuration>
</plugin>
</plugins>
</build>
</project>
README.md
# HTML to Plain Text Extractor (Java) ## Requirements - Ubuntu 22.04 - JDK 17+ - Maven 3.8+ ## Build ```bash mvn -q -DskipTests package assembly:single ``` ## Run ```bash java -cp target/html-to-text-extractor-1.0.0-jar-with-dependencies.jar HtmlToTextExtractor input.html java -cp target/html-to-text-extractor-1.0.0-jar-with-dependencies.jar HtmlToTextExtractor --selector ".content" --extract-urls input.html java -cp target/html-to-text-extractor-1.0.0-jar-with-dependencies.jar HtmlToTextExtractor --batch a.html b.html --output out_dir ``` If no input is passed, sample HTML is generated and analyzed. ## Dependencies - `org.jsoup:jsoup:1.17.2` for robust HTML parsing.
src/main/java/HtmlToTextExtractor.java
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.util.*;
import java.util.regex.Pattern;
public class HtmlToTextExtractor {
private static class Config {
boolean noUrls = false;
boolean extractUrls = false;
boolean batch = false;
String selector = null;
int width = 80;
String output = null;
List<String> inputs = new ArrayList<>();
}
public static void main(String[] args) {
try {
Config cfg = parseArgs(args);
if (cfg.inputs.isEmpty()) {
runSample(cfg);
return;
}
List<String> targets = cfg.batch ? cfg.inputs : Collections.singletonList(cfg.inputs.get(0));
List<Map<String, Object>> filesReport = new ArrayList<>();
Path batchDir = null;
if (cfg.batch && cfg.output != null) {
batchDir = Paths.get(cfg.output).toAbsolutePath();
Files.createDirectories(batchDir);
}
for (String in : targets) {
Path input = Paths.get(in);
if (!Files.exists(input)) {
System.err.println("Warning: missing file " + in);
continue;
}
try {
if (isBinary(input)) throw new IllegalArgumentException("Binary file detected: " + in);
String html = Files.readString(input, StandardCharsets.UTF_8);
Result result = convert(html, cfg);
Path out;
if (cfg.batch) {
Path dir = batchDir != null ? batchDir : input.toAbsolutePath().getParent();
out = dir.resolve(stripExt(input.getFileName().toString()) + ".txt");
} else {
out = cfg.output != null ? Paths.get(cfg.output).toAbsolutePath() : replaceExt(input, ".txt");
}
Files.writeString(out, result.text + System.lineSeparator(), StandardCharsets.UTF_8);
System.out.println(result.text);
if (cfg.extractUrls) {
System.out.println("\nURLs for " + in + ":");
for (String url : result.urls) System.out.println("- " + url);
}
Map<String, Object> row = new LinkedHashMap<>();
row.put("input", input.toString());
row.put("output", out.toString());
row.put("urlsCount", result.urls.size());
filesReport.add(row);
} catch (Exception ex) {
System.err.println("Warning: " + ex.getMessage());
}
}
String json = toJson(filesReport);
Files.writeString(Paths.get("html_extract_report.json"), json, StandardCharsets.UTF_8);
} catch (Exception ex) {
System.err.println("Error: " + ex.getMessage());
System.exit(1);
}
}
private static Config parseArgs(String[] args) {
Config cfg = new Config();
for (int i = 0; i < args.length; i++) {
String a = args[i];
if (!a.startsWith("--")) {
cfg.inputs.add(a);
continue;
}
switch (a) {
case "--no-urls" -> cfg.noUrls = true;
case "--extract-urls" -> cfg.extractUrls = true;
case "--batch" -> cfg.batch = true;
case "--selector" -> cfg.selector = args[++i];
case "--width" -> cfg.width = Integer.parseInt(args[++i]);
case "--output" -> cfg.output = args[++i];
default -> throw new IllegalArgumentException("Unknown option: " + a);
}
}
if (cfg.width < 20) throw new IllegalArgumentException("--width must be >= 20");
return cfg;
}
private static boolean isBinary(Path path) throws IOException {
byte[] bytes = Files.readAllBytes(path);
int n = Math.min(bytes.length, 1024);
for (int i = 0; i < n; i++) {
if (bytes[i] == 0) return true;
}
return false;
}
private static class Result {
String text;
List<String> urls;
Result(String text, List<String> urls) { this.text = text; this.urls = urls; }
}
private static Result convert(String html, Config cfg) {
Document doc = Jsoup.parse(html);
doc.select("script,style").remove();
doc.outputSettings().prettyPrint(false);
Element root = doc.body();
if (cfg.selector != null && !cfg.selector.isBlank()) {
Elements selected = doc.select(cfg.selector);
root = new Element("div");
for (Element e : selected) root.appendChild(e.clone());
}
LinkedHashSet<String> urls = new LinkedHashSet<>();
for (Element e : root.select("[href]")) urls.add(e.attr("href"));
for (Element e : root.select("[src]")) urls.add(e.attr("src"));
for (Element table : root.select("table")) {
table.after(renderTable(table));
table.remove();
}
for (Element ol : root.select("ol")) {
ol.after(renderList(ol, true));
ol.remove();
}
for (Element ul : root.select("ul")) {
ul.after(renderList(ul, false));
ul.remove();
}
for (Element h : root.select("h1,h2,h3,h4,h5,h6")) {
String t = h.text().trim().toUpperCase(Locale.ROOT);
String line = (h.tagName().matches("h1|h2") ? "=" : "-").repeat(Math.max(3, t.length()));
h.after("\n" + t + "\n" + line + "\n");
h.remove();
}
for (Element b : root.select("b,strong")) b.text("*" + b.text() + "*");
for (Element hr : root.select("hr")) {
hr.after("\n" + "-".repeat(40) + "\n");
hr.remove();
}
for (Element a : root.select("a[href]")) {
String text = a.text();
if (!cfg.noUrls) text = text + " [" + a.attr("href") + "]";
a.text(text);
}
String rendered = root.wholeText();
rendered = rendered.replaceAll("[ \t]+", " ");
rendered = rendered.replaceAll("\n{3,}", "\n\n").trim();
rendered = wrap(rendered, cfg.width);
return new Result(rendered, new ArrayList<>(urls));
}
private static String renderTable(Element table) {
List<List<String>> rows = new ArrayList<>();
for (Element tr : table.select("tr")) {
List<String> row = new ArrayList<>();
for (Element c : tr.select("th,td")) row.add(c.text().trim());
if (!row.isEmpty()) rows.add(row);
}
if (rows.isEmpty()) return "";
int cols = rows.stream().mapToInt(List::size).max().orElse(0);
int[] widths = new int[cols];
for (List<String> r : rows) {
for (int i = 0; i < cols; i++) {
String v = i < r.size() ? r.get(i) : "";
widths[i] = Math.max(widths[i], v.length());
}
}
StringBuilder sb = new StringBuilder("\n");
for (int i = 0; i < rows.size(); i++) {
List<String> r = rows.get(i);
sb.append("| ");
for (int c = 0; c < cols; c++) {
String v = c < r.size() ? r.get(c) : "";
sb.append(String.format("%-" + widths[c] + "s", v));
sb.append(c == cols - 1 ? " |\n" : " | ");
}
if (i == 0) {
sb.append("| ");
for (int c = 0; c < cols; c++) {
sb.append("-".repeat(Math.max(3, widths[c])));
sb.append(c == cols - 1 ? " |\n" : " | ");
}
}
}
sb.append("\n");
return sb.toString();
}
private static String renderList(Element list, boolean ordered) {
StringBuilder sb = new StringBuilder("\n");
int i = 1;
for (Element li : list.select("> li")) {
sb.append(ordered ? i + ". " : "- ").append(li.text().trim()).append("\n");
i++;
}
sb.append("\n");
return sb.toString();
}
private static String wrap(String text, int width) {
StringBuilder out = new StringBuilder();
for (String line : text.split("\n")) {
if (line.isBlank()) {
out.append("\n");
continue;
}
String remain = line.trim();
while (remain.length() > width) {
int cut = remain.lastIndexOf(' ', width);
if (cut < 10) cut = width;
out.append(remain, 0, cut).append("\n");
remain = remain.substring(cut).trim();
}
out.append(remain).append("\n");
}
return out.toString().trim();
}
private static void runSample(Config cfg) throws IOException {
String sample = """
<!doctype html>
<html><head><title>Sample</title><style>.x{color:red}</style><script>console.log('x')</script></head>
<body>
<article class='content' id='main'>
<h1>Demo Article</h1>
<p>This is a <strong>sample</strong> paragraph with <a href='https://example.com'>a link</a> & entities — text.</p>
<ul><li>Alpha</li><li>Beta</li></ul>
<ol><li>One</li><li>Two</li></ol>
<hr>
<table><tr><th>Name</th><th>Role</th><th>Score</th></tr><tr><td>Ada</td><td>Engineer</td><td>98</td></tr></table>
</article>
</body></html>
""";
Path htmlPath = Paths.get("sample_page.html");
Files.writeString(htmlPath, sample, StandardCharsets.UTF_8);
Result result = convert(sample, cfg);
Files.writeString(Paths.get("sample_page.txt"), result.text + System.lineSeparator(), StandardCharsets.UTF_8);
System.out.println("===== ORIGINAL HTML =====");
System.out.println(sample);
System.out.println("\n===== EXTRACTED TEXT =====");
System.out.println(result.text);
if (cfg.extractUrls) {
System.out.println("\nURLs:");
for (String url : result.urls) System.out.println("- " + url);
}
}
private static Path replaceExt(Path path, String ext) {
String name = stripExt(path.getFileName().toString()) + ext;
Path parent = path.toAbsolutePath().getParent();
return (parent == null ? Paths.get(name) : parent.resolve(name));
}
private static String stripExt(String fileName) {
int idx = fileName.lastIndexOf('.');
return idx >= 0 ? fileName.substring(0, idx) : fileName;
}
private static String toJson(List<Map<String, Object>> rows) {
StringBuilder sb = new StringBuilder();
sb.append("{\n");
sb.append(" \"generatedAt\": \"").append(Instant.now()).append("\",\n");
sb.append(" \"files\": [\n");
for (int i = 0; i < rows.size(); i++) {
Map<String, Object> r = rows.get(i);
sb.append(" {\"input\":\"").append(escapeJson(String.valueOf(r.get("input")))).append("\",");
sb.append("\"output\":\"").append(escapeJson(String.valueOf(r.get("output")))).append("\",");
sb.append("\"urlsCount\":").append(r.get("urlsCount")).append("}");
if (i < rows.size() - 1) sb.append(',');
sb.append('\n');
}
sb.append(" ]\n");
sb.append("}\n");
return sb.toString();
}
private static String escapeJson(String s) {
return s.replace("\\", "\\\\").replace("\"", "\\\"");
}
}