← All tasks
javacodex/java-t1 #34Not a task: already works

HTML to Plain Text Extractor (java, written by Codex)

envgap__codex__java-t1-34

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

None
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
pom.xml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/java-t1 #34 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: HTML to Plain Text Extractor

Write a program that converts HTML documents to clean plain text, intelligently handling formatting, tables, lists, and links while removing all markup and scripts.

FUNCTIONAL REQUIREMENTS:
- Accept an HTML file path as a command-line argument
- Strip all HTML tags, CSS styles, JavaScript, and comments while preserving readable text content
- Convert HTML formatting to plain text equivalents: headings become UPPERCASE with underlines, bold text is wrapped in *asterisks*, lists become indented with bullets (- ) or numbers (1.), horizontal rules become dashed lines
- Convert HTML tables to aligned plain text tables with column padding and separator rows
- Convert hyperlinks to "text [URL]" format, or optionally strip URLs via --no-urls flag
- Preserve paragraph spacing: consecutive block elements get blank line separators
- Handle HTML entities: decode & < >   — etc. to their text equivalents
- Support extracting text from only specific HTML elements via --selector flag (CSS selector syntax, e.g., --selector "article" or --selector ".content")
- Support extracting and listing all URLs found in the document via --extract-urls flag
- Set maximum line width via --width flag (default: 80 characters) with word wrapping
- Support batch conversion of multiple HTML files via --batch flag
- Print the plain text output to console by default
- Save to a file via --output flag (default: same base name with .txt extension)
- If no input is given, generate a sample HTML page with headings, paragraphs, links, tables, lists, images, inline styles, scripts, and HTML entities, then convert it and display both the original HTML and the extracted text
- Handle errors: malformed HTML (parse gracefully), encoding detection, and binary file detection

Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include:
- Source code
- pom.xml with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
  <modelVersion>4.0.0</modelVersion>
  <groupId>org.tmlr</groupId>
  <artifactId>html-to-text-extractor</artifactId>
  <version>1.0.0</version>
  <properties>
    <maven.compiler.source>17</maven.compiler.source>
    <maven.compiler.target>17</maven.compiler.target>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
  </properties>
  <dependencies>
    <dependency>
      <groupId>org.jsoup</groupId>
      <artifactId>jsoup</artifactId>
      <version>1.17.2</version>
    </dependency>
  </dependencies>
  <build>
    <plugins>
      <plugin>
        <groupId>org.apache.maven.plugins</groupId>
        <artifactId>maven-assembly-plugin</artifactId>
        <version>3.7.1</version>
        <configuration>
          <archive>
            <manifest>
              <mainClass>HtmlToTextExtractor</mainClass>
            </manifest>
          </archive>
          <descriptorRefs>
            <descriptorRef>jar-with-dependencies</descriptorRef>
          </descriptorRefs>
        </configuration>
      </plugin>
    </plugins>
  </build>
</project>
README.md
# HTML to Plain Text Extractor (Java)

## Requirements
- Ubuntu 22.04
- JDK 17+
- Maven 3.8+

## Build
```bash
mvn -q -DskipTests package assembly:single
```

## Run
```bash
java -cp target/html-to-text-extractor-1.0.0-jar-with-dependencies.jar HtmlToTextExtractor input.html
java -cp target/html-to-text-extractor-1.0.0-jar-with-dependencies.jar HtmlToTextExtractor --selector ".content" --extract-urls input.html
java -cp target/html-to-text-extractor-1.0.0-jar-with-dependencies.jar HtmlToTextExtractor --batch a.html b.html --output out_dir
```

If no input is passed, sample HTML is generated and analyzed.

## Dependencies
- `org.jsoup:jsoup:1.17.2` for robust HTML parsing.
src/main/java/HtmlToTextExtractor.java
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.Instant;
import java.util.*;
import java.util.regex.Pattern;

public class HtmlToTextExtractor {
    private static class Config {
        boolean noUrls = false;
        boolean extractUrls = false;
        boolean batch = false;
        String selector = null;
        int width = 80;
        String output = null;
        List<String> inputs = new ArrayList<>();
    }

    public static void main(String[] args) {
        try {
            Config cfg = parseArgs(args);
            if (cfg.inputs.isEmpty()) {
                runSample(cfg);
                return;
            }
            List<String> targets = cfg.batch ? cfg.inputs : Collections.singletonList(cfg.inputs.get(0));
            List<Map<String, Object>> filesReport = new ArrayList<>();

            Path batchDir = null;
            if (cfg.batch && cfg.output != null) {
                batchDir = Paths.get(cfg.output).toAbsolutePath();
                Files.createDirectories(batchDir);
            }

            for (String in : targets) {
                Path input = Paths.get(in);
                if (!Files.exists(input)) {
                    System.err.println("Warning: missing file " + in);
                    continue;
                }
                try {
                    if (isBinary(input)) throw new IllegalArgumentException("Binary file detected: " + in);
                    String html = Files.readString(input, StandardCharsets.UTF_8);
                    Result result = convert(html, cfg);

                    Path out;
                    if (cfg.batch) {
                        Path dir = batchDir != null ? batchDir : input.toAbsolutePath().getParent();
                        out = dir.resolve(stripExt(input.getFileName().toString()) + ".txt");
                    } else {
                        out = cfg.output != null ? Paths.get(cfg.output).toAbsolutePath() : replaceExt(input, ".txt");
                    }

                    Files.writeString(out, result.text + System.lineSeparator(), StandardCharsets.UTF_8);
                    System.out.println(result.text);
                    if (cfg.extractUrls) {
                        System.out.println("\nURLs for " + in + ":");
                        for (String url : result.urls) System.out.println("- " + url);
                    }

                    Map<String, Object> row = new LinkedHashMap<>();
                    row.put("input", input.toString());
                    row.put("output", out.toString());
                    row.put("urlsCount", result.urls.size());
                    filesReport.add(row);
                } catch (Exception ex) {
                    System.err.println("Warning: " + ex.getMessage());
                }
            }

            String json = toJson(filesReport);
            Files.writeString(Paths.get("html_extract_report.json"), json, StandardCharsets.UTF_8);
        } catch (Exception ex) {
            System.err.println("Error: " + ex.getMessage());
            System.exit(1);
        }
    }

    private static Config parseArgs(String[] args) {
        Config cfg = new Config();
        for (int i = 0; i < args.length; i++) {
            String a = args[i];
            if (!a.startsWith("--")) {
                cfg.inputs.add(a);
                continue;
            }
            switch (a) {
                case "--no-urls" -> cfg.noUrls = true;
                case "--extract-urls" -> cfg.extractUrls = true;
                case "--batch" -> cfg.batch = true;
                case "--selector" -> cfg.selector = args[++i];
                case "--width" -> cfg.width = Integer.parseInt(args[++i]);
                case "--output" -> cfg.output = args[++i];
                default -> throw new IllegalArgumentException("Unknown option: " + a);
            }
        }
        if (cfg.width < 20) throw new IllegalArgumentException("--width must be >= 20");
        return cfg;
    }

    private static boolean isBinary(Path path) throws IOException {
        byte[] bytes = Files.readAllBytes(path);
        int n = Math.min(bytes.length, 1024);
        for (int i = 0; i < n; i++) {
            if (bytes[i] == 0) return true;
        }
        return false;
    }

    private static class Result {
        String text;
        List<String> urls;
        Result(String text, List<String> urls) { this.text = text; this.urls = urls; }
    }

    private static Result convert(String html, Config cfg) {
        Document doc = Jsoup.parse(html);
        doc.select("script,style").remove();
        doc.outputSettings().prettyPrint(false);

        Element root = doc.body();
        if (cfg.selector != null && !cfg.selector.isBlank()) {
            Elements selected = doc.select(cfg.selector);
            root = new Element("div");
            for (Element e : selected) root.appendChild(e.clone());
        }

        LinkedHashSet<String> urls = new LinkedHashSet<>();
        for (Element e : root.select("[href]")) urls.add(e.attr("href"));
        for (Element e : root.select("[src]")) urls.add(e.attr("src"));

        for (Element table : root.select("table")) {
            table.after(renderTable(table));
            table.remove();
        }
        for (Element ol : root.select("ol")) {
            ol.after(renderList(ol, true));
            ol.remove();
        }
        for (Element ul : root.select("ul")) {
            ul.after(renderList(ul, false));
            ul.remove();
        }
        for (Element h : root.select("h1,h2,h3,h4,h5,h6")) {
            String t = h.text().trim().toUpperCase(Locale.ROOT);
            String line = (h.tagName().matches("h1|h2") ? "=" : "-").repeat(Math.max(3, t.length()));
            h.after("\n" + t + "\n" + line + "\n");
            h.remove();
        }
        for (Element b : root.select("b,strong")) b.text("*" + b.text() + "*");
        for (Element hr : root.select("hr")) {
            hr.after("\n" + "-".repeat(40) + "\n");
            hr.remove();
        }
        for (Element a : root.select("a[href]")) {
            String text = a.text();
            if (!cfg.noUrls) text = text + " [" + a.attr("href") + "]";
            a.text(text);
        }

        String rendered = root.wholeText();
        rendered = rendered.replaceAll("[ \t]+", " ");
        rendered = rendered.replaceAll("\n{3,}", "\n\n").trim();
        rendered = wrap(rendered, cfg.width);
        return new Result(rendered, new ArrayList<>(urls));
    }

    private static String renderTable(Element table) {
        List<List<String>> rows = new ArrayList<>();
        for (Element tr : table.select("tr")) {
            List<String> row = new ArrayList<>();
            for (Element c : tr.select("th,td")) row.add(c.text().trim());
            if (!row.isEmpty()) rows.add(row);
        }
        if (rows.isEmpty()) return "";
        int cols = rows.stream().mapToInt(List::size).max().orElse(0);
        int[] widths = new int[cols];
        for (List<String> r : rows) {
            for (int i = 0; i < cols; i++) {
                String v = i < r.size() ? r.get(i) : "";
                widths[i] = Math.max(widths[i], v.length());
            }
        }
        StringBuilder sb = new StringBuilder("\n");
        for (int i = 0; i < rows.size(); i++) {
            List<String> r = rows.get(i);
            sb.append("| ");
            for (int c = 0; c < cols; c++) {
                String v = c < r.size() ? r.get(c) : "";
                sb.append(String.format("%-" + widths[c] + "s", v));
                sb.append(c == cols - 1 ? " |\n" : " | ");
            }
            if (i == 0) {
                sb.append("| ");
                for (int c = 0; c < cols; c++) {
                    sb.append("-".repeat(Math.max(3, widths[c])));
                    sb.append(c == cols - 1 ? " |\n" : " | ");
                }
            }
        }
        sb.append("\n");
        return sb.toString();
    }

    private static String renderList(Element list, boolean ordered) {
        StringBuilder sb = new StringBuilder("\n");
        int i = 1;
        for (Element li : list.select("> li")) {
            sb.append(ordered ? i + ". " : "- ").append(li.text().trim()).append("\n");
            i++;
        }
        sb.append("\n");
        return sb.toString();
    }

    private static String wrap(String text, int width) {
        StringBuilder out = new StringBuilder();
        for (String line : text.split("\n")) {
            if (line.isBlank()) {
                out.append("\n");
                continue;
            }
            String remain = line.trim();
            while (remain.length() > width) {
                int cut = remain.lastIndexOf(' ', width);
                if (cut < 10) cut = width;
                out.append(remain, 0, cut).append("\n");
                remain = remain.substring(cut).trim();
            }
            out.append(remain).append("\n");
        }
        return out.toString().trim();
    }

    private static void runSample(Config cfg) throws IOException {
        String sample = """
                <!doctype html>
                <html><head><title>Sample</title><style>.x{color:red}</style><script>console.log('x')</script></head>
                <body>
                <article class='content' id='main'>
                  <h1>Demo Article</h1>
                  <p>This is a <strong>sample</strong> paragraph with <a href='https://example.com'>a link</a> &amp; entities &mdash; text.</p>
                  <ul><li>Alpha</li><li>Beta</li></ul>
                  <ol><li>One</li><li>Two</li></ol>
                  <hr>
                  <table><tr><th>Name</th><th>Role</th><th>Score</th></tr><tr><td>Ada</td><td>Engineer</td><td>98</td></tr></table>
                </article>
                </body></html>
                """;
        Path htmlPath = Paths.get("sample_page.html");
        Files.writeString(htmlPath, sample, StandardCharsets.UTF_8);
        Result result = convert(sample, cfg);
        Files.writeString(Paths.get("sample_page.txt"), result.text + System.lineSeparator(), StandardCharsets.UTF_8);
        System.out.println("===== ORIGINAL HTML =====");
        System.out.println(sample);
        System.out.println("\n===== EXTRACTED TEXT =====");
        System.out.println(result.text);
        if (cfg.extractUrls) {
            System.out.println("\nURLs:");
            for (String url : result.urls) System.out.println("- " + url);
        }
    }

    private static Path replaceExt(Path path, String ext) {
        String name = stripExt(path.getFileName().toString()) + ext;
        Path parent = path.toAbsolutePath().getParent();
        return (parent == null ? Paths.get(name) : parent.resolve(name));
    }

    private static String stripExt(String fileName) {
        int idx = fileName.lastIndexOf('.');
        return idx >= 0 ? fileName.substring(0, idx) : fileName;
    }

    private static String toJson(List<Map<String, Object>> rows) {
        StringBuilder sb = new StringBuilder();
        sb.append("{\n");
        sb.append("  \"generatedAt\": \"").append(Instant.now()).append("\",\n");
        sb.append("  \"files\": [\n");
        for (int i = 0; i < rows.size(); i++) {
            Map<String, Object> r = rows.get(i);
            sb.append("    {\"input\":\"").append(escapeJson(String.valueOf(r.get("input")))).append("\",");
            sb.append("\"output\":\"").append(escapeJson(String.valueOf(r.get("output")))).append("\",");
            sb.append("\"urlsCount\":").append(r.get("urlsCount")).append("}");
            if (i < rows.size() - 1) sb.append(',');
            sb.append('\n');
        }
        sb.append("  ]\n");
        sb.append("}\n");
        return sb.toString();
    }

    private static String escapeJson(String s) {
        return s.replace("\\", "\\\\").replace("\"", "\\\"");
    }
}