← All tasks
javascriptgemini/javascript-t1 #34Not a task: repair changed code

HTML to Plain Text Extractor (javascript, written by Gemini Code Assist)

envgap__gemini__javascript-t1-34

Written by a coding agent; not on GitHubWritten 2026-03-02

01 / FAILURE SIGNATURE

As the study recorded it

SyntaxError: Invalid or unexpected token - broken string literal
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
package.json
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

gemini/javascript-t1 #34 · read the task the agent was given
Gemini Code Assist wrote this javascript project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: HTML to Plain Text Extractor

Write a program that converts HTML documents to clean plain text, intelligently handling formatting, tables, lists, and links while removing all markup and scripts.

FUNCTIONAL REQUIREMENTS:
- Accept an HTML file path as a command-line argument
- Strip all HTML tags, CSS styles, JavaScript, and comments while preserving readable text content
- Convert HTML formatting to plain text equivalents: headings become UPPERCASE with underlines, bold text is wrapped in *asterisks*, lists become indented with bullets (- ) or numbers (1.), horizontal rules become dashed lines
- Convert HTML tables to aligned plain text tables with column padding and separator rows
- Convert hyperlinks to "text [URL]" format, or optionally strip URLs via --no-urls flag
- Preserve paragraph spacing: consecutive block elements get blank line separators
- Handle HTML entities: decode & < >   — etc. to their text equivalents
- Support extracting text from only specific HTML elements via --selector flag (CSS selector syntax, e.g., --selector "article" or --selector ".content")
- Support extracting and listing all URLs found in the document via --extract-urls flag
- Set maximum line width via --width flag (default: 80 characters) with word wrapping
- Support batch conversion of multiple HTML files via --batch flag
- Print the plain text output to console by default
- Save to a file via --output flag (default: same base name with .txt extension)
- If no input is given, generate a sample HTML page with headings, paragraphs, links, tables, lists, images, inline styles, scripts, and HTML entities, then convert it and display both the original HTML and the extracted text
- Handle errors: malformed HTML (parse gracefully), encoding detection, and binary file detection

Create a complete JavaScript project for a clean Ubuntu 22.04 machine with only Node.js 20+ (LTS) installed. Include:
- Source code
- package.json with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

html_extractor.js
const { convert } = require('html-to-text');
const fs = require('fs');
const path = require('path');
const yargs = require('yargs/yargs');
const { hideBin } = require('yargs/helpers');

const argv = yargs(hideBin(process.argv))
    .option('width', { type: 'number', default: 80 })
    .option('no-urls', { type: 'boolean', default: false })
    .argv;

const input = argv._[0];

if (!input) {
    console.log('Usage: node html_extractor.js <input_html> [--width 80] [--no-urls]');
    process.exit(1);
}

try {
    const html = fs.readFileSync(input, 'utf8');
    
    const options = {
        wordwrap: argv.width,
        selectors: [
            { selector: 'a', options: { ignoreHref: argv.noUrls } },
            { selector: 'img', format: 'skip' }
        ]
    };

    const text = convert(html, options);
    
    const outPath = input.replace(/\.html$/, '.txt');
    fs.writeFileSync(outPath, text);
    
    console.log(`Extracted text to ${outPath}`);
    console.log('
--- Preview ---');
    console.log(text.substring(0, 500) + '...');

} catch (e) {
    console.error('Error:', e.message);
}
package.json
{
  "name": "html-extractor",
  "version": "1.0.0",
  "description": "HTML to plain text converter using html-to-text",
  "main": "html_extractor.js",
  "engines": {
    "node": ">=20.0.0"
  },
  "dependencies": {
    "html-to-text": "9.0.5",
    "yargs": "17.7.2"
  },
  "scripts": {
    "start": "node html_extractor.js"
  }
}
README.md
# HTML to Plain Text Extractor (JavaScript)

A tool for converting HTML to plain text using the `html-to-text` library.

## Setup Instructions

1. Ensure Node.js 20+ (LTS) is installed.
2. Install dependencies:
   ```bash
   npm install
   ```

## Run Commands

- **Convert HTML**:
  ```bash
  node html_extractor.js page.html
  ```