← All tasks
javacodex/java-t1 #32Not a task: repair changed code

TF-IDF Search Engine (java, written by Codex)

envgap__codex__java-t1-32

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

Incompatible types: generic type inference mismatch with Map<String Object>
Not a benchmark task.
  • Its repair changed source code, so it is not an environment task.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
pom.xml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/java-t1 #32 · read the task the agent was given
Codex wrote this java project from the task below. It does not run on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: TF-IDF Search Engine

Write a program that builds a TF-IDF (Term Frequency-Inverse Document Frequency) index over a collection of text documents and supports ranked keyword search queries returning the most relevant documents.

FUNCTIONAL REQUIREMENTS:
- Accept a directory of text files as a command-line argument to build the index
- Tokenize documents: split on whitespace and punctuation, convert to lowercase, remove stop words (built-in list of common English stop words like "the", "is", "and", etc.)
- Support optional stemming/lemmatization via --stem flag to group word variants (e.g., "running", "runs", "ran" all map to "run")
- Compute TF-IDF scores for each term in each document using standard formulas: TF = term count / total terms in document, IDF = log(total documents / documents containing term)
- Accept search queries via --query flag and return the top N most relevant documents ranked by cosine similarity between query vector and document vectors (--top flag, default 10)
- Support multi-word queries: compute a query TF-IDF vector and rank documents by similarity
- Support boolean operators in queries via --boolean flag: AND (both terms required), OR (either term), NOT (exclude term)
- Display search results showing: rank, document name, relevance score, and a snippet of the matching text with query terms highlighted
- Save the built index to a file via --save-index flag for reuse without reprocessing
- Load a previously saved index via --load-index flag
- Print index statistics: total documents, total unique terms, average document length, most common terms (top 20)
- Save search results as JSON with --output flag
- If no directory is given, generate a sample corpus of 20 short documents on varied topics (science, sports, technology, cooking, travel), build the index, and demonstrate several search queries with ranked results
- Handle errors: empty documents, binary files in the directory, extremely large documents, and empty queries

Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include:
- Source code
- pom.xml with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd">
  <modelVersion>4.0.0</modelVersion>
  <groupId>org.tmlr.codegen</groupId>
  <artifactId>tfidf-search-engine</artifactId>
  <version>1.0.0</version>
  <name>TF-IDF Search Engine</name>

  <properties>
    <maven.compiler.source>17</maven.compiler.source>
    <maven.compiler.target>17</maven.compiler.target>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
  </properties>

  <dependencyManagement>
    <dependencies>
      <dependency>
        <groupId>com.google.code.gson</groupId>
        <artifactId>gson</artifactId>
        <version>2.11.0</version>
      </dependency>
    </dependencies>
  </dependencyManagement>

  <dependencies>
    <dependency>
      <groupId>com.google.code.gson</groupId>
      <artifactId>gson</artifactId>
    </dependency>
  </dependencies>

  <build>
    <plugins>
      <plugin>
        <groupId>org.apache.maven.plugins</groupId>
        <artifactId>maven-compiler-plugin</artifactId>
        <version>3.13.0</version>
      </plugin>
      <plugin>
        <groupId>org.codehaus.mojo</groupId>
        <artifactId>exec-maven-plugin</artifactId>
        <version>3.3.0</version>
        <configuration>
          <mainClass>TfIdfSearchEngine</mainClass>
        </configuration>
      </plugin>
    </plugins>
  </build>
</project>
README.md
# TF-IDF Search Engine (Java)

Builds a TF-IDF index over text documents and returns ranked search results using cosine similarity, with optional boolean operators.

## Requirements
- Ubuntu 22.04
- JDK 17+
- Maven 3.8+

## Dependencies
- `com.google.code.gson:gson:2.11.0` for index/result serialization

## Build
```bash
mvn -q -DskipTests compile
```

## Run
Build and query:
```bash
mvn -q exec:java -Dexec.args="./docs --query 'machine learning' --top 10"
```

Stemming + boolean:
```bash
mvn -q exec:java -Dexec.args="./docs --stem --boolean --query 'ai AND security NOT malware'"
```

Save/load index:
```bash
mvn -q exec:java -Dexec.args="./docs --save-index tfidf_index.json"
mvn -q exec:java -Dexec.args="--load-index tfidf_index.json --query 'travel budget'"
```

No directory:
```bash
mvn -q exec:java
```
Generates a 20-document sample corpus and runs demo queries.
src/main/java/TfIdfSearchEngine.java
import com.google.gson.Gson;
import com.google.gson.GsonBuilder;
import com.google.gson.reflect.TypeToken;

import java.io.IOException;
import java.lang.reflect.Type;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.time.Instant;
import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Comparator;
import java.util.HashMap;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public final class TfIdfSearchEngine {
    private static final Gson GSON = new GsonBuilder().setPrettyPrinting().create();
    private static final Pattern TOKEN_RX = Pattern.compile("[A-Za-z][A-Za-z0-9']*");
    private static final Set<String> STOP_WORDS = Set.of(
            "a","an","the","is","are","was","were","be","been","being","and","or","but","if","then","else","of","to",
            "in","on","at","for","from","by","with","as","it","its","this","that","these","those","into","about","over",
            "under","between","after","before","during","through","above","below","up","down","out","off","again","further",
            "once","here","there","when","where","why","how","all","any","both","each","few","more","most","other","some",
            "such","no","nor","not","only","own","same","so","than","too","very","can","will","just","do","does","did",
            "doing","have","has","had","having","i","you","he","she","we","they","them","their","our","your","my","me"
    );

    private TfIdfSearchEngine() {}

    private record Config(Path directory, boolean stem, String query, int top, boolean booleanMode, Path saveIndex, Path loadIndex, Path output) {}

    private static final class Document {
        int id;
        String name;
        String path;
        String text;
        int totalTerms;
        Map<String, Integer> counts = new HashMap<>();
        Set<String> termSet = new HashSet<>();
    }

    private static final class Engine {
        boolean useStem;
        List<Document> documents = new ArrayList<>();
        Map<String, Integer> df = new HashMap<>();
        Map<String, Integer> termTotal = new HashMap<>();
        Map<String, Double> idf = new HashMap<>();
        Map<Integer, Map<String, Double>> docVectors = new HashMap<>();
        Map<Integer, Double> docNorm = new HashMap<>();

        Engine(boolean useStem) {
            this.useStem = useStem;
        }

        Map<String, Object> addDocument(String name, String filePath, String text) {
            List<String> terms = tokenize(text, useStem);
            if (terms.isEmpty()) return Map.of("skipped", true, "reason", "empty document");
            Map<String, Integer> counts = new HashMap<>();
            for (String t : terms) counts.put(t, counts.getOrDefault(t, 0) + 1);

            Document d = new Document();
            d.id = documents.size();
            d.name = name;
            d.path = filePath;
            d.text = text;
            d.totalTerms = terms.size();
            d.counts = counts;
            d.termSet = new HashSet<>(counts.keySet());
            documents.add(d);
            return Map.of("skipped", false);
        }

        void buildIndex() {
            df.clear();
            termTotal.clear();
            for (Document d : documents) {
                for (Map.Entry<String, Integer> e : d.counts.entrySet()) {
                    termTotal.put(e.getKey(), termTotal.getOrDefault(e.getKey(), 0) + e.getValue());
                }
                for (String term : d.termSet) {
                    df.put(term, df.getOrDefault(term, 0) + 1);
                }
            }
            int n = documents.size();
            idf.clear();
            for (Map.Entry<String, Integer> e : df.entrySet()) {
                idf.put(e.getKey(), Math.log((double) n / e.getValue()));
            }
            docVectors.clear();
            docNorm.clear();
            for (Document d : documents) {
                Map<String, Double> vec = new HashMap<>();
                double normSq = 0.0;
                for (Map.Entry<String, Integer> e : d.counts.entrySet()) {
                    double tf = (double) e.getValue() / d.totalTerms;
                    double w = tf * idf.getOrDefault(e.getKey(), 0.0);
                    vec.put(e.getKey(), w);
                    normSq += w * w;
                }
                docVectors.put(d.id, vec);
                docNorm.put(d.id, Math.sqrt(normSq));
            }
        }

        Map<String, Object> stats() {
            int totalDocs = documents.size();
            double avgLen = totalDocs == 0 ? 0.0 : documents.stream().mapToInt(d -> d.totalTerms).average().orElse(0.0);
            List<Map<String, Object>> common = termTotal.entrySet().stream()
                    .sorted((a, b) -> {
                        int c = Integer.compare(b.getValue(), a.getValue());
                        if (c != 0) return c;
                        return a.getKey().compareTo(b.getKey());
                    })
                    .limit(20)
                    .map(e -> Map.of("term", e.getKey(), "count", e.getValue()))
                    .toList();
            return Map.of(
                    "total_documents", totalDocs,
                    "total_unique_terms", df.size(),
                    "average_document_length", avgLen,
                    "most_common_terms", common
            );
        }

        List<Map<String, Object>> search(String query, int top, boolean booleanMode) {
            if (query == null || query.trim().isEmpty()) throw new IllegalArgumentException("Empty query is not allowed");
            List<String> qTokens = tokenize(query, useStem);
            if (qTokens.isEmpty()) throw new IllegalArgumentException("Query contains no searchable terms");

            Map<String, Integer> qCounts = new HashMap<>();
            for (String t : qTokens) qCounts.put(t, qCounts.getOrDefault(t, 0) + 1);
            Map<String, Double> qVec = new HashMap<>();
            double qNormSq = 0.0;
            for (Map.Entry<String, Integer> e : qCounts.entrySet()) {
                double tf = (double) e.getValue() / qTokens.size();
                double w = tf * idf.getOrDefault(e.getKey(), 0.0);
                qVec.put(e.getKey(), w);
                qNormSq += w * w;
            }
            double qNorm = Math.sqrt(qNormSq);

            List<Document> candidates = documents;
            if (booleanMode) {
                List<String> postfix = parseBooleanQuery(query, useStem);
                candidates = new ArrayList<>();
                for (Document d : documents) if (evalBooleanPostfix(postfix, d.termSet)) candidates.add(d);
            }

            List<Map<String, Object>> ranked = new ArrayList<>();
            for (Document d : candidates) {
                Map<String, Double> dVec = docVectors.getOrDefault(d.id, Map.of());
                double dot = 0.0;
                for (Map.Entry<String, Double> e : qVec.entrySet()) {
                    dot += e.getValue() * dVec.getOrDefault(e.getKey(), 0.0);
                }
                double dNorm = docNorm.getOrDefault(d.id, 0.0);
                double score = (qNorm == 0.0 || dNorm == 0.0) ? 0.0 : dot / (qNorm * dNorm);
                if (score > 0 || booleanMode) {
                    ranked.add(new LinkedHashMap<>(Map.of(
                            "document", d.name,
                            "path", d.path,
                            "score", score,
                            "snippet", snippet(d.text, qTokens)
                    )));
                }
            }
            ranked.sort(Comparator.<Map<String, Object>, Double>comparing(r -> -((Number) r.get("score")).doubleValue())
                    .thenComparing(r -> (String) r.get("document")));
            return ranked.subList(0, Math.min(top, ranked.size()));
        }

        void saveIndex(Path output) throws IOException {
            Map<String, Object> payload = new LinkedHashMap<>();
            payload.put("use_stem", useStem);
            payload.put("documents", documents);
            payload.put("df", df);
            payload.put("term_total", termTotal);
            payload.put("idf", idf);
            payload.put("doc_norm", docNorm);
            Files.writeString(output, GSON.toJson(payload), StandardCharsets.UTF_8);
        }

        static Engine loadIndex(Path input) throws IOException {
            Type mapType = new TypeToken<Map<String, Object>>() {}.getType();
            Map<String, Object> payload = GSON.fromJson(Files.readString(input, StandardCharsets.UTF_8), mapType);
            boolean stem = Boolean.TRUE.equals(payload.get("use_stem"));
            Engine e = new Engine(stem);

            @SuppressWarnings("unchecked")
            List<Map<String, Object>> docs = (List<Map<String, Object>>) payload.get("documents");
            for (Map<String, Object> raw : docs) {
                Document d = new Document();
                d.id = ((Number) raw.get("id")).intValue();
                d.name = (String) raw.get("name");
                d.path = (String) raw.get("path");
                d.text = (String) raw.get("text");
                d.totalTerms = ((Number) raw.get("totalTerms")).intValue();
                @SuppressWarnings("unchecked")
                Map<String, Number> c = (Map<String, Number>) raw.get("counts");
                for (Map.Entry<String, Number> ce : c.entrySet()) d.counts.put(ce.getKey(), ce.getValue().intValue());
                d.termSet = new HashSet<>(d.counts.keySet());
                e.documents.add(d);
            }
            e.buildIndex();
            return e;
        }
    }

    private static String stemToken(String word) {
        if ("ran".equals(word)) return "run";
        String w = word;
        if (w.endsWith("ies") && w.length() > 4) w = w.substring(0, w.length() - 3) + "y";
        else if (w.endsWith("ing") && w.length() > 5) w = w.substring(0, w.length() - 3);
        else if (w.endsWith("ed") && w.length() > 4) w = w.substring(0, w.length() - 2);
        else if (w.endsWith("es") && w.length() > 4) w = w.substring(0, w.length() - 2);
        else if (w.endsWith("s") && w.length() > 3) w = w.substring(0, w.length() - 1);
        if (w.endsWith("nn")) w = w.substring(0, w.length() - 1);
        return w;
    }

    private static List<String> tokenize(String text, boolean useStem) {
        List<String> out = new ArrayList<>();
        Matcher m = TOKEN_RX.matcher(text);
        while (m.find()) {
            String token = m.group().toLowerCase(Locale.ROOT).replace("'", "");
            if (token.isEmpty() || STOP_WORDS.contains(token)) continue;
            if (useStem) token = stemToken(token);
            if (!token.isEmpty() && !STOP_WORDS.contains(token)) out.add(token);
        }
        return out;
    }

    private static List<String> parseBooleanQuery(String query, boolean useStem) {
        Matcher m = Pattern.compile("\\(|\\)|AND|OR|NOT|[A-Za-z][A-Za-z0-9']*", Pattern.CASE_INSENSITIVE).matcher(query);
        List<String> tokens = new ArrayList<>();
        while (m.find()) {
            String t = m.group();
            if (t.equalsIgnoreCase("AND") || t.equalsIgnoreCase("OR") || t.equalsIgnoreCase("NOT")) {
                tokens.add(t.toUpperCase(Locale.ROOT));
            } else {
                String term = t.toLowerCase(Locale.ROOT).replace("'", "");
                if (useStem) term = stemToken(term);
                tokens.add(term);
            }
        }

        Map<String, Integer> prec = Map.of("OR", 1, "AND", 2, "NOT", 3);
        List<String> output = new ArrayList<>();
        ArrayDeque<String> stack = new ArrayDeque<>();
        for (String t : tokens) {
            if ("(".equals(t)) stack.push(t);
            else if (")".equals(t)) {
                while (!stack.isEmpty() && !"(".equals(stack.peek())) output.add(stack.pop());
                if (!stack.isEmpty() && "(".equals(stack.peek())) stack.pop();
            } else if (prec.containsKey(t)) {
                while (!stack.isEmpty() && prec.containsKey(stack.peek()) && prec.get(stack.peek()) >= prec.get(t)) {
                    output.add(stack.pop());
                }
                stack.push(t);
            } else output.add(t);
        }
        while (!stack.isEmpty()) output.add(stack.pop());
        return output;
    }

    private static boolean evalBooleanPostfix(List<String> postfix, Set<String> terms) {
        ArrayDeque<Boolean> stack = new ArrayDeque<>();
        for (String t : postfix) {
            switch (t) {
                case "NOT" -> {
                    if (stack.isEmpty()) return false;
                    stack.push(!stack.pop());
                }
                case "AND" -> {
                    if (stack.size() < 2) return false;
                    boolean b = stack.pop();
                    boolean a = stack.pop();
                    stack.push(a && b);
                }
                case "OR" -> {
                    if (stack.size() < 2) return false;
                    boolean b = stack.pop();
                    boolean a = stack.pop();
                    stack.push(a || b);
                }
                default -> stack.push(terms.contains(t));
            }
        }
        return stack.size() == 1 && stack.peek();
    }

    private static String snippet(String text, List<String> qTerms) {
        if (qTerms.isEmpty()) return text.replaceAll("\\s+", " ").substring(0, Math.min(180, text.length()));
        String lowerText = text.toLowerCase(Locale.ROOT);
        int best = -1;
        for (String t : qTerms) {
            int p = lowerText.indexOf(t.toLowerCase(Locale.ROOT));
            if (p >= 0 && (best < 0 || p < best)) best = p;
        }
        if (best < 0) return text.replaceAll("\\s+", " ").substring(0, Math.min(180, text.length()));
        int start = Math.max(0, best - 60);
        int end = Math.min(text.length(), best + 120);
        String s = text.substring(start, end).replaceAll("\\s+", " ").trim();
        if (start > 0) s = "..." + s;
        if (end < text.length()) s += "...";
        for (String t : qTerms) s = s.replaceAll("(?i)\\b" + Pattern.quote(t) + "\\b", "**$0**");
        return s;
    }

    private static Config parseArgs(String[] args) {
        Path directory = null;
        boolean stem = false;
        String query = null;
        int top = 10;
        boolean booleanMode = false;
        Path saveIndex = null;
        Path loadIndex = null;
        Path output = Path.of("search_results.json");

        for (int i = 0; i < args.length; i++) {
            String arg = args[i];
            if (!arg.startsWith("--")) {
                if (directory == null) directory = Path.of(arg).toAbsolutePath();
                else throw new IllegalArgumentException("Unexpected argument: " + arg);
                continue;
            }
            switch (arg) {
                case "--stem" -> stem = true;
                case "--boolean" -> booleanMode = true;
                case "--query" -> query = nextValue(args, ++i, "--query");
                case "--top" -> top = Integer.parseInt(nextValue(args, ++i, "--top"));
                case "--save-index" -> saveIndex = Path.of(nextValue(args, ++i, "--save-index")).toAbsolutePath();
                case "--load-index" -> loadIndex = Path.of(nextValue(args, ++i, "--load-index")).toAbsolutePath();
                case "--output" -> output = Path.of(nextValue(args, ++i, "--output")).toAbsolutePath();
                default -> throw new IllegalArgumentException("Unknown option: " + arg);
            }
        }
        if (top <= 0) throw new IllegalArgumentException("--top must be positive");
        return new Config(directory, stem, query, top, booleanMode, saveIndex, loadIndex, output);
    }

    private static String nextValue(String[] args, int idx, String flag) {
        if (idx >= args.length) throw new IllegalArgumentException("Missing value for " + flag);
        return args[idx];
    }

    private static boolean isBinary(byte[] bytes) {
        int limit = Math.min(bytes.length, 1024);
        for (int i = 0; i < limit; i++) if (bytes[i] == 0) return true;
        return false;
    }

    private static Engine buildFromDirectory(Path directory, boolean stem) throws IOException {
        Engine e = new Engine(stem);
        for (Path p : Files.list(directory).sorted().toList()) {
            if (!Files.isRegularFile(p)) continue;
            byte[] bytes = Files.readAllBytes(p);
            if (isBinary(bytes)) {
                System.err.println("Warning: skipped binary file " + p);
                continue;
            }
            if (bytes.length > 10 * 1024 * 1024) {
                System.err.println("Warning: skipped very large file " + p);
                continue;
            }
            String text = new String(bytes, StandardCharsets.UTF_8);
            Map<String, Object> res = e.addDocument(p.getFileName().toString(), p.toString(), text);
            if (Boolean.TRUE.equals(res.get("skipped"))) System.err.println("Warning: skipped " + p + " (" + res.get("reason") + ")");
        }
        e.buildIndex();
        return e;
    }

    private static Path createSampleCorpus() throws IOException {
        Path dir = Path.of("sample_corpus").toAbsolutePath();
        Files.createDirectories(dir);
        Map<String, String> docs = Map.ofEntries(
                Map.entry("science_quantum.txt", "Quantum physics studies particles, waves, uncertainty, and entanglement in tiny systems."),
                Map.entry("science_astronomy.txt", "Astronomy explores stars, galaxies, black holes, and telescopes that map distant planets."),
                Map.entry("science_biology.txt", "Biology examines cells, genes, evolution, and ecosystems in living organisms."),
                Map.entry("science_climate.txt", "Climate science tracks greenhouse gases, weather patterns, and long term temperature changes."),
                Map.entry("sports_football.txt", "Football strategy includes passing, defense, pressing, and midfield control during competition."),
                Map.entry("sports_basketball.txt", "Basketball players practice shooting, dribbling, spacing, and fast breaks to win games."),
                Map.entry("sports_running.txt", "Running performance improves with interval training, nutrition, and recovery routines."),
                Map.entry("sports_tennis.txt", "Tennis matches require serves, volleys, footwork, and tactical shot placement."),
                Map.entry("tech_ai.txt", "Artificial intelligence uses machine learning models, data pipelines, and optimization methods."),
                Map.entry("tech_security.txt", "Cybersecurity protects networks with encryption, monitoring, authentication, and incident response."),
                Map.entry("tech_cloud.txt", "Cloud computing provides scalable storage, virtual machines, and managed application services."),
                Map.entry("tech_web.txt", "Web development combines html css javascript frameworks, testing, and deployment automation."),
                Map.entry("cooking_pasta.txt", "Pasta recipes use olive oil, garlic, tomatoes, basil, and careful timing for sauce texture."),
                Map.entry("cooking_baking.txt", "Baking bread needs flour, yeast, hydration, proofing, and oven temperature control."),
                Map.entry("cooking_spices.txt", "Spice blends balance heat, sweetness, acidity, and aroma in regional cuisine."),
                Map.entry("cooking_salad.txt", "Fresh salad preparation focuses on greens, dressing, crunch, and seasonal produce."),
                Map.entry("travel_mountains.txt", "Mountain travel involves hiking trails, altitude planning, weather safety, and local guides."),
                Map.entry("travel_cities.txt", "City travel highlights museums, transit cards, neighborhoods, and cultural landmarks."),
                Map.entry("travel_beaches.txt", "Beach vacations include snorkeling, tides, sun protection, and coastal food markets."),
                Map.entry("travel_budget.txt", "Budget travel uses hostels, public transport, off season fares, and itinerary planning.")
        );
        for (Map.Entry<String, String> e : docs.entrySet()) {
            Files.writeString(dir.resolve(e.getKey()), e.getValue() + "\n", StandardCharsets.UTF_8);
        }
        return dir;
    }

    private static void printStats(Map<String, Object> stats) {
        System.out.println("Index statistics:");
        System.out.println("  Total documents: " + stats.get("total_documents"));
        System.out.println("  Total unique terms: " + stats.get("total_unique_terms"));
        System.out.printf(Locale.US, "  Average document length: %.2f terms%n", ((Number) stats.get("average_document_length")).doubleValue());
        System.out.println("  Most common terms (top 20):");
        @SuppressWarnings("unchecked")
        List<Map<String, Object>> common = (List<Map<String, Object>>) stats.get("most_common_terms");
        for (Map<String, Object> t : common) System.out.println("    - " + t.get("term") + ": " + t.get("count"));
    }

    private static void printResults(String query, List<Map<String, Object>> results) {
        System.out.println("\nQuery: " + query);
        if (results.isEmpty()) {
            System.out.println("  No matching documents.");
            return;
        }
        for (int i = 0; i < results.size(); i++) {
            Map<String, Object> r = results.get(i);
            System.out.printf(Locale.US, "  %d. %s | score=%.6f%n", i + 1, r.get("document"), ((Number) r.get("score")).doubleValue());
            System.out.println("     " + r.get("snippet"));
        }
    }

    public static void main(String[] args) {
        try {
            Config cfg = parseArgs(args);
            Engine engine;
            boolean usedSample = false;

            if (cfg.loadIndex() != null) {
                engine = Engine.loadIndex(cfg.loadIndex());
            } else {
                Path dir = cfg.directory();
                if (dir == null) {
                    dir = createSampleCorpus();
                    usedSample = true;
                }
                if (!Files.exists(dir) || !Files.isDirectory(dir)) throw new IllegalArgumentException("Directory does not exist: " + dir);
                engine = buildFromDirectory(dir, cfg.stem());
            }

            if (engine.documents.isEmpty()) throw new IllegalStateException("No valid text documents were indexed");
            if (cfg.saveIndex() != null) {
                engine.saveIndex(cfg.saveIndex());
                System.out.println("Saved index: " + cfg.saveIndex());
            }

            Map<String, Object> stats = engine.stats();
            printStats(stats);

            Map<String, Object> payload = new LinkedHashMap<>();
            payload.put("generated_at", Instant.now().toString());
            payload.put("query", cfg.query());
            payload.put("top", cfg.top());
            payload.put("boolean_mode", cfg.booleanMode());
            payload.put("stats", stats);
            payload.put("results", new ArrayList<>());

            if (cfg.query() != null) {
                List<Map<String, Object>> results = engine.search(cfg.query(), cfg.top(), cfg.booleanMode());
                printResults(cfg.query(), results);
                payload.put("results", results);
            } else if (usedSample) {
                List<String> queries = List.of("quantum physics", "pasta recipe", "travel AND budget", "ai AND security NOT malware");
                List<Map<String, Object>> demo = new ArrayList<>();
                for (String q : queries) {
                    boolean booleanMode = q.matches("(?i).*(\\bAND\\b|\\bOR\\b|\\bNOT\\b).*");
                    List<Map<String, Object>> results = engine.search(q, cfg.top(), booleanMode);
                    printResults(q, results);
                    demo.add(Map.of("query", q, "boolean_mode", booleanMode, "items", results));
                }
                payload.put("results", demo);
            } else {
                System.out.println("No query provided. Use --query to search.");
            }

            Files.writeString(cfg.output(), GSON.toJson(payload), StandardCharsets.UTF_8);
            System.out.println("Saved search results JSON: " + cfg.output());
        } catch (Exception ex) {
            System.err.println("Error: " + ex.getMessage());
            System.exit(1);
        }
    }
}