← All tasks
cppcodex/cpp-t1 #32Not a task: already works

TF-IDF Search Engine (cpp, written by Codex)

envgap__codex__cpp-t1-32

Written by a coding agent; not on GitHubWritten 2026-03-03

01 / FAILURE SIGNATURE

As the study recorded it

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
Not freshly verified
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / TASK AND FAILURE

codex/cpp-t1 #32 · read the task the agent was given
Codex wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.

Task given to the agent:

TASK: TF-IDF Search Engine

Write a program that builds a TF-IDF (Term Frequency-Inverse Document Frequency) index over a collection of text documents and supports ranked keyword search queries returning the most relevant documents.

FUNCTIONAL REQUIREMENTS:
- Accept a directory of text files as a command-line argument to build the index
- Tokenize documents: split on whitespace and punctuation, convert to lowercase, remove stop words (built-in list of common English stop words like "the", "is", "and", etc.)
- Support optional stemming/lemmatization via --stem flag to group word variants (e.g., "running", "runs", "ran" all map to "run")
- Compute TF-IDF scores for each term in each document using standard formulas: TF = term count / total terms in document, IDF = log(total documents / documents containing term)
- Accept search queries via --query flag and return the top N most relevant documents ranked by cosine similarity between query vector and document vectors (--top flag, default 10)
- Support multi-word queries: compute a query TF-IDF vector and rank documents by similarity
- Support boolean operators in queries via --boolean flag: AND (both terms required), OR (either term), NOT (exclude term)
- Display search results showing: rank, document name, relevance score, and a snippet of the matching text with query terms highlighted
- Save the built index to a file via --save-index flag for reuse without reprocessing
- Load a previously saved index via --load-index flag
- Print index statistics: total documents, total unique terms, average document length, most common terms (top 20)
- Save search results as JSON with --output flag
- If no directory is given, generate a sample corpus of 20 short documents on varied topics (science, sports, technology, cooking, travel), build the index, and demonstrate several search queries with ranked results
- Handle errors: empty documents, binary files in the directory, extremely large documents, and empty queries

Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]

05 / FILES

The project as the agent wrote it

3 files, exactly as written, before any repair.

CMakeLists.txt
cmake_minimum_required(VERSION 3.22)
project(tfidf_search_engine LANGUAGES CXX)

set(CMAKE_CXX_STANDARD 20)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)

include(FetchContent)

FetchContent_Declare(
  nlohmann_json
  GIT_REPOSITORY https://github.com/nlohmann/json.git
  GIT_TAG v3.11.3
)

FetchContent_MakeAvailable(nlohmann_json)

add_executable(tfidf_search_engine src/main.cpp)
target_link_libraries(tfidf_search_engine PRIVATE nlohmann_json::nlohmann_json)
README.md
# TF-IDF Search Engine (C++)

Builds a TF-IDF index over text documents and performs ranked keyword search with cosine similarity and optional boolean query operators.

## Requirements
- Ubuntu 22.04
- G++ 12+
- CMake 3.22+

## Dependencies (Pinned)
- `nlohmann/json` `v3.11.3` for index/result serialization

## Build
```bash
cmake -S . -B build
cmake --build build -j
```

## Run
Build and query:
```bash
./build/tfidf_search_engine ./docs --query "machine learning" --top 10
```

With stemming and boolean operators:
```bash
./build/tfidf_search_engine ./docs --stem --boolean --query "ai AND security NOT malware"
```

Save/load index:
```bash
./build/tfidf_search_engine ./docs --save-index tfidf_index.json
./build/tfidf_search_engine --load-index tfidf_index.json --query "travel budget"
```

No directory:
```bash
./build/tfidf_search_engine
```
Generates a 20-document sample corpus and runs demo queries.
src/main.cpp
#include <algorithm>
#include <cmath>
#include <cctype>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <map>
#include <numeric>
#include <optional>
#include <regex>
#include <set>
#include <sstream>
#include <stdexcept>
#include <string>
#include <unordered_map>
#include <unordered_set>
#include <utility>
#include <vector>

#include <nlohmann/json.hpp>

namespace fs = std::filesystem;
using json = nlohmann::json;

struct Config {
  std::optional<fs::path> directory;
  bool stem = false;
  std::optional<std::string> query;
  int top = 10;
  bool boolean_mode = false;
  std::optional<fs::path> save_index;
  std::optional<fs::path> load_index;
  fs::path output = "search_results.json";
};

struct Document {
  int id{};
  std::string name;
  std::string path;
  std::string text;
  int total_terms{};
  std::unordered_map<std::string, int> counts;
  std::unordered_set<std::string> term_set;
};

static const std::unordered_set<std::string> STOP_WORDS = {
    "a","an","the","is","are","was","were","be","been","being","and","or","but","if","then","else","of","to",
    "in","on","at","for","from","by","with","as","it","its","this","that","these","those","into","about","over",
    "under","between","after","before","during","through","above","below","up","down","out","off","again","further",
    "once","here","there","when","where","why","how","all","any","both","each","few","more","most","other","some",
    "such","no","nor","not","only","own","same","so","than","too","very","can","will","just","do","does","did",
    "doing","have","has","had","having","i","you","he","she","we","they","them","their","our","your","my","me"
};

static std::string lower(std::string s) {
  for (char& c : s) c = static_cast<char>(std::tolower(static_cast<unsigned char>(c)));
  return s;
}

static std::string trim(const std::string& s) {
  size_t b = 0;
  while (b < s.size() && std::isspace(static_cast<unsigned char>(s[b]))) ++b;
  size_t e = s.size();
  while (e > b && std::isspace(static_cast<unsigned char>(s[e - 1]))) --e;
  return s.substr(b, e - b);
}

static std::string stem_token(std::string w) {
  if (w == "ran") return "run";
  if (w.size() > 4 && w.ends_with("ies")) w = w.substr(0, w.size() - 3) + "y";
  else if (w.size() > 5 && w.ends_with("ing")) w = w.substr(0, w.size() - 3);
  else if (w.size() > 4 && w.ends_with("ed")) w = w.substr(0, w.size() - 2);
  else if (w.size() > 4 && w.ends_with("es")) w = w.substr(0, w.size() - 2);
  else if (w.size() > 3 && w.ends_with("s")) w = w.substr(0, w.size() - 1);
  if (w.size() > 2 && w.ends_with("nn")) w = w.substr(0, w.size() - 1);
  return w;
}

static std::vector<std::string> tokenize(const std::string& text, bool use_stem) {
  std::vector<std::string> out;
  static const std::regex rx(R"([A-Za-z][A-Za-z0-9']*)");
  for (std::sregex_iterator it(text.begin(), text.end(), rx), end; it != end; ++it) {
    std::string token = lower(it->str());
    token.erase(std::remove(token.begin(), token.end(), '\''), token.end());
    if (token.empty() || STOP_WORDS.contains(token)) continue;
    if (use_stem) token = stem_token(token);
    if (!token.empty() && !STOP_WORDS.contains(token)) out.push_back(token);
  }
  return out;
}

static bool is_binary(const std::string& data) {
  size_t n = std::min<size_t>(data.size(), 1024);
  for (size_t i = 0; i < n; ++i) if (data[i] == '\0') return true;
  return false;
}

class Engine {
 public:
  explicit Engine(bool use_stem) : use_stem_(use_stem) {}

  bool add_document(const std::string& name, const std::string& file_path, const std::string& text, std::string& reason) {
    auto terms = tokenize(text, use_stem_);
    if (terms.empty()) {
      reason = "empty document";
      return false;
    }
    Document d;
    d.id = static_cast<int>(documents_.size());
    d.name = name;
    d.path = file_path;
    d.text = text;
    d.total_terms = static_cast<int>(terms.size());
    for (const auto& t : terms) d.counts[t] += 1;
    for (const auto& kv : d.counts) d.term_set.insert(kv.first);
    documents_.push_back(std::move(d));
    return true;
  }

  void build_index() {
    df_.clear();
    term_total_.clear();
    for (const auto& d : documents_) {
      for (const auto& kv : d.counts) term_total_[kv.first] += kv.second;
      for (const auto& term : d.term_set) df_[term] += 1;
    }
    int n = static_cast<int>(documents_.size());
    idf_.clear();
    for (const auto& kv : df_) {
      idf_[kv.first] = std::log(static_cast<double>(n) / kv.second);
    }
    doc_vectors_.clear();
    doc_norm_.clear();
    for (const auto& d : documents_) {
      std::unordered_map<std::string, double> vec;
      double norm_sq = 0.0;
      for (const auto& kv : d.counts) {
        double tf = static_cast<double>(kv.second) / d.total_terms;
        double w = tf * idf_[kv.first];
        vec[kv.first] = w;
        norm_sq += w * w;
      }
      doc_vectors_[d.id] = std::move(vec);
      doc_norm_[d.id] = std::sqrt(norm_sq);
    }
  }

  json stats() const {
    int total_docs = static_cast<int>(documents_.size());
    double avg_len = total_docs == 0 ? 0.0 : std::accumulate(documents_.begin(), documents_.end(), 0.0,
        [](double acc, const Document& d) { return acc + d.total_terms; }) / total_docs;
    std::vector<std::pair<std::string, int>> common(term_total_.begin(), term_total_.end());
    std::sort(common.begin(), common.end(), [](const auto& a, const auto& b) {
      if (a.second != b.second) return a.second > b.second;
      return a.first < b.first;
    });
    json top = json::array();
    for (size_t i = 0; i < std::min<size_t>(20, common.size()); ++i) {
      top.push_back({{"term", common[i].first}, {"count", common[i].second}});
    }
    return {
        {"total_documents", total_docs},
        {"total_unique_terms", static_cast<int>(df_.size())},
        {"average_document_length", avg_len},
        {"most_common_terms", top}
    };
  }

  std::vector<json> search(const std::string& query, int top, bool boolean_mode) const {
    if (trim(query).empty()) throw std::runtime_error("Empty query is not allowed");
    auto q_tokens = tokenize(query, use_stem_);
    if (q_tokens.empty()) throw std::runtime_error("Query contains no searchable terms");

    std::unordered_map<std::string, int> q_counts;
    for (const auto& t : q_tokens) q_counts[t] += 1;
    std::unordered_map<std::string, double> q_vec;
    double q_norm_sq = 0.0;
    for (const auto& kv : q_counts) {
      double tf = static_cast<double>(kv.second) / q_tokens.size();
      double w = tf * (idf_.contains(kv.first) ? idf_.at(kv.first) : 0.0);
      q_vec[kv.first] = w;
      q_norm_sq += w * w;
    }
    double q_norm = std::sqrt(q_norm_sq);

    std::vector<const Document*> candidates;
    if (boolean_mode) {
      auto postfix = parse_boolean_query(query);
      for (const auto& d : documents_) {
        if (eval_boolean_postfix(postfix, d.term_set)) candidates.push_back(&d);
      }
    } else {
      for (const auto& d : documents_) candidates.push_back(&d);
    }

    std::vector<json> ranked;
    for (const auto* d : candidates) {
      double dot = 0.0;
      const auto& d_vec = doc_vectors_.at(d->id);
      for (const auto& kv : q_vec) {
        auto it = d_vec.find(kv.first);
        if (it != d_vec.end()) dot += kv.second * it->second;
      }
      double d_norm = doc_norm_.contains(d->id) ? doc_norm_.at(d->id) : 0.0;
      double score = (q_norm == 0.0 || d_norm == 0.0) ? 0.0 : dot / (q_norm * d_norm);
      if (score > 0.0 || boolean_mode) {
        ranked.push_back({
            {"document", d->name},
            {"path", d->path},
            {"score", score},
            {"snippet", snippet(d->text, q_tokens)}
        });
      }
    }
    std::sort(ranked.begin(), ranked.end(), [](const json& a, const json& b) {
      double sa = a["score"].get<double>();
      double sb = b["score"].get<double>();
      if (sa != sb) return sa > sb;
      return a["document"].get<std::string>() < b["document"].get<std::string>();
    });
    if (static_cast<int>(ranked.size()) > top) ranked.resize(static_cast<size_t>(top));
    return ranked;
  }

  void save_index(const fs::path& out) const {
    json docs = json::array();
    for (const auto& d : documents_) {
      docs.push_back({
          {"id", d.id},
          {"name", d.name},
          {"path", d.path},
          {"text", d.text},
          {"total_terms", d.total_terms},
          {"counts", d.counts}
      });
    }
    json payload{
        {"use_stem", use_stem_},
        {"documents", docs},
        {"df", df_},
        {"term_total", term_total_},
        {"idf", idf_},
        {"doc_norm", doc_norm_}
    };
    std::ofstream f(out);
    if (!f) throw std::runtime_error("Failed to write index: " + out.string());
    f << payload.dump(2) << "\n";
  }

  static Engine load_index(const fs::path& p) {
    std::ifstream f(p);
    if (!f) throw std::runtime_error("Failed to read index file: " + p.string());
    json payload = json::parse(f);
    Engine e(payload["use_stem"].get<bool>());
    for (const auto& raw : payload["documents"]) {
      Document d;
      d.id = raw["id"].get<int>();
      d.name = raw["name"].get<std::string>();
      d.path = raw["path"].get<std::string>();
      d.text = raw["text"].get<std::string>();
      d.total_terms = raw["total_terms"].get<int>();
      d.counts = raw["counts"].get<std::unordered_map<std::string, int>>();
      for (const auto& kv : d.counts) d.term_set.insert(kv.first);
      e.documents_.push_back(std::move(d));
    }
    e.build_index();
    return e;
  }

  bool empty() const { return documents_.empty(); }

 private:
  bool use_stem_;
  std::vector<Document> documents_;
  std::unordered_map<std::string, int> df_;
  std::unordered_map<std::string, int> term_total_;
  std::unordered_map<std::string, double> idf_;
  std::unordered_map<int, std::unordered_map<std::string, double>> doc_vectors_;
  std::unordered_map<int, double> doc_norm_;

  std::vector<std::string> parse_boolean_query(const std::string& query) const {
    std::regex rx(R"(\(|\)|AND|OR|NOT|[A-Za-z][A-Za-z0-9']*)", std::regex::icase);
    std::vector<std::string> tokens;
    for (std::sregex_iterator it(query.begin(), query.end(), rx), end; it != end; ++it) {
      std::string t = it->str();
      std::string u = lower(t);
      if (u == "and" || u == "or" || u == "not") tokens.push_back(std::string{static_cast<char>(std::toupper(u[0])), static_cast<char>(std::toupper(u[1])), static_cast<char>(std::toupper(u[2]))});
      else {
        t = lower(t);
        t.erase(std::remove(t.begin(), t.end(), '\''), t.end());
        if (use_stem_) t = stem_token(t);
        tokens.push_back(t);
      }
    }
    std::unordered_map<std::string, int> prec{{"OR",1},{"AND",2},{"NOT",3}};
    std::vector<std::string> out;
    std::vector<std::string> st;
    for (const auto& t : tokens) {
      if (t == "(") st.push_back(t);
      else if (t == ")") {
        while (!st.empty() && st.back() != "(") { out.push_back(st.back()); st.pop_back(); }
        if (!st.empty() && st.back() == "(") st.pop_back();
      } else if (prec.contains(t)) {
        while (!st.empty() && prec.contains(st.back()) && prec[st.back()] >= prec[t]) {
          out.push_back(st.back()); st.pop_back();
        }
        st.push_back(t);
      } else out.push_back(t);
    }
    while (!st.empty()) { out.push_back(st.back()); st.pop_back(); }
    return out;
  }

  static bool eval_boolean_postfix(const std::vector<std::string>& postfix, const std::unordered_set<std::string>& terms) {
    std::vector<bool> stack;
    for (const auto& t : postfix) {
      if (t == "NOT") {
        if (stack.empty()) return false;
        bool a = stack.back(); stack.pop_back();
        stack.push_back(!a);
      } else if (t == "AND" || t == "OR") {
        if (stack.size() < 2) return false;
        bool b = stack.back(); stack.pop_back();
        bool a = stack.back(); stack.pop_back();
        stack.push_back(t == "AND" ? (a && b) : (a || b));
      } else {
        stack.push_back(terms.contains(t));
      }
    }
    return stack.size() == 1 && stack.back();
  }

  static std::string snippet(const std::string& text, const std::vector<std::string>& q_terms) {
    if (q_terms.empty()) return text.substr(0, std::min<size_t>(180, text.size()));
    std::string lower_text = lower(text);
    int best = -1;
    for (const auto& t : q_terms) {
      auto p = lower_text.find(lower(t));
      if (p != std::string::npos && (best < 0 || static_cast<int>(p) < best)) best = static_cast<int>(p);
    }
    if (best < 0) return text.substr(0, std::min<size_t>(180, text.size()));
    int start = std::max(0, best - 60);
    int end = std::min<int>(static_cast<int>(text.size()), best + 120);
    std::string s = text.substr(static_cast<size_t>(start), static_cast<size_t>(end - start));
    s = std::regex_replace(s, std::regex(R"(\s+)"), " ");
    if (start > 0) s = "..." + s;
    if (end < static_cast<int>(text.size())) s += "...";
    for (const auto& t : q_terms) {
      s = std::regex_replace(s, std::regex("\\b" + t + "\\b", std::regex::icase), "**$&**");
    }
    return s;
  }
};

static Config parse_args(int argc, char** argv) {
  Config cfg;
  for (int i = 1; i < argc; ++i) {
    std::string arg = argv[i];
    if (!arg.starts_with("--")) {
      if (!cfg.directory.has_value()) cfg.directory = fs::absolute(arg);
      else throw std::runtime_error("Unexpected argument: " + arg);
      continue;
    }
    if (arg == "--stem") cfg.stem = true;
    else if (arg == "--boolean") cfg.boolean_mode = true;
    else if (arg == "--query") {
      if (i + 1 >= argc) throw std::runtime_error("Missing value for --query");
      cfg.query = std::string(argv[++i]);
    } else if (arg == "--top") {
      if (i + 1 >= argc) throw std::runtime_error("Missing value for --top");
      cfg.top = std::stoi(argv[++i]);
    } else if (arg == "--save-index") {
      if (i + 1 >= argc) throw std::runtime_error("Missing value for --save-index");
      cfg.save_index = fs::absolute(argv[++i]);
    } else if (arg == "--load-index") {
      if (i + 1 >= argc) throw std::runtime_error("Missing value for --load-index");
      cfg.load_index = fs::absolute(argv[++i]);
    } else if (arg == "--output") {
      if (i + 1 >= argc) throw std::runtime_error("Missing value for --output");
      cfg.output = fs::absolute(argv[++i]);
    } else throw std::runtime_error("Unknown option: " + arg);
  }
  if (cfg.top <= 0) throw std::runtime_error("--top must be a positive integer");
  return cfg;
}

static fs::path create_sample_corpus() {
  fs::path dir = fs::absolute("sample_corpus");
  fs::create_directories(dir);
  std::vector<std::pair<std::string, std::string>> docs{
      {"science_quantum.txt", "Quantum physics studies particles, waves, uncertainty, and entanglement in tiny systems."},
      {"science_astronomy.txt", "Astronomy explores stars, galaxies, black holes, and telescopes that map distant planets."},
      {"science_biology.txt", "Biology examines cells, genes, evolution, and ecosystems in living organisms."},
      {"science_climate.txt", "Climate science tracks greenhouse gases, weather patterns, and long term temperature changes."},
      {"sports_football.txt", "Football strategy includes passing, defense, pressing, and midfield control during competition."},
      {"sports_basketball.txt", "Basketball players practice shooting, dribbling, spacing, and fast breaks to win games."},
      {"sports_running.txt", "Running performance improves with interval training, nutrition, and recovery routines."},
      {"sports_tennis.txt", "Tennis matches require serves, volleys, footwork, and tactical shot placement."},
      {"tech_ai.txt", "Artificial intelligence uses machine learning models, data pipelines, and optimization methods."},
      {"tech_security.txt", "Cybersecurity protects networks with encryption, monitoring, authentication, and incident response."},
      {"tech_cloud.txt", "Cloud computing provides scalable storage, virtual machines, and managed application services."},
      {"tech_web.txt", "Web development combines html css javascript frameworks, testing, and deployment automation."},
      {"cooking_pasta.txt", "Pasta recipes use olive oil, garlic, tomatoes, basil, and careful timing for sauce texture."},
      {"cooking_baking.txt", "Baking bread needs flour, yeast, hydration, proofing, and oven temperature control."},
      {"cooking_spices.txt", "Spice blends balance heat, sweetness, acidity, and aroma in regional cuisine."},
      {"cooking_salad.txt", "Fresh salad preparation focuses on greens, dressing, crunch, and seasonal produce."},
      {"travel_mountains.txt", "Mountain travel involves hiking trails, altitude planning, weather safety, and local guides."},
      {"travel_cities.txt", "City travel highlights museums, transit cards, neighborhoods, and cultural landmarks."},
      {"travel_beaches.txt", "Beach vacations include snorkeling, tides, sun protection, and coastal food markets."},
      {"travel_budget.txt", "Budget travel uses hostels, public transport, off season fares, and itinerary planning."},
  };
  for (const auto& [name, text] : docs) {
    std::ofstream out(dir / name);
    out << text << "\n";
  }
  return dir;
}

static Engine build_from_directory(const fs::path& dir, bool stem) {
  Engine e(stem);
  for (const auto& entry : fs::directory_iterator(dir)) {
    if (!entry.is_regular_file()) continue;
    std::ifstream in(entry.path(), std::ios::binary);
    if (!in) {
      std::cerr << "Warning: unable to read " << entry.path() << "\n";
      continue;
    }
    std::stringstream ss;
    ss << in.rdbuf();
    std::string data = ss.str();
    if (is_binary(data)) {
      std::cerr << "Warning: skipped binary file " << entry.path() << "\n";
      continue;
    }
    if (data.size() > 10 * 1024 * 1024) {
      std::cerr << "Warning: skipped very large file " << entry.path() << "\n";
      continue;
    }
    std::string reason;
    if (!e.add_document(entry.path().filename().string(), entry.path().string(), data, reason)) {
      std::cerr << "Warning: skipped " << entry.path() << " (" << reason << ")\n";
    }
  }
  e.build_index();
  return e;
}

static void print_stats(const json& stats) {
  std::cout << "Index statistics:\n";
  std::cout << "  Total documents: " << stats["total_documents"] << "\n";
  std::cout << "  Total unique terms: " << stats["total_unique_terms"] << "\n";
  std::cout << "  Average document length: " << std::fixed << std::setprecision(2)
            << stats["average_document_length"].get<double>() << " terms\n";
  std::cout << "  Most common terms (top 20):\n";
  for (const auto& item : stats["most_common_terms"]) {
    std::cout << "    - " << item["term"].get<std::string>() << ": " << item["count"] << "\n";
  }
}

static void print_results(const std::string& query, const std::vector<json>& results) {
  std::cout << "\nQuery: " << query << "\n";
  if (results.empty()) {
    std::cout << "  No matching documents.\n";
    return;
  }
  for (size_t i = 0; i < results.size(); ++i) {
    std::cout << "  " << i + 1 << ". " << results[i]["document"].get<std::string>()
              << " | score=" << std::fixed << std::setprecision(6) << results[i]["score"].get<double>() << "\n";
    std::cout << "     " << results[i]["snippet"].get<std::string>() << "\n";
  }
}

int main(int argc, char** argv) {
  try {
    Config cfg = parse_args(argc, argv);
    bool used_sample = false;
    Engine engine(false);

    if (cfg.load_index.has_value()) {
      engine = Engine::load_index(*cfg.load_index);
    } else {
      fs::path dir;
      if (cfg.directory.has_value()) dir = *cfg.directory;
      else {
        dir = create_sample_corpus();
        used_sample = true;
      }
      if (!fs::exists(dir) || !fs::is_directory(dir)) throw std::runtime_error("Directory does not exist: " + dir.string());
      engine = build_from_directory(dir, cfg.stem);
    }

    if (engine.empty()) throw std::runtime_error("No valid text documents were indexed");
    if (cfg.save_index.has_value()) {
      engine.save_index(*cfg.save_index);
      std::cout << "Saved index: " << cfg.save_index->string() << "\n";
    }

    json stats = engine.stats();
    print_stats(stats);

    json payload{
        {"generated_at", "generated"},
        {"query", cfg.query.has_value() ? json(*cfg.query) : json(nullptr)},
        {"top", cfg.top},
        {"boolean_mode", cfg.boolean_mode},
        {"stats", stats},
        {"results", json::array()}
    };

    if (cfg.query.has_value()) {
      auto results = engine.search(*cfg.query, cfg.top, cfg.boolean_mode);
      print_results(*cfg.query, results);
      payload["results"] = results;
    } else if (used_sample) {
      std::vector<std::string> demo{"quantum physics", "pasta recipe", "travel AND budget", "ai AND security NOT malware"};
      json demo_results = json::array();
      for (const auto& q : demo) {
        bool bm = std::regex_search(q, std::regex(R"(\b(AND|OR|NOT)\b)", std::regex::icase));
        auto results = engine.search(q, cfg.top, bm);
        print_results(q, results);
        demo_results.push_back({{"query", q}, {"boolean_mode", bm}, {"items", results}});
      }
      payload["results"] = demo_results;
    } else {
      std::cout << "No query provided. Use --query to search.\n";
    }

    std::ofstream out(cfg.output);
    if (!out) throw std::runtime_error("Failed to write output: " + cfg.output.string());
    out << payload.dump(2) << "\n";
    std::cout << "Saved search results JSON: " << cfg.output.string() << "\n";
    return 0;
  } catch (const std::exception& ex) {
    std::cerr << "Error: " << ex.what() << "\n";
    return 1;
  }
}