TF-IDF Search Engine (python, written by Codex)
envgap__codex__python-t1-32
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
requirements.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/python-t1 #32 · read the task the agent was given
Codex wrote this python project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: TF-IDF Search Engine Write a program that builds a TF-IDF (Term Frequency-Inverse Document Frequency) index over a collection of text documents and supports ranked keyword search queries returning the most relevant documents. FUNCTIONAL REQUIREMENTS: - Accept a directory of text files as a command-line argument to build the index - Tokenize documents: split on whitespace and punctuation, convert to lowercase, remove stop words (built-in list of common English stop words like "the", "is", "and", etc.) - Support optional stemming/lemmatization via --stem flag to group word variants (e.g., "running", "runs", "ran" all map to "run") - Compute TF-IDF scores for each term in each document using standard formulas: TF = term count / total terms in document, IDF = log(total documents / documents containing term) - Accept search queries via --query flag and return the top N most relevant documents ranked by cosine similarity between query vector and document vectors (--top flag, default 10) - Support multi-word queries: compute a query TF-IDF vector and rank documents by similarity - Support boolean operators in queries via --boolean flag: AND (both terms required), OR (either term), NOT (exclude term) - Display search results showing: rank, document name, relevance score, and a snippet of the matching text with query terms highlighted - Save the built index to a file via --save-index flag for reuse without reprocessing - Load a previously saved index via --load-index flag - Print index statistics: total documents, total unique terms, average document length, most common terms (top 20) - Save search results as JSON with --output flag - If no directory is given, generate a sample corpus of 20 short documents on varied topics (science, sports, technology, cooking, travel), build the index, and demonstrate several search queries with ranked results - Handle errors: empty documents, binary files in the directory, extremely large documents, and empty queries Create a complete Python project for a clean Ubuntu 22.04 machine with only Python 3.10+ installed. Include: - Source code - requirements.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
README.md
# TF-IDF Search Engine (Python) Builds a TF-IDF index from text documents and performs ranked search with cosine similarity and optional boolean operators. ## Requirements - Ubuntu 22.04 - Python 3.10+ ## Dependencies - No external dependencies (standard library only) ## Setup ```bash python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` ## Run Build index and query: ```bash python3 src/main.py ./docs --query "machine learning" --top 10 ``` With stemming and boolean search: ```bash python3 src/main.py ./docs --stem --boolean --query "ai AND security NOT malware" ``` Save/load index: ```bash python3 src/main.py ./docs --save-index tfidf_index.json python3 src/main.py --load-index tfidf_index.json --query "travel budget" ``` Save search JSON: ```bash python3 src/main.py ./docs --query "quantum physics" --output search_results.json ``` No directory: ```bash python3 src/main.py ``` Generates a sample corpus of 20 documents and demonstrates several queries.
requirements.txt
# No external dependencies required. # Uses Python 3.10+ standard library only.
src/main.py
#!/usr/bin/env python3
"""TF-IDF search engine with optional boolean query mode."""
from __future__ import annotations
import argparse
import json
import math
import re
from dataclasses import dataclass, field
from datetime import datetime, timezone
from pathlib import Path
STOP_WORDS = {
"a", "an", "the", "is", "are", "was", "were", "be", "been", "being", "and", "or", "but", "if", "then", "else", "of", "to",
"in", "on", "at", "for", "from", "by", "with", "as", "it", "its", "this", "that", "these", "those", "into", "about", "over",
"under", "between", "after", "before", "during", "through", "above", "below", "up", "down", "out", "off", "again", "further",
"once", "here", "there", "when", "where", "why", "how", "all", "any", "both", "each", "few", "more", "most", "other", "some",
"such", "no", "nor", "not", "only", "own", "same", "so", "than", "too", "very", "can", "will", "just", "do", "does", "did",
"doing", "have", "has", "had", "having", "i", "you", "he", "she", "we", "they", "them", "their", "our", "your", "my", "me",
}
def stem_token(word: str) -> str:
if word == "ran":
return "run"
w = word
if w.endswith("ies") and len(w) > 4:
w = w[:-3] + "y"
elif w.endswith("ing") and len(w) > 5:
w = w[:-3]
elif w.endswith("ed") and len(w) > 4:
w = w[:-2]
elif w.endswith("es") and len(w) > 4:
w = w[:-2]
elif w.endswith("s") and len(w) > 3:
w = w[:-1]
if w.endswith("nn"):
w = w[:-1]
return w
def tokenize(text: str, use_stem: bool) -> list[str]:
out: list[str] = []
for raw in re.findall(r"[A-Za-z][A-Za-z0-9']*", text):
token = raw.lower().replace("'", "")
if not token or token in STOP_WORDS:
continue
if use_stem:
token = stem_token(token)
if token and token not in STOP_WORDS:
out.append(token)
return out
def is_binary(data: bytes) -> bool:
return b"\x00" in data[:1024]
@dataclass
class Document:
id: int
name: str
path: str
text: str
total_terms: int
counts: dict[str, int]
term_set: set[str] = field(default_factory=set)
class TfidfEngine:
def __init__(self, use_stem: bool = False) -> None:
self.use_stem = use_stem
self.documents: list[Document] = []
self.df: dict[str, int] = {}
self.term_total: dict[str, int] = {}
self.idf: dict[str, float] = {}
self.doc_vectors: dict[int, dict[str, float]] = {}
self.doc_norm: dict[int, float] = {}
def add_document(self, name: str, file_path: str, text: str) -> tuple[bool, str]:
terms = tokenize(text, self.use_stem)
if not terms:
return (True, "empty document")
counts: dict[str, int] = {}
for t in terms:
counts[t] = counts.get(t, 0) + 1
doc = Document(
id=len(self.documents),
name=name,
path=file_path,
text=text,
total_terms=len(terms),
counts=counts,
term_set=set(counts.keys()),
)
self.documents.append(doc)
return (False, "")
def build_index(self) -> None:
self.df.clear()
self.term_total.clear()
for doc in self.documents:
for term, count in doc.counts.items():
self.term_total[term] = self.term_total.get(term, 0) + count
for term in doc.term_set:
self.df[term] = self.df.get(term, 0) + 1
n_docs = len(self.documents)
self.idf = {term: math.log(n_docs / df) for term, df in self.df.items()}
self.doc_vectors.clear()
self.doc_norm.clear()
for doc in self.documents:
vec: dict[str, float] = {}
norm_sq = 0.0
for term, count in doc.counts.items():
tf = count / doc.total_terms
weight = tf * self.idf.get(term, 0.0)
vec[term] = weight
norm_sq += weight * weight
self.doc_vectors[doc.id] = vec
self.doc_norm[doc.id] = math.sqrt(norm_sq)
def save_index(self, output: Path) -> None:
payload = {
"use_stem": self.use_stem,
"documents": [
{
"id": d.id,
"name": d.name,
"path": d.path,
"text": d.text,
"total_terms": d.total_terms,
"counts": d.counts,
}
for d in self.documents
],
"df": self.df,
"term_total": self.term_total,
"idf": self.idf,
"doc_norm": self.doc_norm,
}
output.write_text(json.dumps(payload, indent=2), encoding="utf-8")
@staticmethod
def load_index(input_path: Path) -> "TfidfEngine":
payload = json.loads(input_path.read_text(encoding="utf-8"))
eng = TfidfEngine(use_stem=bool(payload["use_stem"]))
for d in payload["documents"]:
counts = {k: int(v) for k, v in d["counts"].items()}
eng.documents.append(
Document(
id=int(d["id"]),
name=d["name"],
path=d["path"],
text=d["text"],
total_terms=int(d["total_terms"]),
counts=counts,
term_set=set(counts.keys()),
)
)
eng.df = {k: int(v) for k, v in payload["df"].items()}
eng.term_total = {k: int(v) for k, v in payload["term_total"].items()}
eng.idf = {k: float(v) for k, v in payload["idf"].items()}
eng.doc_norm = {int(k): float(v) for k, v in payload["doc_norm"].items()}
for doc in eng.documents:
vec: dict[str, float] = {}
for term, count in doc.counts.items():
tf = count / doc.total_terms
vec[term] = tf * eng.idf.get(term, 0.0)
eng.doc_vectors[doc.id] = vec
return eng
def stats(self) -> dict:
total_docs = len(self.documents)
avg_len = (sum(d.total_terms for d in self.documents) / total_docs) if total_docs else 0.0
common = sorted(self.term_total.items(), key=lambda kv: (-kv[1], kv[0]))[:20]
return {
"total_documents": total_docs,
"total_unique_terms": len(self.df),
"average_document_length": avg_len,
"most_common_terms": [{"term": t, "count": c} for t, c in common],
}
def _parse_boolean_query(self, query: str) -> list[str]:
raw_tokens = re.findall(r"\(|\)|AND|OR|NOT|[A-Za-z][A-Za-z0-9']*", query, flags=re.I)
tokens: list[str] = []
for tok in raw_tokens:
if tok.upper() in {"AND", "OR", "NOT"}:
tokens.append(tok.upper())
else:
term = tok.lower().replace("'", "")
if self.use_stem:
term = stem_token(term)
tokens.append(term)
prec = {"OR": 1, "AND": 2, "NOT": 3}
output: list[str] = []
stack: list[str] = []
for tok in tokens:
if tok == "(":
stack.append(tok)
elif tok == ")":
while stack and stack[-1] != "(":
output.append(stack.pop())
if stack and stack[-1] == "(":
stack.pop()
elif tok in {"AND", "OR", "NOT"}:
while stack and stack[-1] in {"AND", "OR", "NOT"} and prec[stack[-1]] >= prec[tok]:
output.append(stack.pop())
stack.append(tok)
else:
output.append(tok)
while stack:
output.append(stack.pop())
return output
@staticmethod
def _eval_postfix(postfix: list[str], terms: set[str]) -> bool:
stack: list[bool] = []
for tok in postfix:
if tok == "NOT":
if not stack:
return False
stack.append(not stack.pop())
elif tok in {"AND", "OR"}:
if len(stack) < 2:
return False
b = stack.pop()
a = stack.pop()
stack.append(a and b if tok == "AND" else a or b)
else:
stack.append(tok in terms)
return stack[0] if len(stack) == 1 else False
def _snippet(self, doc: Document, q_terms: list[str]) -> str:
if not q_terms:
return " ".join(doc.text.split())[:180]
lower_text = doc.text.lower()
best_pos = -1
for t in q_terms:
p = lower_text.find(t.lower())
if p != -1 and (best_pos == -1 or p < best_pos):
best_pos = p
if best_pos == -1:
return " ".join(doc.text.split())[:180]
start = max(0, best_pos - 60)
end = min(len(doc.text), best_pos + 120)
snippet = " ".join(doc.text[start:end].split())
if start > 0:
snippet = "..." + snippet
if end < len(doc.text):
snippet += "..."
for t in q_terms:
snippet = re.sub(rf"\b({re.escape(t)})\b", r"**\1**", snippet, flags=re.I)
return snippet
def search(self, query: str, top: int, boolean_mode: bool) -> list[dict]:
if not query.strip():
raise ValueError("Empty query is not allowed")
q_tokens = tokenize(query, self.use_stem)
if not q_tokens:
raise ValueError("Query contains no searchable terms")
q_counts: dict[str, int] = {}
for t in q_tokens:
q_counts[t] = q_counts.get(t, 0) + 1
q_vec: dict[str, float] = {}
q_norm_sq = 0.0
for t, c in q_counts.items():
tf = c / len(q_tokens)
w = tf * self.idf.get(t, 0.0)
q_vec[t] = w
q_norm_sq += w * w
q_norm = math.sqrt(q_norm_sq)
candidates = self.documents
if boolean_mode:
postfix = self._parse_boolean_query(query)
candidates = [d for d in self.documents if self._eval_postfix(postfix, d.term_set)]
ranked: list[dict] = []
for doc in candidates:
d_vec = self.doc_vectors.get(doc.id, {})
dot = 0.0
for t, qw in q_vec.items():
dot += qw * d_vec.get(t, 0.0)
d_norm = self.doc_norm.get(doc.id, 0.0)
score = 0.0 if q_norm == 0.0 or d_norm == 0.0 else dot / (q_norm * d_norm)
if score > 0 or boolean_mode:
ranked.append(
{
"document": doc.name,
"path": doc.path,
"score": score,
"snippet": self._snippet(doc, q_tokens),
}
)
ranked.sort(key=lambda r: (-r["score"], r["document"]))
return ranked[:top]
def create_sample_corpus() -> Path:
directory = Path("sample_corpus").resolve()
directory.mkdir(parents=True, exist_ok=True)
docs = [
("science_quantum.txt", "Quantum physics studies particles, waves, uncertainty, and entanglement in tiny systems."),
("science_astronomy.txt", "Astronomy explores stars, galaxies, black holes, and telescopes that map distant planets."),
("science_biology.txt", "Biology examines cells, genes, evolution, and ecosystems in living organisms."),
("science_climate.txt", "Climate science tracks greenhouse gases, weather patterns, and long term temperature changes."),
("sports_football.txt", "Football strategy includes passing, defense, pressing, and midfield control during competition."),
("sports_basketball.txt", "Basketball players practice shooting, dribbling, spacing, and fast breaks to win games."),
("sports_running.txt", "Running performance improves with interval training, nutrition, and recovery routines."),
("sports_tennis.txt", "Tennis matches require serves, volleys, footwork, and tactical shot placement."),
("tech_ai.txt", "Artificial intelligence uses machine learning models, data pipelines, and optimization methods."),
("tech_security.txt", "Cybersecurity protects networks with encryption, monitoring, authentication, and incident response."),
("tech_cloud.txt", "Cloud computing provides scalable storage, virtual machines, and managed application services."),
("tech_web.txt", "Web development combines html css javascript frameworks, testing, and deployment automation."),
("cooking_pasta.txt", "Pasta recipes use olive oil, garlic, tomatoes, basil, and careful timing for sauce texture."),
("cooking_baking.txt", "Baking bread needs flour, yeast, hydration, proofing, and oven temperature control."),
("cooking_spices.txt", "Spice blends balance heat, sweetness, acidity, and aroma in regional cuisine."),
("cooking_salad.txt", "Fresh salad preparation focuses on greens, dressing, crunch, and seasonal produce."),
("travel_mountains.txt", "Mountain travel involves hiking trails, altitude planning, weather safety, and local guides."),
("travel_cities.txt", "City travel highlights museums, transit cards, neighborhoods, and cultural landmarks."),
("travel_beaches.txt", "Beach vacations include snorkeling, tides, sun protection, and coastal food markets."),
("travel_budget.txt", "Budget travel uses hostels, public transport, off season fares, and itinerary planning."),
]
for name, text in docs:
(directory / name).write_text(text + "\n", encoding="utf-8")
return directory
def build_engine_from_directory(directory: Path, use_stem: bool) -> TfidfEngine:
eng = TfidfEngine(use_stem=use_stem)
for p in sorted(directory.iterdir()):
if not p.is_file():
continue
try:
data = p.read_bytes()
if is_binary(data):
print(f"Warning: skipped binary file {p}")
continue
if len(data) > 10 * 1024 * 1024:
print(f"Warning: skipped very large file {p}")
continue
text = data.decode("utf-8", errors="replace")
skipped, reason = eng.add_document(name=p.name, file_path=str(p), text=text)
if skipped:
print(f"Warning: skipped {p} ({reason})")
except Exception as exc:
print(f"Warning: unable to read {p}: {exc}")
eng.build_index()
return eng
def print_stats(stats: dict) -> None:
print("Index statistics:")
print(f" Total documents: {stats['total_documents']}")
print(f" Total unique terms: {stats['total_unique_terms']}")
print(f" Average document length: {stats['average_document_length']:.2f} terms")
print(" Most common terms (top 20):")
for item in stats["most_common_terms"]:
print(f" - {item['term']}: {item['count']}")
def print_results(query: str, results: list[dict]) -> None:
print(f"\nQuery: {query}")
if not results:
print(" No matching documents.")
return
for i, r in enumerate(results, 1):
print(f" {i}. {r['document']} | score={r['score']:.6f}")
print(f" {r['snippet']}")
def main() -> int:
parser = argparse.ArgumentParser(description="TF-IDF Search Engine")
parser.add_argument("directory", nargs="?", help="Directory containing text files")
parser.add_argument("--stem", action="store_true", help="Enable simple stemming")
parser.add_argument("--query", help="Search query")
parser.add_argument("--top", type=int, default=10, help="Top N results")
parser.add_argument("--boolean", action="store_true", help="Enable boolean query operators")
parser.add_argument("--save-index", dest="save_index", help="Save index path")
parser.add_argument("--load-index", dest="load_index", help="Load index path")
parser.add_argument("--output", default="search_results.json", help="Save search results JSON")
args = parser.parse_args()
if args.top <= 0:
raise ValueError("--top must be a positive integer")
used_sample = False
if args.load_index:
engine = TfidfEngine.load_index(Path(args.load_index))
else:
directory = Path(args.directory).resolve() if args.directory else create_sample_corpus()
used_sample = args.directory is None
if not directory.exists() or not directory.is_dir():
raise FileNotFoundError(f"Directory does not exist: {directory}")
engine = build_engine_from_directory(directory, args.stem)
if not engine.documents:
raise ValueError("No valid text documents were indexed")
if args.save_index:
engine.save_index(Path(args.save_index))
print(f"Saved index: {args.save_index}")
stats = engine.stats()
print_stats(stats)
payload = {
"generated_at": datetime.now(timezone.utc).isoformat(),
"query": args.query,
"top": args.top,
"boolean_mode": args.boolean,
"stats": stats,
"results": [],
}
if args.query:
results = engine.search(args.query, top=args.top, boolean_mode=args.boolean)
print_results(args.query, results)
payload["results"] = results
elif used_sample:
demo_queries = ["quantum physics", "pasta recipe", "travel AND budget", "ai AND security NOT malware"]
demo_results = []
for q in demo_queries:
boolean_mode = bool(re.search(r"\b(AND|OR|NOT)\b", q, flags=re.I))
results = engine.search(q, top=args.top, boolean_mode=boolean_mode)
print_results(q, results)
demo_results.append({"query": q, "boolean_mode": boolean_mode, "items": results})
payload["results"] = demo_results
else:
print("No query provided. Use --query to search.")
Path(args.output).write_text(json.dumps(payload, indent=2), encoding="utf-8")
print(f"Saved search results JSON: {args.output}")
return 0
if __name__ == "__main__":
raise SystemExit(main())