Text Diff Tool (cpp, written by Codex)
envgap__codex__cpp-t1-33
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/cpp-t1 #33 · read the task the agent was given
Codex wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: Text Diff Tool Write a program that compares two text files and produces a detailed diff showing additions, deletions, and modifications with configurable output formats and context control. FUNCTIONAL REQUIREMENTS: - Accept two file paths as command-line arguments (original and modified) - Compute the longest common subsequence (LCS) based diff to identify added, deleted, and changed lines - Support multiple output formats via --format flag: unified diff (default, similar to git diff), side-by-side (two-column view), inline (changes marked within lines), and html (visual diff as an HTML page) - Support configurable context lines around changes via --context flag (default: 3 lines of unchanged context around each change) - Detect and highlight intra-line changes: when a line is modified, show exactly which words or characters changed within the line - Support ignoring whitespace differences via --ignore-whitespace flag - Support ignoring case differences via --ignore-case flag - Support ignoring blank lines via --ignore-blank-lines flag - Compute and display diff statistics: total lines in each file, lines added, lines deleted, lines modified, and a similarity percentage - Support comparing directories via --recursive flag: compare all matching files in two directory trees and report which files are added, deleted, modified, or identical - Apply color coding in console output: green for additions, red for deletions, yellow for modifications - Save the diff output to a file via --output flag - If no input files are given, generate two sample text files (original and modified version with insertions, deletions, modifications, and moved blocks), then compute and display the diff in all supported formats - Handle errors: binary files (detect and skip with warning), missing files, encoding mismatches, and very large files Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include: - Source code - CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.22) project(text_diff_tool LANGUAGES CXX) set(CMAKE_CXX_STANDARD 20) set(CMAKE_CXX_STANDARD_REQUIRED ON) set(CMAKE_CXX_EXTENSIONS OFF) include(FetchContent) FetchContent_Declare( nlohmann_json GIT_REPOSITORY https://github.com/nlohmann/json.git GIT_TAG v3.11.3 ) FetchContent_MakeAvailable(nlohmann_json) add_executable(text_diff_tool src/main.cpp) target_link_libraries(text_diff_tool PRIVATE nlohmann_json::nlohmann_json)
README.md
# Text Diff Tool (C++) Compares two text files using an LCS-based diff with unified, side-by-side, inline, and HTML outputs. ## Requirements - Ubuntu 22.04 - G++ 12+ - CMake 3.22+ ## Dependencies (Pinned) - `nlohmann/json` `v3.11.3` for JSON report output ## Build ```bash cmake -S . -B build cmake --build build -j ``` ## Run Unified diff: ```bash ./build/text_diff_tool original.txt modified.txt --format unified --context 3 ``` Inline diff with ignore options: ```bash ./build/text_diff_tool original.txt modified.txt --format inline --ignore-whitespace --ignore-case ``` HTML diff: ```bash ./build/text_diff_tool original.txt modified.txt --format html --output diff.html ``` Recursive directory comparison: ```bash ./build/text_diff_tool dir_a dir_b --recursive ``` No input files: ```bash ./build/text_diff_tool ``` Generates sample files and displays all supported formats.
src/main.cpp
#include <algorithm>
#include <cctype>
#include <filesystem>
#include <fstream>
#include <iomanip>
#include <iostream>
#include <map>
#include <optional>
#include <regex>
#include <set>
#include <sstream>
#include <stdexcept>
#include <string>
#include <unordered_map>
#include <unordered_set>
#include <utility>
#include <vector>
#include <nlohmann/json.hpp>
namespace fs = std::filesystem;
using json = nlohmann::json;
static constexpr const char* RED = "\x1b[31m";
static constexpr const char* GREEN = "\x1b[32m";
static constexpr const char* YELLOW = "\x1b[33m";
static constexpr const char* RESET = "\x1b[0m";
struct Config {
std::string format = "unified";
int context = 3;
bool ignore_whitespace = false;
bool ignore_case = false;
bool ignore_blank_lines = false;
bool recursive = false;
std::optional<std::string> output;
std::vector<std::string> inputs;
};
struct ReadResult {
bool binary = false;
std::vector<std::string> lines;
std::vector<std::string> normalized;
};
static std::string trim(const std::string& s) {
size_t b = 0;
while (b < s.size() && std::isspace(static_cast<unsigned char>(s[b]))) ++b;
size_t e = s.size();
while (e > b && std::isspace(static_cast<unsigned char>(s[e - 1]))) --e;
return s.substr(b, e - b);
}
static std::string lower(std::string s) {
for (char& c : s) c = static_cast<char>(std::tolower(static_cast<unsigned char>(c)));
return s;
}
static std::string colorize(const std::string& text, const char* color) {
return std::string(color) + text + RESET;
}
static Config parse_args(int argc, char** argv) {
Config cfg;
for (int i = 1; i < argc; ++i) {
std::string arg = argv[i];
if (!arg.starts_with("--")) {
cfg.inputs.push_back(arg);
continue;
}
if (arg == "--ignore-whitespace") cfg.ignore_whitespace = true;
else if (arg == "--ignore-case") cfg.ignore_case = true;
else if (arg == "--ignore-blank-lines") cfg.ignore_blank_lines = true;
else if (arg == "--recursive") cfg.recursive = true;
else if (arg == "--format") {
if (i + 1 >= argc) throw std::runtime_error("Missing value for --format");
cfg.format = argv[++i];
} else if (arg == "--context") {
if (i + 1 >= argc) throw std::runtime_error("Missing value for --context");
cfg.context = std::stoi(argv[++i]);
} else if (arg == "--output") {
if (i + 1 >= argc) throw std::runtime_error("Missing value for --output");
cfg.output = std::string(argv[++i]);
} else {
throw std::runtime_error("Unknown option: " + arg);
}
}
if (cfg.context < 0) throw std::runtime_error("--context must be non-negative");
if (cfg.format != "unified" && cfg.format != "side-by-side" && cfg.format != "inline" && cfg.format != "html") {
throw std::runtime_error("--format must be one of: unified, side-by-side, inline, html");
}
return cfg;
}
static std::string maybe_normalize(std::string line, const Config& cfg) {
if (cfg.ignore_whitespace) {
line = std::regex_replace(line, std::regex(R"(\s+)"), " ");
line = trim(line);
}
if (cfg.ignore_case) line = lower(line);
return line;
}
static bool is_binary(const std::string& bytes) {
size_t n = std::min<size_t>(bytes.size(), 1024);
for (size_t i = 0; i < n; ++i) if (bytes[i] == '\0') return true;
return false;
}
static ReadResult read_lines(const fs::path& p, const Config& cfg) {
std::ifstream in(p, std::ios::binary);
if (!in) throw std::runtime_error("Unreadable file: " + p.string());
std::stringstream ss;
ss << in.rdbuf();
std::string data = ss.str();
if (is_binary(data)) return {.binary = true};
if (data.size() > 20 * 1024 * 1024) throw std::runtime_error("Very large file skipped: " + p.string());
data = std::regex_replace(data, std::regex("\r\n"), "\n");
std::vector<std::string> lines;
std::stringstream ls(data);
std::string line;
while (std::getline(ls, line)) lines.push_back(line);
if (cfg.ignore_blank_lines) {
lines.erase(std::remove_if(lines.begin(), lines.end(), [](const std::string& l) { return trim(l).empty(); }), lines.end());
}
std::vector<std::string> normalized;
normalized.reserve(lines.size());
for (const auto& l : lines) normalized.push_back(maybe_normalize(l, cfg));
return {.binary = false, .lines = lines, .normalized = normalized};
}
struct Op {
std::string type; // equal, add, del, mod
int ai = -1;
int bj = -1;
};
static std::vector<Op> lcs_diff(const std::vector<std::string>& a, const std::vector<std::string>& b) {
int n = static_cast<int>(a.size());
int m = static_cast<int>(b.size());
std::vector<std::vector<int>> dp(n + 1, std::vector<int>(m + 1, 0));
for (int i = n - 1; i >= 0; --i) {
for (int j = m - 1; j >= 0; --j) {
dp[i][j] = a[i] == b[j] ? dp[i + 1][j + 1] + 1 : std::max(dp[i + 1][j], dp[i][j + 1]);
}
}
std::vector<Op> ops;
int i = 0, j = 0;
while (i < n && j < m) {
if (a[i] == b[j]) { ops.push_back({"equal", i, j}); ++i; ++j; }
else if (dp[i + 1][j] >= dp[i][j + 1]) { ops.push_back({"del", i, -1}); ++i; }
else { ops.push_back({"add", -1, j}); ++j; }
}
while (i < n) { ops.push_back({"del", i, -1}); ++i; }
while (j < m) { ops.push_back({"add", -1, j}); ++j; }
return ops;
}
static std::vector<Op> pair_mods(const std::vector<Op>& ops) {
std::vector<Op> out;
size_t i = 0;
while (i < ops.size()) {
if (ops[i].type != "del") {
out.push_back(ops[i]);
++i;
continue;
}
size_t di = i;
while (di < ops.size() && ops[di].type == "del") ++di;
size_t ai = di;
while (ai < ops.size() && ops[ai].type == "add") ++ai;
if (di < ai) {
auto dels_begin = ops.begin() + static_cast<long>(i);
auto dels_end = ops.begin() + static_cast<long>(di);
auto adds_begin = ops.begin() + static_cast<long>(di);
auto adds_end = ops.begin() + static_cast<long>(ai);
std::vector<Op> dels(dels_begin, dels_end);
std::vector<Op> adds(adds_begin, adds_end);
size_t k = std::min(dels.size(), adds.size());
for (size_t t = 0; t < k; ++t) out.push_back({"mod", dels[t].ai, adds[t].bj});
for (size_t t = k; t < dels.size(); ++t) out.push_back(dels[t]);
for (size_t t = k; t < adds.size(); ++t) out.push_back(adds[t]);
i = ai;
} else {
out.push_back(ops[i]);
++i;
}
}
return out;
}
static std::vector<std::string> tokenize_inline(const std::string& line) {
std::vector<std::string> out;
std::regex rx(R"((\s+|[^\w\s]+|\w+))");
for (std::sregex_iterator it(line.begin(), line.end(), rx), end; it != end; ++it) out.push_back(it->str());
return out;
}
static std::pair<std::string, std::string> inline_word_diff(const std::string& old_line, const std::string& new_line) {
auto a = tokenize_inline(old_line);
auto b = tokenize_inline(new_line);
auto ops = lcs_diff(a, b);
std::string oa;
std::string ob;
for (const auto& op : ops) {
if (op.type == "equal") { oa += a[static_cast<size_t>(op.ai)]; ob += b[static_cast<size_t>(op.bj)]; }
else if (op.type == "del") oa += "[-" + a[static_cast<size_t>(op.ai)] + "-]";
else if (op.type == "add") ob += "[+" + b[static_cast<size_t>(op.bj)] + "+]";
}
return {oa, ob};
}
static json diff_stats(const std::vector<Op>& ops, int a_len, int b_len) {
int added = 0, deleted = 0, modified = 0, equal = 0;
for (const auto& op : ops) {
if (op.type == "add") added++;
else if (op.type == "del") deleted++;
else if (op.type == "mod") modified++;
else equal++;
}
double similarity = std::max(a_len, b_len) == 0 ? 100.0 : (equal * 100.0 / std::max(a_len, b_len));
return {
{"totalOriginal", a_len},
{"totalModified", b_len},
{"added", added},
{"deleted", deleted},
{"modified", modified},
{"similarityPct", similarity}
};
}
static std::string render_unified(const std::vector<Op>& ops, const std::vector<std::string>& left,
const std::vector<std::string>& right, int context) {
std::vector<int> changed;
for (size_t i = 0; i < ops.size(); ++i) if (ops[i].type != "equal") changed.push_back(static_cast<int>(i));
if (changed.empty()) return "No differences.\n";
std::set<int> keep;
for (int idx : changed) {
for (int k = std::max(0, idx - context); k <= std::min(static_cast<int>(ops.size()) - 1, idx + context); ++k) keep.insert(k);
}
std::ostringstream out;
out << "--- original\n+++ modified\n";
bool skipping = false;
for (size_t i = 0; i < ops.size(); ++i) {
if (!keep.contains(static_cast<int>(i))) {
if (!skipping) out << "@@ ... @@\n";
skipping = true;
continue;
}
skipping = false;
const auto& op = ops[i];
if (op.type == "equal") out << " " << left[static_cast<size_t>(op.ai)] << "\n";
else if (op.type == "add") out << colorize("+" + right[static_cast<size_t>(op.bj)] + "\n", GREEN);
else if (op.type == "del") out << colorize("-" + left[static_cast<size_t>(op.ai)] + "\n", RED);
else {
auto [oa, ob] = inline_word_diff(left[static_cast<size_t>(op.ai)], right[static_cast<size_t>(op.bj)]);
out << colorize("~" + oa + "\n", YELLOW);
out << colorize("~" + ob + "\n", YELLOW);
}
}
return out.str();
}
static std::string render_side_by_side(const std::vector<Op>& ops, const std::vector<std::string>& left, const std::vector<std::string>& right) {
std::ostringstream out;
const int width = 60;
for (const auto& op : ops) {
std::string l = op.ai < 0 ? "" : left[static_cast<size_t>(op.ai)];
std::string r = op.bj < 0 ? "" : right[static_cast<size_t>(op.bj)];
if (op.type == "add") out << colorize((std::ostringstream{} << std::left << std::setw(width) << "" << " | + " << r << "\n").str(), GREEN);
else if (op.type == "del") out << colorize((std::ostringstream{} << std::left << std::setw(width) << l << " | - " << "\n").str(), RED);
else if (op.type == "mod") {
auto [oa, ob] = inline_word_diff(l, r);
out << colorize((std::ostringstream{} << std::left << std::setw(width) << oa << " | ~ " << ob << "\n").str(), YELLOW);
} else out << (std::ostringstream{} << std::left << std::setw(width) << l << " | " << r << "\n").str();
}
return out.str();
}
static std::string render_inline(const std::vector<Op>& ops, const std::vector<std::string>& left, const std::vector<std::string>& right) {
std::ostringstream out;
for (const auto& op : ops) {
if (op.type == "equal") out << " " << left[static_cast<size_t>(op.ai)] << "\n";
else if (op.type == "add") out << colorize("+ " + right[static_cast<size_t>(op.bj)] + "\n", GREEN);
else if (op.type == "del") out << colorize("- " + left[static_cast<size_t>(op.ai)] + "\n", RED);
else {
auto [oa, ob] = inline_word_diff(left[static_cast<size_t>(op.ai)], right[static_cast<size_t>(op.bj)]);
out << colorize("~ " + oa + "\n", YELLOW);
out << colorize("~ " + ob + "\n", YELLOW);
}
}
return out.str();
}
static std::string esc(const std::string& s) {
std::string out = s;
out = std::regex_replace(out, std::regex("&"), "&");
out = std::regex_replace(out, std::regex("<"), "<");
out = std::regex_replace(out, std::regex(">"), ">");
out = std::regex_replace(out, std::regex("\""), """);
out = std::regex_replace(out, std::regex("'"), "'");
return out;
}
static std::string render_html(const std::vector<Op>& ops, const std::vector<std::string>& left, const std::vector<std::string>& right) {
std::ostringstream rows;
for (const auto& op : ops) {
if (op.type == "equal") rows << "<tr class='eq'><td>" << esc(left[static_cast<size_t>(op.ai)]) << "</td><td>" << esc(right[static_cast<size_t>(op.bj)]) << "</td></tr>";
else if (op.type == "add") rows << "<tr class='add'><td></td><td>" << esc(right[static_cast<size_t>(op.bj)]) << "</td></tr>";
else if (op.type == "del") rows << "<tr class='del'><td>" << esc(left[static_cast<size_t>(op.ai)]) << "</td><td></td></tr>";
else {
auto [oa, ob] = inline_word_diff(left[static_cast<size_t>(op.ai)], right[static_cast<size_t>(op.bj)]);
rows << "<tr class='mod'><td>" << esc(oa) << "</td><td>" << esc(ob) << "</td></tr>";
}
}
return "<!doctype html><html><head><meta charset='utf-8'><title>Diff</title>"
"<style>body{font-family:Arial,sans-serif;margin:1rem}table{width:100%;border-collapse:collapse}"
"td{border:1px solid #ddd;padding:.4rem;vertical-align:top;white-space:pre-wrap;font-family:Consolas,monospace}"
".add td{background:#e8ffe8}.del td{background:#ffe8e8}.mod td{background:#fff7db}</style></head>"
"<body><h1>Diff</h1><table>" + rows.str() + "</table></body></html>\n";
}
static json diff_two_files(const fs::path& left_path, const fs::path& right_path, const Config& cfg) {
auto left = read_lines(left_path, cfg);
auto right = read_lines(right_path, cfg);
if (left.binary || right.binary) {
return {{"warning", "Binary file detected; diff skipped"}, {"output", ""}, {"stats", diff_stats({}, 0, 0)}, {"ops", json::array()}};
}
auto ops = pair_mods(lcs_diff(left.normalized, right.normalized));
json stats = diff_stats(ops, static_cast<int>(left.lines.size()), static_cast<int>(right.lines.size()));
std::string output;
if (cfg.format == "unified") output = render_unified(ops, left.lines, right.lines, cfg.context);
else if (cfg.format == "side-by-side") output = render_side_by_side(ops, left.lines, right.lines);
else if (cfg.format == "inline") output = render_inline(ops, left.lines, right.lines);
else output = render_html(ops, left.lines, right.lines);
json jops = json::array();
for (const auto& op : ops) jops.push_back({{"type", op.type}, {"ai", op.ai < 0 ? json(nullptr) : json(op.ai)}, {"bj", op.bj < 0 ? json(nullptr) : json(op.bj)}});
return {{"warning", nullptr}, {"output", output}, {"stats", stats}, {"ops", jops}};
}
static std::map<std::string, fs::path> walk_files(const fs::path& root) {
std::map<std::string, fs::path> out;
for (const auto& e : fs::recursive_directory_iterator(root)) {
if (e.is_regular_file()) out[e.path().lexically_relative(root).string()] = e.path();
}
return out;
}
static json diff_directories(const fs::path& a, const fs::path& b, const Config& cfg) {
auto files_a = walk_files(a);
auto files_b = walk_files(b);
std::set<std::string> all;
for (const auto& kv : files_a) all.insert(kv.first);
for (const auto& kv : files_b) all.insert(kv.first);
std::vector<std::string> added, deleted, modified, identical;
for (const auto& rel : all) {
bool in_a = files_a.contains(rel);
bool in_b = files_b.contains(rel);
if (!in_a) { added.push_back(rel); continue; }
if (!in_b) { deleted.push_back(rel); continue; }
auto res = diff_two_files(files_a[rel], files_b[rel], cfg);
bool changed = !res["warning"].is_null();
if (!changed) {
for (const auto& op : res["ops"]) {
if (op["type"].get<std::string>() != "equal") { changed = true; break; }
}
}
if (changed) modified.push_back(rel); else identical.push_back(rel);
}
std::ostringstream out;
out << "Directory diff summary\n";
out << "Added: " << added.size() << "\n";
for (const auto& p : added) out << colorize(" + " + p + "\n", GREEN);
out << "Deleted: " << deleted.size() << "\n";
for (const auto& p : deleted) out << colorize(" - " + p + "\n", RED);
out << "Modified: " << modified.size() << "\n";
for (const auto& p : modified) out << colorize(" ~ " + p + "\n", YELLOW);
out << "Identical: " << identical.size() << "\n";
for (const auto& p : identical) out << " " << p << "\n";
return {{"summary", {{"added", added}, {"deleted", deleted}, {"modified", modified}, {"identical", identical}}}, {"output", out.str()}};
}
static std::pair<fs::path, fs::path> create_samples() {
fs::path left = fs::absolute("sample_original.txt");
fs::path right = fs::absolute("sample_modified.txt");
std::ofstream l(left), r(right);
l << "Project Delta Status Report\n"
"The team completed phase one on Monday.\n"
"We tested the API and database integration.\n"
"Performance baseline is 220 requests per second.\n"
"Risks include deployment timing and data migration.\n"
"Action: finalize rollback plan.\n";
r << "Project Delta Status Report\n"
"The team completed phase one on Tuesday.\n"
"We tested API integration and caching layer.\n"
"Performance baseline is 260 requests per second.\n"
"Action: finalize rollback plan and run rehearsal.\n"
"New note: monitor latency in production.\n";
return {left, right};
}
int main(int argc, char** argv) {
try {
Config cfg = parse_args(argc, argv);
fs::path left, right;
bool sample_mode = false;
if (cfg.inputs.size() >= 2) {
left = fs::absolute(cfg.inputs[0]);
right = fs::absolute(cfg.inputs[1]);
} else {
auto [l, r] = create_samples();
left = l;
right = r;
sample_mode = true;
}
if (cfg.recursive) {
if (!fs::is_directory(left) || !fs::is_directory(right)) throw std::runtime_error("For --recursive both inputs must be directories.");
auto res = diff_directories(left, right, cfg);
std::cout << res["output"].get<std::string>();
if (cfg.output.has_value()) {
std::ofstream o(*cfg.output);
o << res["output"].get<std::string>();
}
return 0;
}
if (!fs::exists(left) || !fs::exists(right)) throw std::runtime_error("Input files not found.");
if (fs::is_directory(left) || fs::is_directory(right)) throw std::runtime_error("Use --recursive for directories.");
std::vector<std::string> formats = sample_mode ? std::vector<std::string>{"unified", "side-by-side", "inline", "html"} : std::vector<std::string>{cfg.format};
json format_reports = json::array();
std::ostringstream write_content;
for (const auto& fmt : formats) {
Config run = cfg;
run.format = fmt;
auto res = diff_two_files(left, right, run);
std::string block = sample_mode ? "\n===== FORMAT: " + fmt + " =====\n" + res["output"].get<std::string>() : res["output"].get<std::string>();
if (!res["warning"].is_null()) std::cout << "Warning: " << res["warning"].get<std::string>() << "\n";
std::cout << block;
if (!sample_mode) {
auto s = res["stats"];
std::cout << "Stats: original=" << s["totalOriginal"] << ", modified=" << s["totalModified"]
<< ", +" << s["added"] << ", -" << s["deleted"] << ", ~" << s["modified"]
<< ", similarity=" << std::fixed << std::setprecision(2) << s["similarityPct"].get<double>() << "%\n";
}
write_content << block << "\n";
format_reports.push_back({{"format", fmt}, {"warning", res["warning"]}, {"stats", res["stats"]}});
}
if (cfg.output.has_value()) {
std::ofstream out(*cfg.output);
out << write_content.str();
}
json report{
{"generatedAt", "generated"},
{"original", left.string()},
{"modified", right.string()},
{"recursive", false},
{"formats", format_reports}
};
std::ofstream rep("diff_report.json");
rep << report.dump(2) << "\n";
return 0;
} catch (const std::exception& ex) {
std::cerr << "Error: " << ex.what() << "\n";
return 1;
}
}