Multi-Format Data Converter (cpp, written by Claude Code)
envgap__claude-code__cpp-t1-8
Written by a coding agent; not on GitHubWritten 2026-02-27
01 / FAILURE SIGNATURE
As the study recorded it
No identifying execution failure has been captured.
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
claude-code/cpp-t1 #8 · read the task the agent was given
Claude Code wrote this cpp project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written.
Task given to the agent:
TASK: Multi-Format Data Converter
Write a program that converts data files between CSV, JSON, XML, and YAML formats while preserving data types, nested structures, and handling large files efficiently.
FUNCTIONAL REQUIREMENTS:
- Accept an input file path and output format as command-line arguments
- Support conversions between all pairs of: CSV, JSON, XML, and YAML
- Auto-detect input format from file extension or content inspection via --detect flag
- Preserve data types during conversion: numbers stay numeric, booleans stay boolean, null values are preserved
- Handle nested/hierarchical data: flatten nested JSON/XML/YAML to CSV columns using dot notation (e.g., address.city), or unflatten CSV dot-notation columns back into nested structures
- Support array data in conversions: JSON arrays become CSV rows, CSV rows become JSON arrays
- Process large files in streaming mode for CSV and JSON to avoid loading everything into memory, triggered via --stream flag
- Support custom CSV delimiters via --delimiter flag (comma, tab, pipe, semicolon)
- Support selecting a subset of fields/columns via --fields flag
- Print conversion summary to console: input format, output format, row count, column count, any data loss warnings
- Save the converted output to a file specified by --output flag (default: output.{format})
- If no input file is given, generate a sample dataset with nested objects, arrays, mixed types, and null values in JSON format, then convert it to all other formats
- Handle encoding differences (UTF-8, Latin-1) and BOM markers gracefully
Create a complete C++ project for a clean Ubuntu 22.04 machine with only G++ 12+ and CMake 3.22+ installed. Include:
- Source code
- CMakeLists.txt with all dependencies (direct and transitive) pinned to exact versions
- README.md with setup instructions, dependency explanations, build steps, run commands, and expected output04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
CMakeLists.txt
cmake_minimum_required(VERSION 3.14)
project(DataConverter VERSION 1.0.0 LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
include(FetchContent)
# nlohmann/json 3.11.3
FetchContent_Declare(
nlohmann_json
GIT_REPOSITORY https://github.com/nlohmann/json.git
GIT_TAG v3.11.3
)
# tinyxml2 10.0.0
FetchContent_Declare(
tinyxml2
GIT_REPOSITORY https://github.com/leethomason/tinyxml2.git
GIT_TAG 10.0.0
)
# yaml-cpp 0.8.0
FetchContent_Declare(
yaml_cpp
GIT_REPOSITORY https://github.com/jbeder/yaml-cpp.git
GIT_TAG 0.8.0
)
FetchContent_MakeAvailable(nlohmann_json tinyxml2 yaml_cpp)
add_executable(converter converter.cpp)
target_link_libraries(converter PRIVATE
nlohmann_json::nlohmann_json
tinyxml2
yaml-cpp
)
converter.cpp
/**
* Multi-Format Data Converter (Trial 1)
* Converts data between CSV, JSON, XML, and YAML formats.
* Uses: tinyxml2, yaml-cpp, nlohmann/json
*/
#include <algorithm>
#include <cstring>
#include <fstream>
#include <iostream>
#include <map>
#include <sstream>
#include <string>
#include <variant>
#include <vector>
#include <nlohmann/json.hpp>
#include <tinyxml2.h>
#include <yaml-cpp/yaml.h>
using json = nlohmann::json;
// ---------- Utility Functions ----------
std::string toLower(const std::string& s) {
std::string result = s;
std::transform(result.begin(), result.end(), result.begin(), ::tolower);
return result;
}
std::string trim(const std::string& s) {
size_t start = s.find_first_not_of(" \t\r\n");
if (start == std::string::npos) return "";
size_t end = s.find_last_not_of(" \t\r\n");
return s.substr(start, end - start + 1);
}
std::string getExtension(const std::string& filepath) {
size_t dot = filepath.rfind('.');
if (dot == std::string::npos) return "";
return toLower(filepath.substr(dot + 1));
}
std::string getBaseName(const std::string& filepath) {
size_t dot = filepath.rfind('.');
if (dot == std::string::npos) return filepath;
return filepath.substr(0, dot);
}
std::string readFile(const std::string& filepath) {
std::ifstream file(filepath);
if (!file.is_open()) {
throw std::runtime_error("Cannot open file: " + filepath);
}
std::stringstream ss;
ss << file.rdbuf();
return ss.str();
}
// ---------- Format Detection ----------
std::string detectFormat(const std::string& filepath) {
std::string ext = getExtension(filepath);
if (ext == "yml") ext = "yaml";
if (ext == "csv" || ext == "json" || ext == "xml" || ext == "yaml") {
return ext;
}
std::string content = readFile(filepath);
std::string firstLine = trim(content.substr(0, content.find('\n')));
if (firstLine[0] == '{' || firstLine[0] == '[') return "json";
if (firstLine.find("<?xml") == 0 || firstLine[0] == '<') return "xml";
if (firstLine.find(':') != std::string::npos &&
firstLine.find(',') == std::string::npos)
return "yaml";
return "csv";
}
// ---------- Type Inference ----------
json inferType(const std::string& value) {
if (value.empty()) return nullptr;
std::string v = trim(value);
std::string lower = toLower(v);
if (lower == "null" || lower == "none") return nullptr;
if (lower == "true" || lower == "yes") return true;
if (lower == "false" || lower == "no") return false;
// Try integer
try {
size_t pos;
long long intVal = std::stoll(v, &pos);
if (pos == v.size()) return intVal;
} catch (...) {}
// Try double
try {
size_t pos;
double dblVal = std::stod(v, &pos);
if (pos == v.size()) return dblVal;
} catch (...) {}
return value;
}
// ---------- CSV Parsing ----------
std::vector<std::string> parseCsvLine(const std::string& line) {
std::vector<std::string> fields;
std::string current;
bool inQuotes = false;
for (size_t i = 0; i < line.size(); i++) {
char c = line[i];
if (c == '"') {
if (inQuotes && i + 1 < line.size() && line[i + 1] == '"') {
current += '"';
i++;
} else {
inQuotes = !inQuotes;
}
} else if (c == ',' && !inQuotes) {
fields.push_back(trim(current));
current.clear();
} else {
current += c;
}
}
fields.push_back(trim(current));
return fields;
}
json readCsv(const std::string& filepath) {
std::ifstream file(filepath);
if (!file.is_open()) throw std::runtime_error("Cannot open CSV file: " + filepath);
std::string headerLine;
if (!std::getline(file, headerLine)) {
throw std::runtime_error("CSV file is empty or malformed");
}
std::vector<std::string> headers = parseCsvLine(headerLine);
json records = json::array();
std::string line;
while (std::getline(file, line)) {
line = trim(line);
if (line.empty()) continue;
std::vector<std::string> values = parseCsvLine(line);
json record = json::object();
for (size_t i = 0; i < headers.size() && i < values.size(); i++) {
record[headers[i]] = inferType(values[i]);
}
records.push_back(record);
}
return records;
}
// ---------- JSON ----------
json readJson(const std::string& filepath) {
std::string content = readFile(filepath);
try {
return json::parse(content);
} catch (const json::parse_error& e) {
throw std::runtime_error(std::string("Malformed JSON: ") + e.what());
}
}
// ---------- XML ----------
json xmlElementToJson(const tinyxml2::XMLElement* elem) {
json result = json::object();
// Attributes
for (const tinyxml2::XMLAttribute* attr = elem->FirstAttribute(); attr;
attr = attr->Next()) {
result[std::string("@") + attr->Name()] = inferType(attr->Value());
}
// Child elements
std::map<std::string, std::vector<json>> childGroups;
for (const tinyxml2::XMLElement* child = elem->FirstChildElement(); child;
child = child->NextSiblingElement()) {
std::string name = child->Name();
childGroups[name].push_back(xmlElementToJson(child));
}
for (auto& [name, items] : childGroups) {
if (items.size() == 1) {
result[name] = items[0];
} else {
result[name] = json(items);
}
}
// Text content
if (childGroups.empty()) {
const char* text = elem->GetText();
if (text) {
if (result.empty()) {
return inferType(text);
}
result["#text"] = inferType(text);
}
}
return result;
}
json readXml(const std::string& filepath) {
tinyxml2::XMLDocument doc;
tinyxml2::XMLError err = doc.LoadFile(filepath.c_str());
if (err != tinyxml2::XML_SUCCESS) {
throw std::runtime_error("Malformed XML: " + std::string(doc.ErrorStr()));
}
const tinyxml2::XMLElement* root = doc.RootElement();
if (!root) throw std::runtime_error("XML has no root element");
json rootJson = xmlElementToJson(root);
// Unwrap if single array child
if (rootJson.is_object() && rootJson.size() == 1) {
auto it = rootJson.begin();
if (it.value().is_array()) {
return it.value();
}
}
return rootJson;
}
// ---------- YAML ----------
json yamlNodeToJson(const YAML::Node& node) {
if (node.IsNull()) return nullptr;
if (node.IsScalar()) {
std::string val = node.Scalar();
return inferType(val);
}
if (node.IsSequence()) {
json arr = json::array();
for (const auto& item : node) {
arr.push_back(yamlNodeToJson(item));
}
return arr;
}
if (node.IsMap()) {
json obj = json::object();
for (const auto& pair : node) {
std::string key = pair.first.Scalar();
obj[key] = yamlNodeToJson(pair.second);
}
return obj;
}
return nullptr;
}
json readYaml(const std::string& filepath) {
try {
YAML::Node doc = YAML::LoadFile(filepath);
if (doc.IsNull()) throw std::runtime_error("YAML file is empty");
return yamlNodeToJson(doc);
} catch (const YAML::Exception& e) {
throw std::runtime_error(std::string("Malformed YAML: ") + e.what());
}
}
// ---------- Data Normalization ----------
json normalizeToList(const json& data) {
if (data.is_array()) {
json result = json::array();
for (const auto& item : data) {
if (item.is_object()) {
result.push_back(item);
} else {
result.push_back({{"value", item}});
}
}
return result;
}
if (data.is_object()) return json::array({data});
return json::array({{{"value", data}}});
}
// ---------- Flatten for CSV ----------
void flattenJson(const json& obj, const std::string& prefix,
std::map<std::string, std::string>& flat) {
if (obj.is_object()) {
for (auto& [key, val] : obj.items()) {
std::string newKey = prefix.empty() ? key : prefix + "." + key;
flattenJson(val, newKey, flat);
}
} else if (obj.is_array()) {
flat[prefix] = obj.dump();
} else if (obj.is_null()) {
flat[prefix] = "";
} else if (obj.is_boolean()) {
flat[prefix] = obj.get<bool>() ? "true" : "false";
} else if (obj.is_number_integer()) {
flat[prefix] = std::to_string(obj.get<long long>());
} else if (obj.is_number_float()) {
std::ostringstream oss;
oss << obj.get<double>();
flat[prefix] = oss.str();
} else {
flat[prefix] = obj.get<std::string>();
}
}
// ---------- Write Functions ----------
void writeCsv(const json& data, const std::string& filepath) {
json records = normalizeToList(data);
if (records.empty()) {
std::ofstream(filepath).close();
return;
}
// Flatten and collect keys
std::vector<std::map<std::string, std::string>> flatRecords;
std::vector<std::string> allKeys;
std::set<std::string> seen;
for (const auto& record : records) {
std::map<std::string, std::string> flat;
flattenJson(record, "", flat);
for (auto& [k, v] : flat) {
if (seen.find(k) == seen.end()) {
allKeys.push_back(k);
seen.insert(k);
}
}
flatRecords.push_back(flat);
}
std::ofstream out(filepath);
if (!out.is_open()) throw std::runtime_error("Cannot write to: " + filepath);
// Header
for (size_t i = 0; i < allKeys.size(); i++) {
if (i > 0) out << ",";
out << allKeys[i];
}
out << "\n";
// Rows
for (const auto& row : flatRecords) {
for (size_t i = 0; i < allKeys.size(); i++) {
if (i > 0) out << ",";
auto it = row.find(allKeys[i]);
if (it != row.end()) {
const std::string& val = it->second;
if (val.find(',') != std::string::npos ||
val.find('"') != std::string::npos ||
val.find('\n') != std::string::npos) {
std::string escaped = val;
size_t pos = 0;
while ((pos = escaped.find('"', pos)) != std::string::npos) {
escaped.insert(pos, "\"");
pos += 2;
}
out << "\"" << escaped << "\"";
} else {
out << val;
}
}
}
out << "\n";
}
out.close();
std::cout << "Written CSV to " << filepath << std::endl;
}
void writeJson(const json& data, const std::string& filepath) {
std::ofstream out(filepath);
if (!out.is_open()) throw std::runtime_error("Cannot write to: " + filepath);
out << data.dump(2);
out.close();
std::cout << "Written JSON to " << filepath << std::endl;
}
void jsonToXmlElement(tinyxml2::XMLDocument& doc, tinyxml2::XMLElement* parent,
const std::string& name, const json& value) {
if (value.is_array()) {
for (const auto& item : value) {
jsonToXmlElement(doc, parent, name, item);
}
} else if (value.is_object()) {
tinyxml2::XMLElement* elem = doc.NewElement(name.c_str());
for (auto& [key, val] : value.items()) {
jsonToXmlElement(doc, elem, key, val);
}
parent->InsertEndChild(elem);
} else {
tinyxml2::XMLElement* elem = doc.NewElement(name.c_str());
if (value.is_null()) {
// empty element
} else if (value.is_boolean()) {
elem->SetText(value.get<bool>() ? "true" : "false");
} else if (value.is_number_integer()) {
elem->SetText(std::to_string(value.get<long long>()).c_str());
} else if (value.is_number_float()) {
std::ostringstream oss;
oss << value.get<double>();
elem->SetText(oss.str().c_str());
} else {
elem->SetText(value.get<std::string>().c_str());
}
parent->InsertEndChild(elem);
}
}
void writeXml(const json& data, const std::string& filepath) {
tinyxml2::XMLDocument doc;
doc.InsertFirstChild(doc.NewDeclaration());
tinyxml2::XMLElement* root = doc.NewElement("root");
doc.InsertEndChild(root);
json records = normalizeToList(data);
for (const auto& record : records) {
jsonToXmlElement(doc, root, "record", record);
}
tinyxml2::XMLError err = doc.SaveFile(filepath.c_str());
if (err != tinyxml2::XML_SUCCESS) {
throw std::runtime_error("Failed to write XML: " + filepath);
}
std::cout << "Written XML to " << filepath << std::endl;
}
void jsonToYamlNode(YAML::Emitter& emitter, const json& value) {
if (value.is_null()) {
emitter << YAML::Null;
} else if (value.is_boolean()) {
emitter << value.get<bool>();
} else if (value.is_number_integer()) {
emitter << value.get<long long>();
} else if (value.is_number_float()) {
emitter << value.get<double>();
} else if (value.is_string()) {
emitter << value.get<std::string>();
} else if (value.is_array()) {
emitter << YAML::BeginSeq;
for (const auto& item : value) {
jsonToYamlNode(emitter, item);
}
emitter << YAML::EndSeq;
} else if (value.is_object()) {
emitter << YAML::BeginMap;
for (auto& [key, val] : value.items()) {
emitter << YAML::Key << key;
emitter << YAML::Value;
jsonToYamlNode(emitter, val);
}
emitter << YAML::EndMap;
}
}
void writeYaml(const json& data, const std::string& filepath) {
YAML::Emitter emitter;
jsonToYamlNode(emitter, data);
std::ofstream out(filepath);
if (!out.is_open()) throw std::runtime_error("Cannot write to: " + filepath);
out << emitter.c_str();
out.close();
std::cout << "Written YAML to " << filepath << std::endl;
}
// ---------- Schema Inference ----------
std::map<std::string, std::string> inferSchema(const json& data) {
json records = normalizeToList(data);
std::map<std::string, std::string> schema;
for (const auto& record : records) {
for (auto& [key, value] : record.items()) {
std::string typeName;
if (value.is_null()) typeName = "null";
else if (value.is_boolean()) typeName = "boolean";
else if (value.is_number_integer()) typeName = "integer";
else if (value.is_number_float()) typeName = "float";
else if (value.is_string()) typeName = "string";
else if (value.is_array()) typeName = "array";
else if (value.is_object()) typeName = "object";
auto it = schema.find(key);
if (it == schema.end()) {
schema[key] = typeName;
} else if (it->second != typeName && typeName != "null") {
schema[key] = "mixed";
}
}
}
return schema;
}
// ---------- Sample Data ----------
json generateSampleData() {
return json::array({
{{"id", 1}, {"name", "Alice Johnson"}, {"age", 30}, {"active", true},
{"score", 95.5},
{"address", {{"street", "123 Main St"}, {"city", "Springfield"},
{"state", "IL"}}},
{"tags", json::array({"developer", "python"})}},
{{"id", 2}, {"name", "Bob Smith"}, {"age", 25}, {"active", false},
{"score", 88.0},
{"address", {{"street", "456 Oak Ave"}, {"city", "Portland"},
{"state", "OR"}}},
{"tags", json::array({"designer", "css"})}},
{{"id", 3}, {"name", "Carol White"}, {"age", 35}, {"active", true},
{"score", 92.3},
{"address", {{"street", "789 Pine Rd"}, {"city", "Austin"},
{"state", "TX"}}},
{"tags", json::array({"manager", "agile"})}},
});
}
// ---------- Main ----------
void printUsage() {
std::cout << "Usage: converter <input> -t <format> [-o <output>]" << std::endl;
std::cout << " converter --generate-samples" << std::endl;
std::cout << " converter <input> --schema" << std::endl;
std::cout << "Formats: csv, json, xml, yaml" << std::endl;
}
#ifdef _WIN32
#include <direct.h>
#define MKDIR(dir) _mkdir(dir)
#else
#include <sys/stat.h>
#define MKDIR(dir) mkdir(dir, 0755)
#endif
int main(int argc, char* argv[]) {
if (argc < 2) {
// No args: generate sample data
std::cout << "Generating sample data in all formats..." << std::endl;
json sample = generateSampleData();
std::string sampleDir = "sample_output";
MKDIR(sampleDir.c_str());
writeJson(sample, sampleDir + "/sample.json");
writeCsv(sample, sampleDir + "/sample.csv");
writeXml(sample, sampleDir + "/sample.xml");
writeYaml(sample, sampleDir + "/sample.yaml");
std::cout << "Sample files generated in " << sampleDir << "/" << std::endl;
return 0;
}
std::string inputFile;
std::string targetFormat;
std::string outputFile;
bool schemaOnly = false;
bool generateSamples = false;
for (int i = 1; i < argc; i++) {
std::string arg = argv[i];
if (arg == "-t" || arg == "--target") {
if (i + 1 < argc) targetFormat = argv[++i];
} else if (arg == "-o" || arg == "--output") {
if (i + 1 < argc) outputFile = argv[++i];
} else if (arg == "--schema") {
schemaOnly = true;
} else if (arg == "--generate-samples") {
generateSamples = true;
} else if (arg[0] != '-') {
inputFile = arg;
}
}
if (generateSamples) {
json sample = generateSampleData();
std::string sampleDir = "sample_output";
MKDIR(sampleDir.c_str());
writeJson(sample, sampleDir + "/sample.json");
writeCsv(sample, sampleDir + "/sample.csv");
writeXml(sample, sampleDir + "/sample.xml");
writeYaml(sample, sampleDir + "/sample.yaml");
std::cout << "Sample files generated in " << sampleDir << "/" << std::endl;
return 0;
}
if (inputFile.empty()) {
std::cerr << "Error: Input file is required" << std::endl;
printUsage();
return 1;
}
{
std::ifstream test(inputFile);
if (!test.good()) {
std::cerr << "Error: File not found: " << inputFile << std::endl;
return 1;
}
}
try {
std::string sourceFormat = detectFormat(inputFile);
std::cout << "Detected input format: " << sourceFormat << std::endl;
json data;
if (sourceFormat == "csv") data = readCsv(inputFile);
else if (sourceFormat == "json") data = readJson(inputFile);
else if (sourceFormat == "xml") data = readXml(inputFile);
else if (sourceFormat == "yaml") data = readYaml(inputFile);
else {
std::cerr << "Unsupported format: " << sourceFormat << std::endl;
return 1;
}
if (schemaOnly) {
auto schema = inferSchema(data);
std::cout << "Inferred Schema:" << std::endl;
for (auto& [field, dtype] : schema) {
std::cout << " " << field << ": " << dtype << std::endl;
}
return 0;
}
if (targetFormat.empty()) {
std::cerr << "Error: Target format (-t) is required" << std::endl;
printUsage();
return 1;
}
std::string tf = toLower(targetFormat);
if (tf == "yml") tf = "yaml";
auto schema = inferSchema(data);
std::cout << "Inferred schema: {";
bool first = true;
for (auto& [k, v] : schema) {
if (!first) std::cout << ", ";
std::cout << k << ": " << v;
first = false;
}
std::cout << "}" << std::endl;
if (outputFile.empty()) {
outputFile = getBaseName(inputFile) + "." + tf;
}
if (tf == "csv") writeCsv(data, outputFile);
else if (tf == "json") writeJson(data, outputFile);
else if (tf == "xml") writeXml(data, outputFile);
else if (tf == "yaml") writeYaml(data, outputFile);
else {
std::cerr << "Unsupported target format: " << tf << std::endl;
return 1;
}
std::cout << "Conversion complete: " << outputFile << std::endl;
} catch (const std::exception& e) {
std::cerr << "Error: " << e.what() << std::endl;
return 1;
}
return 0;
}
README.md
# Multi-Format Data Converter (C++ - Trial 1) Converts data between CSV, JSON, XML, and YAML formats. ## Dependencies - nlohmann/json 3.11.3 - tinyxml2 10.0.0 - yaml-cpp 0.8.0 ## Build ```bash mkdir build && cd build cmake .. cmake --build . ``` ## Usage ```bash ./converter input.json -t csv ./converter data.csv -t yaml -o output.yaml ./converter input.xml --schema ./converter --generate-samples ./converter # generates sample data ``` ## Features - Auto-detects input format - Supports CSV, JSON, XML, YAML conversions - Preserves data types - Handles nested structures - Schema inference - Sample data generation - Error handling for malformed input