File Deduplicator (java, written by Codex)
envgap__codex__java-t1-39
Written by a coding agent; not on GitHubWritten 2026-03-03
01 / FAILURE SIGNATURE
As the study recorded it
None
Not a benchmark task.
- The project already builds and runs before the fix, so there is nothing to repair.
02 / ENVIRONMENT RECIPE
- Base commit
Not freshly verified- Manifest
pom.xml- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / TASK AND FAILURE
codex/java-t1 #39 · read the task the agent was given
Codex wrote this java project from the task below. It installed and ran on a clean Ubuntu 22.04 machine as written. Task given to the agent: TASK: File Deduplicator Write a program that finds and manages duplicate files across directories using content-based hashing, supporting multiple deduplication strategies and detailed reporting. FUNCTIONAL REQUIREMENTS: - Accept one or more directory paths as command-line arguments - Find duplicate files by comparing SHA-256 content hashes, using a two-phase approach: first compare file sizes to narrow candidates, then hash only size-matched files - Support configurable minimum file size via --min-size flag (default: 1 byte) to skip tiny files - Support file type filtering via --include and --exclude flags with glob patterns - Group duplicates into sets showing all copies with their full paths, sizes, and modification dates - Support multiple deduplication actions via --action flag: report (default, just list duplicates), delete (remove duplicates keeping the oldest/newest based on --keep flag), hardlink (replace duplicates with hard links to save space), symlink (replace with symbolic links) - Support a --dry-run flag to preview what would be done without actually modifying files - Scan directories recursively by default, with --no-recursive flag to disable - Display a progress bar during scanning showing files processed and duplicates found so far - Print summary to console: total files scanned, total unique files, duplicate sets found, total wasted space, space that would be recovered - Save the full deduplication report as JSON with --output flag (default: dedup_report.json) - If no directories are given, create a sample directory with intentional duplicates (exact copies, files with same content but different names, and unique files), run deduplication analysis, and display the results - Handle errors: permission denied, broken symlinks, files modified during scan, and cross-filesystem hard links Create a complete Java project for a clean Ubuntu 22.04 machine with only JDK 17+ installed. Include: - Source code - pom.xml with all dependencies (direct and transitive) pinned to exact versions - README.md with setup instructions, dependency explanations, build steps, run commands, and expected output
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]
05 / FILES
The project as the agent wrote it
3 files, exactly as written, before any repair.
pom.xml
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>org.tmlr</groupId>
<artifactId>file-deduplicator</artifactId>
<version>1.0.0</version>
<properties>
<maven.compiler.source>17</maven.compiler.source>
<maven.compiler.target>17</maven.compiler.target>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
</properties>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-assembly-plugin</artifactId>
<version>3.7.1</version>
<configuration>
<archive>
<manifest>
<mainClass>FileDeduplicator</mainClass>
</manifest>
</archive>
<descriptorRefs>
<descriptorRef>jar-with-dependencies</descriptorRef>
</descriptorRefs>
</configuration>
</plugin>
</plugins>
</build>
</project>
README.md
# File Deduplicator (Java) ## Requirements - Ubuntu 22.04 - JDK 17+ - Maven 3.8+ ## Build ```bash mvn -q -DskipTests package assembly:single ``` ## Run ```bash java -cp target/file-deduplicator-1.0.0-jar-with-dependencies.jar FileDeduplicator dir1 dir2 java -cp target/file-deduplicator-1.0.0-jar-with-dependencies.jar FileDeduplicator dir1 --action delete --keep oldest --dry-run java -cp target/file-deduplicator-1.0.0-jar-with-dependencies.jar FileDeduplicator dir1 --action hardlink --keep newest --include "*.txt" ``` Writes JSON report to `dedup_report.json` by default.
src/main/java/FileDeduplicator.java
import java.io.IOException;
import java.nio.file.*;
import java.nio.file.attribute.BasicFileAttributes;
import java.security.MessageDigest;
import java.time.Instant;
import java.util.*;
public class FileDeduplicator {
private static class Config {
List<String> dirs = new ArrayList<>();
long minSize = 1;
List<String> include = new ArrayList<>();
List<String> exclude = new ArrayList<>();
String action = "report";
String keep = "oldest";
boolean dryRun = false;
boolean recursive = true;
String output = "dedup_report.json";
}
private static class FileInfo {
Path path;
long size;
long mtime;
String hash;
FileInfo(Path path, long size, long mtime) {
this.path = path;
this.size = size;
this.mtime = mtime;
}
}
public static void main(String[] args) {
try {
Config cfg = parseArgs(args);
if (cfg.dirs.isEmpty()) cfg.dirs = makeSample();
List<FileInfo> scanned = scan(cfg);
List<Map<String, Object>> dupSets = findDuplicates(scanned);
List<Map<String, Object>> actions = applyActions(dupSets, cfg);
long wasted = 0;
long dupFileCount = 0;
for (Map<String, Object> ds : dupSets) {
@SuppressWarnings("unchecked")
List<Map<String, Object>> files = (List<Map<String, Object>>) ds.get("files");
long size = ((Number) ds.get("size")).longValue();
wasted += size * Math.max(0, files.size() - 1);
dupFileCount += files.size();
}
Map<String, Object> summary = new LinkedHashMap<>();
summary.put("totalFilesScanned", scanned.size());
summary.put("duplicateSetsFound", dupSets.size());
summary.put("totalWastedSpace", wasted);
summary.put("totalUniqueFilesEstimate", scanned.size() - (dupFileCount - dupSets.size()));
Map<String, Object> report = new LinkedHashMap<>();
report.put("generatedAt", Instant.now().toString());
report.put("summary", summary);
report.put("duplicateSets", dupSets);
report.put("actions", actions);
System.out.println("Duplicate sets: " + dupSets.size());
System.out.println("Wasted space: " + wasted + " bytes");
Files.writeString(Paths.get(cfg.output), toJson(report));
} catch (Exception ex) {
System.err.println("Error: " + ex.getMessage());
System.exit(1);
}
}
private static Config parseArgs(String[] args) {
Config cfg = new Config();
for (int i = 0; i < args.length; i++) {
String a = args[i];
if (!a.startsWith("--")) { cfg.dirs.add(a); continue; }
switch (a) {
case "--min-size" -> cfg.minSize = Long.parseLong(args[++i]);
case "--include" -> cfg.include.add(args[++i]);
case "--exclude" -> cfg.exclude.add(args[++i]);
case "--action" -> cfg.action = args[++i];
case "--keep" -> cfg.keep = args[++i];
case "--dry-run" -> cfg.dryRun = true;
case "--no-recursive" -> cfg.recursive = false;
case "--output" -> cfg.output = args[++i];
default -> throw new IllegalArgumentException("Unknown option: " + a);
}
}
if (!List.of("report", "delete", "hardlink", "symlink").contains(cfg.action)) throw new IllegalArgumentException("Invalid --action");
if (!List.of("oldest", "newest").contains(cfg.keep)) throw new IllegalArgumentException("Invalid --keep");
return cfg;
}
private static List<FileInfo> scan(Config cfg) throws IOException {
List<FileInfo> out = new ArrayList<>();
for (String dir : cfg.dirs) {
Path root = Paths.get(dir);
if (!Files.exists(root)) {
System.err.println("Warning: missing directory " + dir);
continue;
}
if (cfg.recursive) {
Files.walkFileTree(root, new SimpleFileVisitor<>() {
@Override
public FileVisitResult visitFile(Path file, BasicFileAttributes attrs) {
tryAdd(file, cfg, out);
if (out.size() % 250 == 0) System.out.print("\rScanned " + out.size() + " files...");
return FileVisitResult.CONTINUE;
}
});
} else {
try (DirectoryStream<Path> ds = Files.newDirectoryStream(root)) {
for (Path p : ds) if (Files.isRegularFile(p)) tryAdd(p, cfg, out);
}
}
}
System.out.println();
return out;
}
private static void tryAdd(Path file, Config cfg, List<FileInfo> out) {
try {
long size = Files.size(file);
if (size < cfg.minSize) return;
String name = file.getFileName().toString();
if (!matches(name, cfg.include, true)) return;
if (matches(name, cfg.exclude, false)) return;
out.add(new FileInfo(file.toAbsolutePath(), size, Files.getLastModifiedTime(file).toMillis()));
} catch (Exception ignored) {
}
}
private static boolean matches(String name, List<String> patterns, boolean defaultValueIfEmpty) {
if (patterns.isEmpty()) return defaultValueIfEmpty;
for (String p : patterns) {
PathMatcher m = FileSystems.getDefault().getPathMatcher("glob:" + p);
if (m.matches(Paths.get(name))) return true;
}
return false;
}
private static String sha256(Path p) throws Exception {
MessageDigest md = MessageDigest.getInstance("SHA-256");
byte[] bytes = Files.readAllBytes(p);
byte[] dig = md.digest(bytes);
StringBuilder sb = new StringBuilder();
for (byte b : dig) sb.append(String.format("%02x", b));
return sb.toString();
}
private static List<Map<String, Object>> findDuplicates(List<FileInfo> files) throws Exception {
Map<Long, List<FileInfo>> bySize = new LinkedHashMap<>();
for (FileInfo fi : files) bySize.computeIfAbsent(fi.size, k -> new ArrayList<>()).add(fi);
List<Map<String, Object>> dupSets = new ArrayList<>();
for (Map.Entry<Long, List<FileInfo>> e : bySize.entrySet()) {
if (e.getValue().size() < 2) continue;
Map<String, List<FileInfo>> byHash = new LinkedHashMap<>();
for (FileInfo fi : e.getValue()) {
fi.hash = sha256(fi.path);
byHash.computeIfAbsent(fi.hash, k -> new ArrayList<>()).add(fi);
}
for (List<FileInfo> group : byHash.values()) {
if (group.size() < 2) continue;
Map<String, Object> ds = new LinkedHashMap<>();
ds.put("size", e.getKey());
ds.put("hash", group.get(0).hash);
List<Map<String, Object>> arr = new ArrayList<>();
for (FileInfo fi : group) {
Map<String, Object> row = new LinkedHashMap<>();
row.put("path", fi.path.toString());
row.put("size", fi.size);
row.put("mtime", fi.mtime);
row.put("hash", fi.hash);
arr.add(row);
}
ds.put("files", arr);
dupSets.add(ds);
}
}
return dupSets;
}
private static List<Map<String, Object>> applyActions(List<Map<String, Object>> dupSets, Config cfg) {
List<Map<String, Object>> actions = new ArrayList<>();
for (Map<String, Object> ds : dupSets) {
@SuppressWarnings("unchecked")
List<Map<String, Object>> files = (List<Map<String, Object>>) ds.get("files");
files.sort(Comparator.comparingLong(f -> ((Number) f.get("mtime")).longValue()));
Map<String, Object> keeper = "oldest".equals(cfg.keep) ? files.get(0) : files.get(files.size() - 1);
for (Map<String, Object> f : files) {
if (Objects.equals(f.get("path"), keeper.get("path"))) continue;
Map<String, Object> act = new LinkedHashMap<>();
act.put("type", cfg.action);
act.put("source", keeper.get("path"));
act.put("target", f.get("path"));
act.put("status", "planned");
if (!"report".equals(cfg.action) && !cfg.dryRun) {
try {
Path source = Paths.get(String.valueOf(keeper.get("path")));
Path target = Paths.get(String.valueOf(f.get("path")));
if ("delete".equals(cfg.action)) {
Files.deleteIfExists(target);
} else if ("hardlink".equals(cfg.action)) {
Files.deleteIfExists(target);
Files.createLink(target, source);
} else if ("symlink".equals(cfg.action)) {
Files.deleteIfExists(target);
Files.createSymbolicLink(target, target.getParent().relativize(source));
}
act.put("status", "done");
} catch (Exception ex) {
act.put("status", "failed");
act.put("error", ex.getMessage());
}
} else if (cfg.dryRun && !"report".equals(cfg.action)) {
act.put("status", "dry-run");
}
actions.add(act);
}
}
return actions;
}
private static List<String> makeSample() throws IOException {
Path root = Paths.get("sample_dedup_data").toAbsolutePath();
if (Files.exists(root)) deleteRec(root);
Files.createDirectories(root.resolve("a"));
Files.createDirectories(root.resolve("b"));
Files.writeString(root.resolve("a/x1.txt"), "hello duplicate\n");
Files.writeString(root.resolve("a/x2.txt"), "hello duplicate\n");
Files.writeString(root.resolve("b/x3.txt"), "hello duplicate\n");
Files.writeString(root.resolve("b/unique.txt"), "unique\n");
return List.of(root.toString());
}
private static void deleteRec(Path p) throws IOException {
if (Files.isDirectory(p)) {
try (DirectoryStream<Path> ds = Files.newDirectoryStream(p)) {
for (Path c : ds) deleteRec(c);
}
}
Files.deleteIfExists(p);
}
private static String toJson(Object obj) {
if (obj == null) return "null";
if (obj instanceof String s) return '"' + s.replace("\\", "\\\\").replace("\"", "\\\"") + '"';
if (obj instanceof Number || obj instanceof Boolean) return obj.toString();
if (obj instanceof Map<?, ?> m) {
StringBuilder sb = new StringBuilder("{");
boolean first = true;
for (Map.Entry<?, ?> e : m.entrySet()) {
if (!first) sb.append(',');
first = false;
sb.append(toJson(String.valueOf(e.getKey()))).append(':').append(toJson(e.getValue()));
}
sb.append('}');
return sb.toString();
}
if (obj instanceof Iterable<?> it) {
StringBuilder sb = new StringBuilder("[");
boolean first = true;
for (Object x : it) {
if (!first) sb.append(',');
first = false;
sb.append(toJson(x));
}
sb.append(']');
return sb.toString();
}
return toJson(String.valueOf(obj));
}
}