EnvGap
Can an AI coding agent get real software running on a clean machine?
Overview
EnvGap measures whether an AI coding agent can get real software to build and run on a clean machine by fixing only its environment: the dependency manifests, build files and setup scripts, never the source code. EnvGap checks that a project installs and runs in the first place, and whether the packages the agent chose are safe to install.
Each task is a real failure in Python, JavaScript, Java or C++: a dependency that no longer resolves, a system library the build silently assumes, two requirements that cannot coexist, or a package that was removed from its registry. Most come from public open-source projects; Lite tasks are small projects that AI coding agents wrote.
The agent receives the project as it was on the day of the failure, the error it produced, and a container with nothing but the language toolchain. It may edit environment files only; a change to any source file is rejected. When it is done, a fresh container installs exactly what it declared and runs the failing command again. Package registries are limited to what had been published by that date wherever the ecosystem allows it.
A task is resolved when that command runs. The leaderboard ranks by the agent's first attempt, made before our checker has told it anything, as a user's agent would have to work. Runs also give the agent up to five attempts, each after reading the checker's output; that score is shown beside it, but it is not the ranking. Two further scores say how well it was done. Manifest F1 compares the dependencies the agent declared with the files a system-call trace shows the program actually loading, so missing and unnecessary packages both cost points. Safe install is the share of declared packages that exist, are pinned to an exact version and have no known advisory.
A task is admitted only after its failure has been reproduced from scratch and a known correct environment fix has passed the same checks. Candidates that could not be reproduced are kept with the reason and are not counted.
Boards
Tasks come in three sets, ranked separately because they test different things.
- Reported
- Environment failures that users reported in GitHub issues. Some of those reports no longer reproduce on a clean machine today. They stay in as controls: the right answer is to leave a working environment working, so an agent that changes nothing still gets them. That share is the board's floor; a useful agent has to beat it.
- Repaired
- Breakages that a project's maintainers later repaired by changing only dependency or build files. Every task fails in a clean container before that repair and passes after it, so there are no controls and the floor is 0%.
- All
- Every set together. Results are usually read per board.
- Lite
- Small projects that AI coding agents wrote from a task description, which did not run on a clean machine as written. Each one fails in a clean container as written and runs once the environment part of the repair made in the study is applied, with the source left exactly as the agent wrote it. See Lite projects.
- Secure no results yet
- Tasks where the failure or its fix involves a risky package: one that was removed, never existed, has a known vulnerability or is left unpinned. Ranked by safe install first, then resolved.
- Stochastic no results yet
- Not a separate task set but a way of running one: each task is attempted three times from a fresh start, to see whether a model produces the same environment every time. Ranked by how consistent the three attempts are.
Updates
- 2026-10-06 Lite results: eight models on the 129 projects coding agents wrote. On the first attempt, before any feedback from our checker, they resolve 17% to 98%; with up to five attempts that read the checker's output, 36% to 100%. The leaderboard now ranks every board by the first attempt. One model passed seven Java tasks by having the build write and run a placeholder program; the checker now rejects any class that is not built from the project's own sources, and those seven count as unresolved. Leaderboard
- 2026-10 The EnvGap paper is on arXiv. Read it
- 2026-09-25 We ran the same models on the new Repaired tasks. Ten are ranked, resolving 37% to 89% of 19 scored tasks within five attempts (yamllint-387 is excluded: it passes with no change at all). The four best resolve 84% to 89%, yet their manifest F1 is only 0.13 to 0.15, against 0.27 to 0.35 on the Reported board. Llama 3.3 and Gemma 3 are not ranked on either board because neither made a single working tool call. Repaired board
- 2026-09-24 Added 20 tasks that maintainers had fixed in merged pull requests. We started from 1,937 candidates whose issue, pull request and fix we confirmed on GitHub. Of those, 745 already worked before the fix and 1,172 could not be reproduced in a clean container. Tasks
- 2026-09-23 First leaderboard: twelve models on Amazon Bedrock, all run with the same tool loop. The ten ranked models resolve 58% to 95% of 19 scored tasks within five attempts, but manifest F1 stays between 0.27 and 0.35 and safe install below 0.2 for all of them. A task where the agent saw the project's later fix counts as unresolved. Leaderboard
- 2026-09-22
Pilot on 4 tasks: Claude Code resolves 12 of 12 trials at k = 3; the
noopand pattern baselines resolve none. - 2026-07 Our earlier study appears at ACM REP '26: 68.3% of 300 projects generated by agents ran out of the box (Python 89.2%, JavaScript 61.9%, Java 44.0%). Paper
Resources
- Dataset on Hugging Face: the tasks, with their repository, base commit, failing command, environment recipe and labels. You can also browse it in the dataset viewer.
- leaderboards.json: every listed run, with D¹, D², D³, F1, safe install and timing for each task.
- Task pages: the issue, the error it produced, the dated checkout and the commands used to reproduce it.
- Protocol: how a run is scored, step by step, and what each metric means.
- Harness source:
build-images,prepare,run --agent claude_code,eval,report,site.