Reported. Environment failures that users reported in GitHub issues. Some no longer reproduce on a clean machine and stay in as controls, where the right answer is to change nothing. How the boards differ
| # | Model | Agent | % Resolved | With feedback | Underspec. | Misspec. | Conflict | Security | Manifest F1 | Safe install | Tasks | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 |
Claude Opus 4.5
Anthropic
|
bedrock | 88.89 | 88.89 | 75.00 | 71.43 | 100.00 | 80.00 | 0.34 | 0.12 | 18 | 2026-09-23 |
| 2 |
gpt-oss (120b)
OpenAI
|
bedrock | 83.33 | 94.44 | 75.00 | 71.43 | 0.00 | 80.00 | 0.34 | 0.12 | 18 | 2026-09-23 |
| 3 |
Devstral 2 (123B)
Mistral AI
|
bedrock | 72.22 | 94.44 | 25.00 | 57.14 | 100.00 | 80.00 | 0.36 | 0.12 | 18 | 2026-09-23 |
| 4 |
Claude Sonnet 4.6
Anthropic
|
bedrock | 66.67 | 83.33 | 50.00 | 42.86 | 0.00 | 40.00 | 0.35 | 0.17 | 18 | 2026-09-23 |
| 5 |
Claude Haiku 4.5
Anthropic
|
bedrock | 61.11 | 83.33 | 0.00 | 42.86 | 100.00 | 60.00 | 0.36 | 0.15 | 18 | 2026-09-23 |
| 6 |
Claude Sonnet 4.5
Anthropic
|
bedrock | 61.11 | 77.78 | 75.00 | 57.14 | 100.00 | 40.00 | 0.36 | 0.15 | 18 | 2026-09-23 |
| 7 |
Grok 4.6
xAI
|
bedrock | 50.00 | 61.11 | 50.00 | 71.43 | 100.00 | 80.00 | 0.34 | 0.12 | 18 | 2026-09-23 |
| 8 |
Mistral Large 3 (675B instruct)
Mistral AI
|
bedrock | 50.00 | 83.33 | 0.00 | 57.14 | 100.00 | 80.00 | 0.35 | 0.13 | 18 | 2026-09-23 |
| 9 |
Nemotron Super 3 (120B)
NVIDIA
|
bedrock | 50.00 | 72.22 | 0.00 | 42.86 | 100.00 | 60.00 | 0.36 | 0.12 | 18 | 2026-09-23 |
| 10 |
Llama 4 Maverick (17B instruct)
Meta
|
bedrock | 44.44 | 55.56 | 25.00 | 0.00 | 0.00 | 0.00 | 0.25 | 0.20 | 18 | 2026-09-23 |
| No models match this search. | ||||||||||||
% Resolved: tasks the agent fixed on its first attempt, before our checker told it anything (the ranking). With feedback: within five attempts, each after reading the checker's output. Underspec. / Misspec. / Conflict / Security: % Resolved on the tasks of each failure kind. Manifest F1: declared dependencies against the ones the program loaded. Safe install: declared packages that exist, are pinned and have no known advisory.
Repaired. Breakages that maintainers later repaired by changing only dependency or build files. Every task fails before that repair and passes after it. How the boards differ
| # | Model | Agent | % Resolved | With feedback | Underspec. | Misspec. | Conflict | Security | Manifest F1 | Safe install | Tasks | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 |
Claude Sonnet 4.6
Anthropic
|
bedrock | 100.00 | 100.00 | 100.00 | 100.00 | – | 100.00 | 0.14 | 0.00 | 13 | 2026-09-25 |
| 2 |
Claude Opus 4.5
Anthropic
|
bedrock | 84.62 | 92.31 | 100.00 | 91.67 | – | 100.00 | 0.14 | 0.00 | 13 | 2026-09-25 |
| 3 |
Claude Haiku 4.5
Anthropic
|
bedrock | 69.23 | 76.92 | 50.00 | 66.67 | – | 100.00 | 0.14 | 0.00 | 13 | 2026-09-25 |
| 4 |
Devstral 2 (123B)
Mistral AI
|
bedrock | 69.23 | 92.31 | 100.00 | 66.67 | – | 0.00 | 0.14 | 0.00 | 13 | 2026-09-25 |
| 5 |
Claude Sonnet 4.5
Anthropic
|
bedrock | 61.54 | 69.23 | 100.00 | 66.67 | – | 0.00 | 0.14 | 0.00 | 13 | 2026-09-25 |
| 6 |
Mistral Large 3 (675B instruct)
Mistral AI
|
bedrock | 61.54 | 69.23 | 50.00 | 58.33 | – | 0.00 | 0.20 | 0.00 | 13 | 2026-09-25 |
| 7 |
gpt-oss (120b)
OpenAI
|
bedrock | 53.85 | 92.31 | 100.00 | 58.33 | – | 0.00 | 0.19 | 0.01 | 13 | 2026-09-25 |
| 8 |
Nemotron Super 3 (120B)
NVIDIA
|
bedrock | 38.46 | 76.92 | 50.00 | 41.67 | – | 100.00 | 0.20 | 0.01 | 13 | 2026-09-25 |
| 9 |
Grok 4.6
xAI
|
bedrock | 30.77 | 30.77 | 0.00 | 33.33 | – | 100.00 | 0.14 | 0.00 | 13 | 2026-09-25 |
| 10 |
Llama 4 Maverick (17B instruct)
Meta
|
bedrock | 15.38 | 46.15 | 100.00 | 16.67 | – | 0.00 | 0.00 | 0.23 | 13 | 2026-09-25 |
| No models match this search. | ||||||||||||
% Resolved: tasks the agent fixed on its first attempt, before our checker told it anything (the ranking). With feedback: within five attempts, each after reading the checker's output. Underspec. / Misspec. / Conflict / Security: % Resolved on the tasks of each failure kind. Manifest F1: declared dependencies against the ones the program loaded. Safe install: declared packages that exist, are pinned and have no known advisory.
All. Every set together. Results are normally read per board. No results yet How the boards differ
| # | Model | Agent | % Resolved | With feedback | Underspec. | Misspec. | Conflict | Security | Manifest F1 | Safe install | Tasks | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| No checked runs on All yet. How to submit a run. |
||||||||||||
Lite. Projects that AI coding agents wrote which did not run on a clean machine. Each one fails as written and runs once the environment part of the repair made in the study is applied, with the source left as the agent wrote it. How the boards differ
| # | Model | Agent | % Resolved | With feedback | Underspec. | Misspec. | Conflict | Security | Manifest F1 | Safe install | Tasks | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 |
Grok 4.6
xAI
|
bedrock | 98.45 | 100.00 | 98.18 | 98.70 | – | – | 0.50 | 0.73 | 129 | 2026-10-02 |
| 2 |
Claude Sonnet 4.6
Anthropic
|
bedrock | 94.57 | 98.45 | 98.18 | 90.91 | – | – | 0.41 | 0.63 | 129 | 2026-10-02 |
| 3 |
Claude Haiku 4.5
Anthropic
|
bedrock | 89.15 | 96.12 | 98.18 | 83.12 | – | – | 0.43 | 0.62 | 129 | 2026-10-02 |
| 4 |
Devstral 2 (123B)
Mistral AI
|
bedrock | 62.02 | 86.82 | 92.73 | 41.56 | – | – | 0.43 | 0.58 | 129 | 2026-10-02 |
| 5 |
Mistral Large 3 (675B instruct)
Mistral AI
|
bedrock | 60.47 | 88.37 | 76.36 | 48.05 | – | – | 0.43 | 0.62 | 129 | 2026-10-02 |
| 6 |
Nemotron Super 3 (120B)
NVIDIA
|
bedrock | 48.06 | 77.52 | 89.09 | 19.48 | – | – | 0.43 | 0.59 | 129 | 2026-10-02 |
| 7 |
gpt-oss (120b)
OpenAI
|
bedrock | 43.41 | 82.17 | 56.36 | 35.06 | – | – | 0.43 | 0.63 | 129 | 2026-10-02 |
| 8 |
Llama 4 Maverick (17B instruct)
Meta
|
bedrock | 17.05 | 36.43 | 40.00 | 0.00 | – | – | 0.39 | 0.68 | 129 | 2026-10-02 |
| No models match this search. | ||||||||||||
% Resolved: tasks the agent fixed on its first attempt, before our checker told it anything (the ranking). With feedback: within five attempts, each after reading the checker's output. Underspec. / Misspec. / Conflict / Security: % Resolved on the tasks of each failure kind. Manifest F1: declared dependencies against the ones the program loaded. Safe install: declared packages that exist, are pinned and have no known advisory.
Secure. Tasks involving a risky package: removed, nonexistent, vulnerable or unpinned. Ranked by safe install first. No results yet How the boards differ
| # | Model | Agent | % Resolved | With feedback | Underspec. | Misspec. | Conflict | Security | Manifest F1 | Safe install | Tasks | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| No checked runs on Secure yet. Secure membership comes from the registry audit of each task's original manifest. How to submit a run. |
||||||||||||
Stochastic. Each task attempted three times from a fresh start, ranked by how consistent the attempts are. No results yet How the boards differ
| # | Model | Agent | consistency@3 | With feedback | Jaccard | Safe install | Tasks | Date | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| No checked runs on Stochastic yet. Runs with --k 3 appear here.
How to submit a run. |
||||||||||||
Explore EnvGap
About the benchmarkTasks
Real environment failures in Python, JavaScript, Java and C++ projects, each with its original report and how we reproduced it.
Browse tasksDataset
Every task with its dated checkout, failing command and environment recipe, on Hugging Face.
Open on Hugging FaceTest your agent
Run it through the same harness and clean containers we use.
Get started