Evaluation protocol

Reported. Environment failures that users reported in GitHub issues. Some no longer reproduce on a clean machine and stay in as controls, where the right answer is to change nothing. How the boards differ

10 models JSON
Reported benchmark results.
# Model Agent % Resolved With feedback Underspec.Misspec.ConflictSecurity Manifest F1 Safe install Tasks Date
1
Claude Opus 4.5 Anthropic
bedrock 88.89 88.89 75.0071.43100.0080.00 0.34 0.12 18 2026-09-23
2
gpt-oss (120b) OpenAI
bedrock 83.33 94.44 75.0071.430.0080.00 0.34 0.12 18 2026-09-23
3
Devstral 2 (123B) Mistral AI
bedrock 72.22 94.44 25.0057.14100.0080.00 0.36 0.12 18 2026-09-23
4
Claude Sonnet 4.6 Anthropic
bedrock 66.67 83.33 50.0042.860.0040.00 0.35 0.17 18 2026-09-23
5
Claude Haiku 4.5 Anthropic
bedrock 61.11 83.33 0.0042.86100.0060.00 0.36 0.15 18 2026-09-23
6
Claude Sonnet 4.5 Anthropic
bedrock 61.11 77.78 75.0057.14100.0040.00 0.36 0.15 18 2026-09-23
7
Grok 4.6 xAI
bedrock 50.00 61.11 50.0071.43100.0080.00 0.34 0.12 18 2026-09-23
8
Mistral Large 3 (675B instruct) Mistral AI
bedrock 50.00 83.33 0.0057.14100.0080.00 0.35 0.13 18 2026-09-23
9
Nemotron Super 3 (120B) NVIDIA
bedrock 50.00 72.22 0.0042.86100.0060.00 0.36 0.12 18 2026-09-23
10
Llama 4 Maverick (17B instruct) Meta
bedrock 44.44 55.56 25.000.000.000.00 0.25 0.20 18 2026-09-23
Re-run by us in clean containersAgent: Amazon Bedrock · 1 context

% Resolved: tasks the agent fixed on its first attempt, before our checker told it anything (the ranking). With feedback: within five attempts, each after reading the checker's output. Underspec. / Misspec. / Conflict / Security: % Resolved on the tasks of each failure kind. Manifest F1: declared dependencies against the ones the program loaded. Safe install: declared packages that exist, are pinned and have no known advisory.

Explore EnvGap

About the benchmark