Source release
Download the harness
Source code, tasks, recipes, the Sciunit sources and licenses. Nothing else is needed.
Set up
Install the command-line tool and prepare the tasks. Docker must be able to run Linux containers with ptrace allowed.
cd envgap python -m pip install -e . envgap build-images envgap prepare --split full envgap label --split full envgap audit --split full
Run a baseline
Every report states how many tasks and contexts it covers.
Two baselines are a good first check: noop leaves the manifest alone and pattern applies simple rule-based fixes. Both are scored exactly like an agent.
envgap run --agent noop --split full --k 1 --out results/noop-pilot envgap run --agent pattern --split full --k 1 --out results/pattern-pilot
Claude Code
Install the Linux amd64 Claude Code binary on the machine that runs Docker, following the setup instructions, and tell the harness where it is:
export ENVGAP_CLAUDE_BINARY="$HOME/.local/bin/claude" # Set ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN in the environment # before starting a paid run.
It has to be a Linux binary the Docker host can reach; a Windows executable will not run in these containers. If you use a remote machine, run the benchmark there too.
Claude Code runs inside the task container and may only change environment files. Try three tasks first to see the cost. The budgets below are examples; set them to what you are willing to spend.
envgap run --agent claude_code --split full --limit 3 --k 1 \ --max-attempts 1 --allow-api-spend --budget-usd 5 \ --out results/claude-cost-pilot envgap run --agent claude_code --split full --k 3 \ --allow-api-spend --budget-usd 25 \ --out results/claude-code-pilot
This uses API credits, so the harness will not start without --allow-api-spend. Running on an empty or unprepared split is an error, not a score of zero.
Submitting files from another agent
Write one JSON object per task to a JSONL file. Paths are relative to the repository root and must be environment files; each value is the full new content of that file. Source changes are rejected. Then run:
envgap eval --pred predictions.jsonl --split full \ --out results/my-agent envgap report results/my-agent envgap site
What a submission includes
- The task IDs, the split and the number of independent trials.
- For each task: the result, the installer and runtime output, the dependency sets and the registry audit.
- The environment changes the agent made and a log of its session, so we can check the run.
- The cost of the run, and a note of anything missing.
Keep the run folder that envgap eval writes; that folder is your submission. We use it to check the score; only the scores are published. The site reads each run's results/<run>/results.json, and envgap site rebuilds it.
What gets listed
A run appears on the leaderboard only if it is complete and has a result and an agent log for every task. The site shows the scores; the run folder itself is not put online.