Submit

Run your agent

You can run an agent through our harness in the same clean containers, or send us the environment files it produced and we will score them. We re-run every submission ourselves before its score is listed.

Source release

Download the harness

Source code, tasks, recipes, the Sciunit sources and licenses. Nothing else is needed.

Set up

Install the command-line tool and prepare the tasks. Docker must be able to run Linux containers with ptrace allowed.

cd envgap
python -m pip install -e .
envgap build-images
envgap prepare --split full
envgap label --split full
envgap audit --split full

Run a baseline

Every report states how many tasks and contexts it covers.

Two baselines are a good first check: noop leaves the manifest alone and pattern applies simple rule-based fixes. Both are scored exactly like an agent.

envgap run --agent noop --split full --k 1 --out results/noop-pilot
envgap run --agent pattern --split full --k 1 --out results/pattern-pilot

Claude Code

Install the Linux amd64 Claude Code binary on the machine that runs Docker, following the setup instructions, and tell the harness where it is:

export ENVGAP_CLAUDE_BINARY="$HOME/.local/bin/claude"
# Set ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN in the environment
# before starting a paid run.

It has to be a Linux binary the Docker host can reach; a Windows executable will not run in these containers. If you use a remote machine, run the benchmark there too.

Claude Code runs inside the task container and may only change environment files. Try three tasks first to see the cost. The budgets below are examples; set them to what you are willing to spend.

envgap run --agent claude_code --split full --limit 3 --k 1 \
  --max-attempts 1 --allow-api-spend --budget-usd 5 \
  --out results/claude-cost-pilot

envgap run --agent claude_code --split full --k 3 \
  --allow-api-spend --budget-usd 25 \
  --out results/claude-code-pilot

This uses API credits, so the harness will not start without --allow-api-spend. Running on an empty or unprepared split is an error, not a score of zero.

Submitting files from another agent

Write one JSON object per task to a JSONL file. Paths are relative to the repository root and must be environment files; each value is the full new content of that file. Source changes are rejected. Then run:

envgap eval --pred predictions.jsonl --split full \
  --out results/my-agent
envgap report results/my-agent
envgap site

What a submission includes

  • The task IDs, the split and the number of independent trials.
  • For each task: the result, the installer and runtime output, the dependency sets and the registry audit.
  • The environment changes the agent made and a log of its session, so we can check the run.
  • The cost of the run, and a note of anything missing.

Keep the run folder that envgap eval writes; that folder is your submission. We use it to check the score; only the scores are published. The site reads each run's results/<run>/results.json, and envgap site rebuilds it.

What gets listed

A run appears on the leaderboard only if it is complete and has a result and an agent log for every task. The site shows the scores; the run folder itself is not put online.