Steps
- Checkout. Check out the task's base commit and keep the log. The agent's checkout holds only that commit and its ancestors; any network access to the upstream repository is logged separately. See later history for how earlier runs were handled.
- Container. Start the language's Ubuntu 22.04 image with Sciunit installed, and record which date limits on the registry are actually enforced.
- Apply the submission. Copy in the agent's environment files: manifests, setup instructions, Dockerfile and README. Any change to a source file is rejected.
- Install. Run the language's installer and record the installed dependency set D², the exit code, the time taken and the full output.
- Trace. Run the task's command under Sciunit and turn the files it opened into D³. The seven standard C++ base libraries are left out.
- Audit. Check every declared package: does it exist, was it removed, is it pinned, does the pinned version have an OSV advisory, is its name one typo away from a popular package, and does it run install scripts.
- Score. The task is resolved if both the install and the traced run succeed. We also save dependency precision, recall and F1, the safe-install share, the number of attempts and the cost.
Metrics
- resolved%
resolved / nShare of trials resolved, shown as a percentage.
- pass@k
1 - C(n-c, k)/C(n, k)The usual unbiased estimator, reported at k = 1 and k = 3, where n is the number of independent samples for a task and c the number that succeeded. With too few samples, no value is reported.
- consistency@k
|tasks with identical D¹ across k runs| / nCompares the normalized declared dependency sets from separate fresh contexts.
- manifest_F1
F1(D¹, D³) from prf()Averaged within each language, then across languages, using the same normalization code as our earlier work.
- safe_install
|{d : exists, not removed, not vulnerable, pinned}| / |D¹|Registry and OSV responses are cached by language, name, version and date. A package we could not check does not count as safe.
- attempts
mean attempts over resolved tasks- cost
USD spent on model calls, mean per task
Boards
All holds every admitted task, including controls. Lite holds small projects that AI coding agents wrote, each of which fails in a clean container as written and runs once the environment part of the repair made in the study is applied; see Lite projects.
Secure holds tasks whose original manifest was flagged by the audit or that carry a security label; it is ranked by safe install, then by resolved. Stochastic runs Lite three times from fresh contexts and reports consistency@3 and mean pairwise Jaccard.
Newest is not the same as safe
For each pinned package we compare the declared version with the newest one on the registry today and the newest one published before the issue was filed, where the registry lets us. An older version is not automatically vulnerable, and a newer one is not automatically safe.
Loosening a pin or turning off install scripts can make a build pass while making it less reproducible or changing what runs. That is why safe install counts only pinned packages: a build that passes after a pin was loosened still loses points.
Repeated runs disagree
Repeated trials start from the same dated checkout in separate contexts. Identical normalized D¹ sets count toward consistency@3, and mean pairwise Jaccard captures partial overlap. Both describe which dependencies an agent picks, separately from whether its environment works.
Limits of dating a registry
Pointing pip at a different index URL does not give you PyPI as it was. npm's publish-date cutoff cannot bring back removed packages, and lockfiles and direct URLs still have to be checked. Maven's update policy controls when it refreshes, not which releases it sees. For C++, system packages and compiler versions need their own snapshots. Wherever we could not enforce a date, the task notes say so.