The methodology is public before the results are.
Every target, hotspot, seed count, and hit threshold was declared and signed in advance. Runs are not published yet; results land here once the pipeline reproduces cleanly.
The protocol is public and frozen. The runs are not published yet — the execution pipeline is still being validated end-to-end, and we will not show a number we cannot reproduce from a clean environment. The leaderboard appears here the day the first run set clears that gate, against thresholds that were written down long before we saw them.
- Pre-registered
- Targets, hotspots, seed counts, hit thresholds, and statistics were fixed and published before any run was scored. Nothing can be tuned after seeing results.
- Signed & versioned
- Methodology v1.0 is git-signed and immutable after publication. Changing a threshold, target, or statistic forks a new version with a public diff.
- External advisory panel
- Three external members — structural biology, experimental protein engineering, and ML methodology — review methodology changes and adjudicate reproducibility disputes.
- 7-day public comment
- Every proposed new target goes through a seven-day public comment window before the panel votes on admission.
A result is only a result if you can re-run it.
- Pinned tool versions, container digests, and weight references
- Per-target hotspot files, configs, and the full seed range
- Raw outputs and scores for every reported cell
- A clean-environment rerun that reproduces the published aggregates
Until all four hold for a run set, this page shows the protocol and nothing else. That is the guarantee, not a placeholder.
How we measure.
Declared in advance: the controls and the evidence a run set must produce before any number is published here.
Review the protocol draft- 01Targets
The proposed panel uses seven canonical targets spanning surface chemistries, sizes, and design difficulty. Hotspots will be declared once per target, frozen before execution, and reused across every tool.
- 02Validation pipeline
The proposed pipeline folds each generated design with Boltz-2 against the target receptor, then records ipTM, interface pAE, pDockQ, and ipSAE. The first live release must publish the exact scorer versions and justify the primary ranker from cited validation data.
- 03Filtering thresholds
The draft Tier-A rule is ipSAE ≥ 0.35, ipTM ≥ 0.60, and interface pLDDT ≥ 70. These thresholds remain proposal values until their calibration sources and sensitivity analysis ship with the benchmark.
- 04Cost methodology
GPU time will be measured wall-clock on a declared instance class. Cost-per-success will use captured run cost divided by Tier-A hits; closed-weight API charges and open-weight GPU costs will be reported separately.
- 05Reproducibility
Promotion to a live benchmark requires committed configs and seeds, pinned images and weights, public target inputs and hotspot files, and a clean-environment rerun from the published harness.
- 06Statistics
The planned report includes 95% Wilson confidence intervals and paired bootstrap comparisons. Rank claims will be withheld when the declared uncertainty criteria are not met.
- 07Submission process
The submission process is still being validated. Open-weight tools will require a pinned runnable repository; closed-weight tools will require an auditable API endpoint that accepts the standard inputs.
Built a binder design tool? Register it for the first run set.
Share your tool name, repository, paper, and contact email. We’ll follow up as the reproducible submission harness and review process are validated.
- Public, open-weight tools added free of charge
- Closed-weight tools require an API endpoint we can call
- Published results include configs, seeds, and an audit trail
- Release cadence begins after the first reproducible run set