How the benchmark will be run.
This is the pre-registered protocol, published before any result exists. Nothing here is evidence that a run has been reproduced - a published release requires these exact targets, hotspots, seeds, scoring pipeline, and hardware controls.
Targets
Five PDB structures spanning helical interfaces, viral antigens, flexible E3 handles, cytokine receptors, and a canonical PPI.
Tools tested
Proposed evaluation set. Versions will be frozen before execution; runnable containers, weight references, and commit hashes must ship with the first live release.
Metrics
We rank with ipSAE and report the rest for context.
Interprotein Score from Aligned Errors. Reweights PAE so only across-chain residue pairs near the predicted interface contribute. It is a candidate primary ranker, but within-target performance varies and the historical pooled AUC is not evidence of transfer to a new target.
Interface predicted-TM score from AlphaFold2 / Boltz. Used as a hard confidence floor before ipSAE reranking, not as a final ranker.
Continuous CAPRI-style score in [0, 1] comparing predicted interface to a reference complex. Reported only for targets with experimental complex structures.
Energy-style transform of pTM proposed in 2025 literature. Tracked as an exploratory secondary signal; not used to break ties on the public leaderboard.
Protocol
- 01Generate
Each tool will run N=50 seeds per target with author-recommended defaults plus a fixed hotspot panel per PDB. Sequence length sweeps are declared up-front, never tuned post-hoc.
- 02Validate
All designs will be predicted with Boltz-2 at recycles=3, samples=5. We keep the top-1 by ipTM as the canonical structure.
- 03Rerank
Candidates passing ipTM ≥ 0.50 are reranked by ipSAE. Top-K per tool will advance to the leaderboard.
- 04Report
Per-tool means, medians, and full distributions will be published. Cost (USD) and GPU-seconds will be measured on the same A100-80GB instance class.
Reproducibility
- - Exact tool versions & commit hashes
- - Per-target hotspot YAMLs
- - All seeds, params, and configs
- - Versioned prediction and rerank pipeline
- - Raw outputs for every reported cell
Status: planned Release gate: clean-environment rerun Artifacts: configs, seeds, outputs, scores Repository: linked when runnable
We don’t claim significance we can’t defend.
- In silico metrics are not wet-lab binding. The historical pooled ipSAE estimate is target-confounded and is not a guarantee on any target here.
- Per-target rankings can flip with N=50 seeds. Differences inside the noise band are shown, not hidden.
- Cost and speed depend on instance class, queue, and retries. A live release must publish raw runs so both can be re-derived.
- When a tool fails on a target, we say so on the leaderboard. We do not silently drop runs.