Benchmark · Methodology

How the benchmark will be run.

This is the pre-registered protocol, published before any result exists. Nothing here is evidence that a run has been reproduced - a published release requires these exact targets, hotspots, seeds, scoring pipeline, and hardware controls.

Targets

Five PDB structures spanning helical interfaces, viral antigens, flexible E3 handles, cytokine receptors, and a canonical PPI.

5O45PD-L1 (atezolizumab Fab interface)Helical-bias hotspot panel; canonical immuno-oncology target.
7BWJSARS-CoV-2 RBDACE2-mimetic interface; well-characterised epitope.
2LGVRBX1 (RING-box 1)E3 ligase substrate handle; small, flexible target.
3DI3IL-7Rα extracellular domainCytokine receptor; published binder benchmark exists.
1YCRMDM2–p53 interfaceClassical PPI hotspot; strong literature ground truth.

Tools tested

Proposed evaluation set. Versions will be frozen before execution; runnable containers, weight references, and commit hashes must ship with the first live release.

BindCraftv1.5.0
BoltzGenv0.4.1
RFdiffusionv1.1.0
RFdiffusion-AAv0.2.0
ProteinMPNNv1.0.1
LigandMPNNv0.3.0
SolubleMPNNv1.0.0
Chai-1v0.5.0
Boltz-2v2.0.0
AlphaFold2v2.3.2
ESMFoldv1.0.3
ESM2 (650M)esm2_t33_650M_UR50D

Metrics

We rank with ipSAE and report the rest for context.

ipSAE
Primary ranker
pooled historical estimate

Interprotein Score from Aligned Errors. Reweights PAE so only across-chain residue pairs near the predicted interface contribute. It is a candidate primary ranker, but within-target performance varies and the historical pooled AUC is not evidence of transfer to a new target.

ipTM
Confidence floor
AUC 0.628

Interface predicted-TM score from AlphaFold2 / Boltz. Used as a hard confidence floor before ipSAE reranking, not as a final ranker.

DockQ
Structural agreement
reference only

Continuous CAPRI-style score in [0, 1] comparing predicted interface to a reference complex. Reported only for targets with experimental complex structures.

pTMEnergy
Secondary signal
exploratory

Energy-style transform of pTM proposed in 2025 literature. Tracked as an exploratory secondary signal; not used to break ties on the public leaderboard.

Protocol

  1. 01
    Generate

    Each tool will run N=50 seeds per target with author-recommended defaults plus a fixed hotspot panel per PDB. Sequence length sweeps are declared up-front, never tuned post-hoc.

  2. 02
    Validate

    All designs will be predicted with Boltz-2 at recycles=3, samples=5. We keep the top-1 by ipTM as the canonical structure.

  3. 03
    Rerank

    Candidates passing ipTM ≥ 0.50 are reranked by ipSAE. Top-K per tool will advance to the leaderboard.

  4. 04
    Report

    Per-tool means, medians, and full distributions will be published. Cost (USD) and GPU-seconds will be measured on the same A100-80GB instance class.

Reproducibility

Live release requirements
  • - Exact tool versions & commit hashes
  • - Per-target hotspot YAMLs
  • - All seeds, params, and configs
  • - Versioned prediction and rerank pipeline
  • - Raw outputs for every reported cell
Public harness
Status: planned
Release gate: clean-environment rerun
Artifacts: configs, seeds, outputs, scores
Repository: linked when runnable
Limitations

We don’t claim significance we can’t defend.

  • In silico metrics are not wet-lab binding. The historical pooled ipSAE estimate is target-confounded and is not a guarantee on any target here.
  • Per-target rankings can flip with N=50 seeds. Differences inside the noise band are shown, not hidden.
  • Cost and speed depend on instance class, queue, and retries. A live release must publish raw runs so both can be re-derived.
  • When a tool fails on a target, we say so on the leaderboard. We do not silently drop runs.