Reading these tables
Each row is one model on one task. The three score columns are the same question asked of three different surfaces of the model, all with a readout of identical size, so a difference between them is a fact about the model rather than about the measuring instrument.
B, behavioural
The output the model documentation tells you to use. This is what a conventional benchmark scores.
P, probe
The best internal layer, read with the same size of readout. Where P sits well above B, the published score understates the model.
F, elicited
What light fine-tuning recovers, under a pinned budget recorded on the card. Measured on the fold task, where it lands above both other surfaces. Shown as n/a where it was not run, and withheld where the run collapsed onto a single class rather than reported as a number.
- Best baseline is the strongest result obtainable without the model at all. A score that does not clear it is not evidence of capability.
- Cheap-readout reference is the best a conventional supervised readout achieves on the cheap features alone — amino-acid composition for proteins, raw expression for cells. Scoring above it means the representation adds something those features do not already carry. This column was previously labelled "ceiling", which was a misnomer we are correcting: it is a second and harder baseline, not an upper bound on the task, which is why most models sit above it. A genuine task ceiling, from label noise or from fully fine-tuning the same architecture, is not yet measured for any task here.
- Null p is the fraction of covariate-matched label shuffles that scored at least as well. A real effect should be rare under shuffling.
- CI95 is a 95% interval from resampling whole groups, donors or sequence clusters, rather than individual items, because items within a group are not independent.
- Gap P−B is differenced within each bootstrap replicate, so its interval accounts for the two surfaces moving together.
Results
Coverage
Result cards per modality and capability axis. A dot is an honest gap, and most of this grid is dots.