SONDE Technical record

Every result in full

This is the unsummarised view of the same result cards the summary page is built from: every score with the baseline it had to beat, the ceiling it was measured against, its confidence interval, and the null test it had to survive.

It is written for someone who wants to check the work rather than read the finding. If you want the finding, the summary page is the better door.

Reading these tables

Each row is one model on one task. The three score columns are the same question asked of three different surfaces of the model, all with a readout of identical size, so a difference between them is a fact about the model rather than about the measuring instrument.

B, behavioural

The output the model documentation tells you to use. This is what a conventional benchmark scores.

P, probe

The best internal layer, read with the same size of readout. Where P sits well above B, the published score understates the model.

F, elicited

What light fine-tuning recovers, under a pinned budget recorded on the card. Measured on the fold task, where it lands above both other surfaces. Shown as n/a where it was not run, and withheld where the run collapsed onto a single class rather than reported as a number.

  • Best baseline is the strongest result obtainable without the model at all. A score that does not clear it is not evidence of capability.
  • Cheap-readout reference is the best a conventional supervised readout achieves on the cheap features alone — amino-acid composition for proteins, raw expression for cells. Scoring above it means the representation adds something those features do not already carry. This column was previously labelled "ceiling", which was a misnomer we are correcting: it is a second and harder baseline, not an upper bound on the task, which is why most models sit above it. A genuine task ceiling, from label noise or from fully fine-tuning the same architecture, is not yet measured for any task here.
  • Null p is the fraction of covariate-matched label shuffles that scored at least as well. A real effect should be rare under shuffling.
  • CI95 is a 95% interval from resampling whole groups, donors or sequence clusters, rather than individual items, because items within a group are not independent.
  • Gap P−B is differenced within each bootstrap replicate, so its interval accounts for the two surfaces moving together.

Results

Coverage

Result cards per modality and capability axis. A dot is an honest gap, and most of this grid is dots.