SONDE Where this sits in the literature

What others have already found

An evaluation programme that never checks whether it is repeating other people's work is not much of an evaluation programme. This page is that check, run against the published literature, and it did not come out flattering.

Most of what the summary page reports has been found before, sometimes on far more evidence than we have. One of our claims has published evidence pointing the other way. The measurement we treat as our centrepiece has already been made on a biological model by someone else. All of that is below, with sources.

Things we measured that were already known

We measured these independently, on our own task and protocol, and we still think that is worth doing — a replication on different data is evidence. But they are replications, and presenting them as discoveries would be wrong.

That protein models stop improving with size

Reported before us, more than once. In 2023 the team behind the Ankh model wrote that ESM-2 at 15 billion parameters "did not outperform the smaller models from the ESM family in all of the tasks", and added that it gave inconsistent results between runs. In 2025 a study in Scientific Reports found that on one of its two benchmark sets every model of 150 million parameters or more performed comparably, and that ESM-2 150M edged out ESM-2 3B on mutational-scanning data while being twenty times smaller. The PFMBench preprint reports the same flatness across a much larger set of tasks than ours.

What we add is not the finding but the apparatus around it: a no-model reference, resampled intervals and a shuffled-label test attached to every point, on a function task rather than a mutational one.

That the biggest model is not the best one

Same source, and it is the cleanest prior claim of the three: the Ankh authors said it explicitly in 2023, including the run-to-run instability of that particular checkpoint. That last part matters for us, because our own 15 billion result is a single run. Until we measure seed variance for it, our number should be read as one draw rather than a settled value.

We have since seen that instability from a second direction, which we did not go looking for. Fine-tuning the same checkpoint on fold recognition ran the full budget and fell apart: accuracy climbed to 0.381 and then collapsed to 0.013, under settings that every smaller model in the ladder handled without trouble. We report it as corroboration of their claim rather than as a finding of ours — they described the checkpoint as unstable first, and this is what that looks like on our instrument.

That the best layer is not the final one

Here the literature disagrees with itself, and we should not pretend otherwise. A 2026 preprint that swept layers across 13 protein language models, 15 tasks and 11 datasets found the deepest layer best in only 17.9% of cases, which supports us. The Scientific Reports study above swept layers on a different model and found the last layer most commonly best, which contradicts us. Both sets of authors conclude that the answer depends on the dataset and on whether the task concerns a whole protein or its individual residues.

This is the claim of ours most exposed to being wrong. Ours is one task on one dataset, which makes it a data point in an unsettled argument rather than a resolution of it.

Our headline task itself

Telling transcription factors from other proteins with this family of models was published in 2025, with a stricter control than ours: sequences clustered at 25% identity so that near-duplicates cannot straddle the training and test sets. Ours merged almost nothing, so on the one dimension where a direct comparison is possible, their protocol is the more rigorous. They also reach a different conclusion from the one we might have drawn from their table: their own text names ESM2-650M the best pre-trained model, and the smaller-beats-larger difference visible in their numbers sits inside their error bars.

The single-cell results are established negatives

Our record shows Geneformer failing to clear a cheap reference on cell-type transfer, and both single-cell models making donor identity easier to read off rather than harder. Anyone who follows this field already expects that.

  • A 2025 Genome Biology study evaluated these models without fine-tuning and found simple highly-variable-gene selection outperforming both Geneformer and scGPT across its metrics, with Geneformer ranking last on batch mixing and showing more batch-driven variance in its embeddings than the untransformed data.
  • A 2024 Nature Machine Intelligence Matters Arising found L1 logistic regression scoring 0.811 accuracy against scBERT's 0.766 on cell-type annotation, and a version with no pre-training at all performing no worse than the pre-trained model to within noise.
  • A 2025 Nature Methods Brief Communication benchmarked five single-cell foundation models plus two other deep-learning methods on predicting gene-perturbation effects against deliberately simple baselines, and reported that none outperformed them.

One honest qualification, since it would be easy to overclaim the agreement. These are not the same experiments as ours. The Genome Biology work scores unsupervised embedding quality and variance explained by batch; we decode donor identity with a classifier. The direction is the same and the conclusion is compatible, but "consistent with" is the accurate phrase, not "reproduces". The Matters Arising concerns scBERT, not the models we ran.

The measurement we call our centrepiece has been made before

This is the finding that cost us the most, and it is the reason this page exists.

SONDE's organising idea is to score a model's sanctioned output against what a probe of its internal layers can read, under a readout of identical size, and to treat the difference as the interesting quantity. In October 2025 a collaboration between Scale, Princeton, the University of Maryland, SecureBio and the Center for AI Safety published BioRiskEval, which does exactly that on a biological model — and goes further than we currently do.

  • It measures all three surfaces, not two. Behavioural scoring by log-likelihood, linear probes, and fine-tuning from 25 to 2000 steps. We have not measured the fine-tuning surface on a single model yet.
  • Its layer sweep is more thorough than ours: probes trained on every layer with the layer chosen afterwards, rather than a fixed choice.
  • It reports the gap head-to-head. On mutational-effect prediction, Evo 2 scored by log-likelihood reaches a correlation of about 0.03, while a linear probe on the same frozen model reaches about 0.16.
  • It answers a sharper question than ours. Evo 2 had eukaryotic viral data deliberately excluded from its training, and the paper shows that around 50 fine-tuning steps — under an hour on a single GPU — restores what the filtering removed.

We had previously described that work, in our own internal scan, as behavioural only. That was our error and it flattered us. The corrected position: measuring an elicitation gap on a biological model is not new, and neither is finding one. What remains open is doing it as standing instrumentation across many models and keeping the record current, rather than once, to make a point.

The idea has a longer history outside biology, too. Work on large language models has shown that a model's internal states can encode a correct answer while its output gives a wrong one, and separate work on deliberately capability-hidden models has shown that fine-tuning recovers what prompting cannot. We should be careful with that second strand, though: it establishes that behaviour understates capability, not that reading internals is the way to find it out.

What that leaves

Narrower than we would like, and worth stating plainly rather than dressing up.

Rules the schema enforces

Homology-aware splits, mandatory no-model baselines and group-level intervals are established practice in the protein benchmarks, not our inventions. What is unusual is that here a result which lacks its baseline, its reference and its null cannot be emitted at all — the failure is a crash, not a footnote a reviewer might miss.

Standing rather than one-off

The studies above are papers, published once. Re-running the same protocol as each new model ships, and keeping a public record that includes the models nobody got round to, is a role rather than a result — and it is currently vacant.

Caveats that cut against us

Every card carries the objections the harness could find to its own number, including the ones that undermine our headline. This page is the same instinct applied to the programme as a whole.

What we should stop claiming: that flat scaling above a few hundred million parameters is our finding, that the 15 billion reversal is news, that the best-layer effect is settled, or that measuring an elicitation gap on a biological model had not been done. None of those survive contact with the literature.

Three more defects, all found in our own favour

The count is now twenty-two. Every one of them, without exception, was biasing a number in the direction that flattered us. We keep finding them the same way: by checking the thing we were about to publish.

A trend fitted on models that failed their own test

We measured how much light fine-tuning adds, on a third task — viral fitness — and got a flat line so tight it appeared to contradict the two tasks that then showed a rise. It was ready to publish.

Three of the five models in that line had failed the null test printed on their own result cards. At those sizes the model was not doing the task at all, so the quantity being fitted was the difference between two scores that a shuffle could have produced. Two usable points cannot make a line. The analysis had every one of those p-values in front of it — our own pipeline computed them, wrote them to the card, and loaded them back — and never looked. It now refuses to report a slope unless at least three models beat their own null, and prints the count alongside any slope it does report.

A finding filed as a limitation of our tooling

ESM2-15B, the largest model we run, produced no fine-tuning number. The card explained why: "adapter does not support it." That was false. Fine-tuning ran the full budget; the card carries the curve. It diverged — accuracy climbed to 0.381 and then collapsed to 0.013.

The detector correctly withheld the number, and then the card blamed our software. That swaps the most interesting thing we saw — the largest model in the ladder falls apart under a budget every smaller model handled — for a note about our harness that any reader would skip. Cards now say which of the two happened, and name the peak and the fall.

A missing model was holding up the answer

Two runs had died of out-of-memory errors on the largest models. We treated that as an inconvenience and read the trends off the ladders we had. On enzyme function, four models gave a clear rise in how much fine-tuning adds with size — the strongest such result we had.

Recovering the missing 3-billion model changed it into a flat line. The recovered point sits below the model a fifth its size, and the interval that had excluded zero now spans it comfortably. The rung that crashed was the rung that would have contradicted us — which is not a coincidence so much as a warning: an out-of-memory error removes the biggest model, and the biggest model is exactly where a trend that is not a trend gets caught.

What that leaves us claiming

One task of three now shows fine-tuning recovering more as models grow, and it clears zero by a whisker. Repeating it under three different random seeds gives the same answer three times, so it is not noise. But two tasks say nothing, and every model in these comparisons received the same fixed training budget — which a larger model can simply do more with.

So this is not on the site as a result, and that is the point of saying it here. The measurement designed to remove that budget objection turned out to have a flaw of its own, larger than the effect it was meant to settle.

None of these defects changed a published number. All three would have, had they been found a day later, and that is the only reason this section exists rather than a correction notice.

Sources

Every source below was retrieved and checked directly rather than taken from a search summary, and each licence was read from the publisher's own statement. Where a work is openly licensed we say so; where it is not, we describe it in our own words and link out rather than reproducing anything.

SourceWhat it isLicence
Elnaggar et al., Ankh (2023) Protein language model; reports ESM-2 15B not beating smaller ESM models, and its run-to-run inconsistency CC BY-NC-SA 4.0 · preprint
Vieira, Handojo & Wilke (2025) Scientific Reports. Medium-sized protein models transfer well; finds the last layer most commonly best, against us CC BY 4.0 · open
Joeres et al. (2026) Layer sweep over 13 protein models, 15 tasks, 11 datasets; deepest layer best in 17.9% of cases CC BY 4.0 · preprint
Gao et al., PFMBench (2025) Protein foundation model benchmark across many tasks with identity-clustered splits arXiv licence · preprint, not peer reviewed
Cheng et al. (2024) NeurIPS. Compute-optimal protein language models; the token budget confound in the ESM-2 ladder Author copyright · prose and numbers only
Hassan et al. (2025) Transcription-factor identification with protein language models, clustered at 25% identity — our task, done first All rights reserved · description only
Kedzierska et al. (2025) Genome Biology. Single-cell foundation models underperform simple gene selection without fine-tuning CC BY 4.0 · open
Boiarsky et al. (2024) Nature Machine Intelligence. Logistic regression matches a single-cell foundation model on annotation Not open access · description only
Ahlmann-Eltze, Huber & Anders (2025) Nature Methods. Deep perturbation-effect prediction does not yet beat simple baselines CC BY 4.0 · open
Wei et al., BioRiskEval (2025) Behavioural, probe and fine-tuned surfaces measured on Evo 2; data filtering is not tamper-resistant CC BY 4.0 · preprint
Orgad et al. (2025) ICLR. Language model internals can encode a correct answer the output does not give Open preprint
Greenblatt et al. (2024) NeurIPS. Password-locked models; fine-tuning elicits capability that prompting does not reveal CC BY 4.0 · preprint

A note on names, since it will otherwise cause confusion. SONDE is also the name of an established protein representation benchmark — Ünsal et al., Nature Machine Intelligence 4:227–245 (2022), "Protein RepresentatiOn BEnchmark" — which is well cited and still maintained by its group. Our programme is unrelated to it and does not build on it. We picked the name without knowing theirs existed, which was our failure of homework. We are keeping it for now, since their leaderboard is not accepting active submissions and the two projects do different things, but the credit for the name belongs to them and this note is here so that nobody is misled about which is which.