Bigger protein models do read biological function better. The effect is real but modest, and the largest model in a family is never the best one.
One family, ESM2, so the only thing changing along the axis is size. The task is spotting a transcription factor from its amino-acid sequence alone. The red line is what you get with no model at all.
How to read this, and what it does not show
Across a nearly two-thousandfold range of size the score rises about 4 points in total, roughly 1.3 points for every tenfold increase, and most of that arrives below 150 million parameters. Above it the line is flat and noisy: the 3 billion model is the best of the family and the 15 billion model sits below it, back at roughly the 150 million level. The intervals overlap heavily, so read the ordering at the top as unresolved rather than as a ranking.
Every number on this page moved when we fixed our own measurement. Scores on this task fell by 1.5 to 3.6 points once the probe layer stopped being chosen on the test set and the homology split started clustering at the identity threshold the field uses. The earlier, higher numbers were our error, not the models' achievement. What we got wrong, and who found what first →
The flattening at the top is not our discovery either. The Ankh team reported in 2023 that ESM-2 15B did not beat smaller ESM models and behaved inconsistently between runs; a 2025 study in Scientific Reports found every model above 150 million performing comparably.
There is also a confound in the axis: every ESM2 model here saw roughly the same number of training tokens regardless of size, so this ladder does not test scale under a budget that grows with the model. A NeurIPS 2024 study that re-allocated that compute beat ESM-2 3B. So this ladder stops paying; protein models in general have not stopped scaling.
It is learned biology, not just more parameters
The obvious sceptical reading of any scaling curve is that bigger models are better at everything simply for having more parameters, with nothing biological about it. That is testable, so we tested it: the same architectures at the same sizes, with the training thrown away.
Untrained models score 0.79 here against 0.88 to 0.92 for trained ones, and they do not improve with size at all: 0.07 points per tenfold increase against 2.1 for trained models. Whatever scale buys, it is bought with what the models learned, not with their parameter count.
Newer is not automatically better
Held at comparable size, the strongest result on this task belongs to a model from the end of 2020. On this measure the field has been broadly flat for four years while the models themselves changed a great deal.
Five research groups' models, all measured the same way, with size held between 300 and 700 million parameters so the axis is vintage rather than scale.
How firmly to hold that: not very
These intervals overlap almost entirely, so this is an ordering of best estimates rather than a demonstrated ranking. Our own Standard says to report how likely a ranking is to flip rather than publish a bare order, and we have not computed that here yet, so by our own rule the claim is underpowered. Read it as "nothing here has clearly moved" rather than "2020 beat 2023".
An earlier build of this page showed the 2023 model far below the others, which would have read as a sharp decline. That was a bug in how its sequences were spelled for its tokeniser, not a property of the model. It is recorded here because it is exactly the kind of plausible artefact this programme exists to catch.
Every model we have measured
Three questions, put to every model the same way. The striped rows are what you get with no model at all.
Two things worth tracking
A model can be useful and dangerous through the same understanding.
What these models make possible
Recognising protein function, recovering regulatory structure, reading cell identity.
What they make easier to misuse
A released checkpoint is a starting point, not a fixed capability. Gated on the disclosure process in the Standard.
What we can now measure on the risk side
Our first numbers here. Three viral proteins — influenza, HIV, SARS-CoV-2 — where laboratories have measured what thousands of single mutations do, and the models are asked to predict it from sequence alone.
They clear a no-model comparison only narrowly, by a few points on a scale where counting amino acids already gets about 0.51 to 0.55. On influenza the three smallest models fail the shuffled-label test outright: once you control for which residue was changed, their scores are not distinguishable from chance. Only the two largest learn anything on that virus at all.
Why we are not summarising this as one number
The three viruses disagree with each other. On SARS-CoV-2 the internal layers read more than the model's own output does, on every model. On influenza and HIV they read less. Same protocol, same models, opposite sign — so any single assay would characterise the virus as much as the method, and we report all three rather than the average.
These are published laboratory measurements of mutations other people have already made. Nothing here designs a sequence or proposes one, and the numbers are model-level aggregates rather than per-mutation predictions.
What others have shown
Filtering training data does not remove the capability
Evo 2 excluded viral sequences from training. Fifty fine-tuning steps, under an hour on one GPU, put the capability back.
Design tools can route around screening
Design software produced variants of proteins of concern that DNA synthesis screening did not reliably catch.
The question we are built to answer
As models get bigger, does the gap between what a model says and what it internally knows widen, or close? If it closes, judging models by their outputs gets safer as the field advances. If it widens, output-only assessment becomes steadily more misleading exactly as models become more capable.
Answer, on the evidence so far: it depends on the task, and we cannot yet say why. On two of five tasks the gap clearly widens as models grow — the direction that makes output-only assessment less trustworthy, not more. On two others there is no change we can detect. On a genome task it appears to close.
This page said no measurable change until 20 August 2026, on all six tasks. That was wrong, and the reason it was wrong is the most useful thing we have learned: each task had been run once, and a single run of an experiment like this has a margin of error wide enough to contain almost anything. Repeating each task six times changed the answer on two of them and left it unchanged on the rest.
The five tasks, and what each one says
Each number below is six independent repeats of the whole experiment, on model families spanning roughly a thousandfold in size. "Room to improve" is how far the models are from a perfect score — a gap cannot grow where there is nowhere left to go.
| Task | Room to improve | Does the gap change with size? |
|---|---|---|
| Fold recognition | 0.54 | Widens — all six repeats agree |
| Enzyme function | 0.66 | Widens — all six repeats agree |
| Transcription-factor family | 0.33 | No change detectable |
| Transcription-factor identity | 0.12 | No change detectable |
| Enhancer classes (genome) | 0.52 | Appears to close — but repeats disagree |
The two clear results are the strongest thing on this site. Fold recognition and enzyme function share no data, no categories and no method of splitting train from test, and they agree to within a thousandth. Twelve repeats out of twelve point the same way.
Room to improve explains part of the pattern and not all of it. The task showing nothing is the one nearest a perfect score, where a gap is squeezed shut by arithmetic rather than by anything the model is doing. But the genome task has as much room as fold recognition and points the other way. So something else separates them — modality, architecture, or the particular models — and our current design cannot tell which. We would rather say that than pick the flattering explanation.
Two of the numbers are weak and we mark them so. The genome result and a second protein family both clear zero by less than a hundredth while their repeats disagree in direction. We have twice watched a margin that narrow fail to survive more repeats, in both directions, and see no reason to trust these ones further.
Why a single run is not enough, in detail
Our single run of fold recognition gave a trend of +0.018 with a margin of −0.011 to +0.046, which we reported as no change. Six runs gave +0.020 with a margin of +0.009 to +0.031. The measurement barely moved; only the uncertainty did. The first number was a statement about how little we had run, dressed as a statement about models.
It holds up to the checks we could think of: dropping the largest model, the smallest, or the strongest performer each leaves a trend still clear of zero, and all six repeats point the same way.
The habit was learned the hard way, in the opposite direction. A second family of models appeared, on one run, to show the gap closing sharply. Run four times, the four disagreed entirely and the effect vanished. A single run misleads in whichever direction it happens to fall, which is why every figure above rests on six.
This matters beyond our own results. Benchmark papers in this field routinely report a single run across four to six model sizes. If our headline could reverse under repetition, so can others.
Why that answer was worth more than it sounded
An earlier build of this page said the gap shrank with scale. It did not. That apparent trend was an artefact of two things: choosing the best internal layer by its score on the very data used to report it, which inflated the gap by about 92% on the one model where we isolated it, and a confidence interval that covered 90% while claiming 95%. Both are fixed, and the trend disappeared with them.
The gap is also much larger where there is room for it. On the easy task it is a fraction of a point; on the hard fold-recognition task it reaches 10 points. That is consistent with the small gaps on the easy task having been squeezed shut by the ceiling rather than reflecting anything about the models, which is why we now report every gap twice, raw and relative to the room remaining.
A benchmark normally scores what a model puts out. SONDE also reads what the model holds internally, through a readout of the same size for both, so a difference is a fact about the model rather than the instrument. How the measurement works →
The third surface, what light fine-tuning recovers, is now measured properly on four models, and it lands where an elicitation measurement should: above both the output and the best internal layer, on every one. An earlier build reported it far below both. That was our error — the budget was too small to move a freshly initialised classifier at all, so the number measured our optimiser rather than the models — and those figures are withdrawn.
The interesting part is how fast. At 25 adaptation steps the models sit near the floor; by 200 the mid-sized one has already matched what the best reading of its frozen internals achieves, and by 400 it has passed it. That speed is the quantity that matters for deciding whether releasing a model's weights is releasing a fixed capability or a starting point.
Whether that recovery gets easier or harder as models grow is the question we most want to answer, and four models is not enough to answer it. The trend across them points upward and is not distinguishable from zero on a properly calibrated interval, so we are not reporting it as a trend. More rungs on the ladder, not a better story, is what settles it.
Genome models barely beat counting DNA words
The first measurements here on raw DNA rather than protein. Four sizes of one genome model family, asked to pick out regulatory switches in the human genome, with whole chromosomes held out.
They score between 0.48 and 0.50 against 0.45 for counting short DNA words with no model at all — and they do not improve with size. The largest, at half a billion parameters, is the weakest of the four. On this task the models are close to the cheap alternative, which is a different picture from the protein side, where the gap over a no-model reference is wide.
How much weight this carries
One task, one family, four sizes. It was chosen by measuring five candidate tasks and taking the one least solvable by cheap tricks: two of the others turned out to be nearly answerable by counting four-letter DNA words, and one was essentially a detector for how GC-rich a sequence is wearing a biological label. So this is the hardest of the five, and the models still barely clear the cheap alternative.
It matters beyond capability. Genome models are where the biosecurity question actually lives, and until now this programme could not measure one at all.
How far this reaches
The kinds of model this programme has reached so far, and the kinds it has not.
What this page still cannot tell you
- Every score on this page dropped by 1.5 to 3.6 points when we fixed our own measurement. The corrections are described on the prior-work page; the earlier numbers should not be cited.
- Confidence intervals on the scaling curves overlap heavily, so the ordering of the larger models is unresolved rather than measured.
- The size trend covers one model family. The vintage comparison covers five, enough to be suggestive and not enough to be a trend line.
- The biorisk track has no numbers of ours. It is gated on process, not capability. What is known so far on that side is other people's work, credited above.
- Fine-tuning is measured on four models but under a budget too small to be informative; it currently reads as a lower bound, not a capability.
- Our second task turned out to score seven of its fourteen classes, because a protein family and a homology group are nearly the same thing, so holding homologues out removes whole classes. It is reported as seven.
- The fold task has no genuine ceiling, so we do not know how much of the remaining distance is even reachable.
- Genome language models are not covered yet, and the single-cell track is mid-rerun.
- Coverage is a handful of tasks on one dataset per modality, and the layer result in particular is known to vary by dataset.
Look underneath
Every score keeps the baseline it had to beat and the test that could have killed it.
Prior work
Which findings were already published, by whom, and where the evidence goes against ours.