← Gautam Parab

Nine Billion Predictions, and the Margin Is 0.02

DeepMind released AlphaGenome Atlas this morning: predicted molecular effects for β€œ9 billion single-nucleotide variants β€” every single-letter change possible β€” in the human genome.” A petabyte, more than thirty times the AlphaFold Database, free through a web portal for academic use, commercial access on Google Cloud β€œsoon.”

The same page ends with this: β€œAlphaGenome has not been validated for, and is not approved for, any clinical use.”

Both sentences are true. A petabyte of predictions about human disease that nobody may act on is a strange object, and the thing being celebrated is not quite the thing that was built.

AlphaGenome against the previous best model on four splicing benchmarks Paired dot plot of area under the precision-recall curve on four splicing variant benchmarks, from Avsec and colleagues, Nature, January 2026. On ClinVar deep intronic and synonymous variants the previous best model, Pangolin, scores 0.64 and AlphaGenome scores 0.66. On ClinVar splice region variants Pangolin scores 0.55 and AlphaGenome 0.57. On ClinVar missense variants DeltaSplice and Pangolin score 0.16 and AlphaGenome 0.18. Each of those three margins is 0.02. On MFASS, a minigene reporter assay of experimentally measured splice disruption, the order reverses: Pangolin scores 0.54 and AlphaGenome 0.51. SPLICING VARIANT BENCHMARKS Β· auPRC Β· AVSEC ET AL., NATURE, 28 JAN 2026 previous best model AlphaGenome composite ClinVar: deep intronic + synonymous ClinVar: splice region ClinVar: missense MFASS: lab-measured splice disruption 0.64 β€” Pangolin, ClinVar deep intronic and synonymous 0.55 β€” Pangolin, ClinVar splice region 0.16 β€” DeltaSplice and Pangolin, ClinVar missense 0.54 β€” Pangolin, MFASS 0.66 β€” AlphaGenome, ClinVar deep intronic and synonymous 0.57 β€” AlphaGenome, ClinVar splice region 0.18 β€” AlphaGenome, ClinVar missense 0.51 β€” AlphaGenome, MFASS 0.64 0.66 0.55 0.57 0.16 0.18 0.51 0.54 the one benchmark built from a laboratory measurement 0.0 0.2 0.4 0.6 Six of seven benchmarks in the paper go to AlphaGenome; four are shown. Higher is better.
The three ClinVar wins are 0.02 apart. The reversal is on MFASS, the only one of the four that measures splicing in cells rather than inferring it from population data.

The margins

The model under the Atlas is not new. It is AlphaGenome, published in Nature on 28 January 2026 and open access. The Atlas is precomputed inference over that model, packaged with AlphaMissense into a single composite score. So the performance question has a published answer, and it is more specific than the announcement’s β€œbest-in-class performance across many variant pathogenicity and rare disease benchmarks,” which attaches no number to anything.

The paper attaches numbers. On classifying pathogenic against benign ClinVar variants by splicing effect, AlphaGenome’s composite scores beat the previous best in all three categories: deep intronic and synonymous variants at 0.66 area under the precision–recall curve against Pangolin’s 0.64; splice region at 0.57 against 0.55; missense at 0.18 against 0.16. Those are the wins. The margin in each case is 0.02.

Then there is the seventh benchmark. MFASS is a multiplexed functional assay of splicing: a minigene reporter experiment that measures whether rare variants actually disrupt splicing, in cells, rather than inferring it from population association data. On that one, per the paper: β€œit was outperformed by Pangolin (auPRC 0.54 versus 0.51).”

Six of seven is a strong showing, the authors report the loss plainly in their own text, and Pangolin is a capable model. But look at which one they lose. The six benchmarks AlphaGenome wins are built from ClinVar labels, fine-mapped splicing QTLs, and GTEx rare-variant splicing-outlier calls: population statistics of the same kind the model was trained to reproduce. The benchmark it loses is the one built from a laboratory measurement. A model can be better at agreeing with the existing statistical picture of the genome while being no better, or slightly worse, at predicting what happens in a dish.

That gap is not a scandal. It is the ordinary condition of the field right now, and the paper is candid about it.

What 22% means

The Atlas announcement offers one end-to-end result. Gareth Hawkes at Exeter applied it to whole-genome data from more than 54,000 UK Biobank participants, and, in DeepMind’s words, β€œuncovered 22% more non-coding genetic associations, which would otherwise have not been detectable in the statistical noise.”

That is a discovery-rate number. It says the analysis surfaced 22% more candidate associations, not that 22% more real biology was found. Whether those additional signals are true is a separate measurement, and it is not in the release.

The paper’s equivalent number makes the distinction concrete. At a score threshold tuned to 90% accuracy on the direction of an expression QTL’s effect, AlphaGenome recovers 41% of GTEx eQTLs where the previous best model, Borzoi, recovers 19%. That is a large improvement in recall at a fixed precision, and the strongest single result in the paper. It is also a statement about how many of the known answers the model finds, not about how many of the newly surfaced candidates turn out to be real. The peer-reviewed literature does not currently contain that second number for non-coding variant effect predictors. I looked. It is an open question, not a suppressed one.

This is the same structure I keep running into: proposing a candidate is cheap and confirming it is not, and a metric that looks like accuracy is often recall wearing a different label.

How much evidence weight a computational prediction can carry Four stacked tiers of ACMG/AMP evidence strength, weakest at the bottom: supporting, moderate, strong, very strong. For missense variant predictors, which the ClinGen calibration of Pejaver and colleagues covers for thirteen tools, the reachable tiers are supporting, moderate and strong, with very strong reached by one tool for benign classification only. For non-coding predictors, including AlphaGenome, only the supporting tier is shown as reachable: no calibration of AlphaGenome to these tiers has been published, and outside splicing, which has its own separate ClinGen recommendations from 2023, the only other published non-coding calibration covers 5 prime cis-regulatory promoter variants scored by CADD and REMM (Villani and colleagues, 2024), which reached moderate evidence; chromatin accessibility predictors and AlphaGenome itself remain uncalibrated. A note records that supporting evidence cannot classify a variant on its own under the combining rules. ACMG/AMP EVIDENCE STRENGTH Β· PP3/BP4 Β· CLINGEN CALIBRATION, PEJAVER ET AL. 2022 Missense predictors Non-coding predictors 13 tools calibrated AlphaGenome, uncalibrated Very strong Strong Moderate Supporting Several calibrated missense tools reach strong evidence β€” Pejaver et al., AJHG, December 2022 Multiple calibrated missense tools reach moderate evidence β€” Pejaver et al., AJHG, December 2022 Supporting is the default strength for PP3/BP4 β€” ACMG/AMP, 2015 Non-coding predictions enter at supporting strength, the ACMG/AMP default for PP3 and BP4 one tool, benign only not calibrated to these tiers not calibrated to these tiers not calibrated to these tiers Supporting evidence cannot classify a variant alone; the combining rules require corroboration. Splicing (Walker et al. 2023) and promoter variants (Villani et al. 2024) have separate calibrations; AlphaGenome is in none.
The calibration that lifts a computational score above the weakest tier was built for coding variants. The Atlas exists for the other 98% of the genome, where the ladder it has been fitted to has one rung.

What the score is allowed to count for

Suppose you are a clinical geneticist and a patient’s non-coding variant of uncertain significance comes back with a striking Atlas score. What can you do with it?

Under the ACMG/AMP framework that governs clinical variant classification, a computational prediction enters as PP3 (supporting evidence of pathogenicity) or BP4 (supporting evidence of benignity). β€œSupporting” is the weakest tier, and by design a supporting line of evidence cannot classify a variant on its own, because the combining rules require corroboration from population data, segregation, or functional assays.

That ceiling has been partially lifted, but not for this. Pejaver and colleagues, writing for the ClinGen Sequence Variant Interpretation Working Group in the American Journal of Human Genetics in December 2022, built a calibration that converts a tool’s raw scores into evidence strengths using local positive predictive value. Under it, β€œmultiple tools reached score thresholds justifying moderate and several reached strong evidence levels,” and one reached very strong, for benign classification only.

The calibration covers thirteen missense variant interpretation tools. Missense variants are coding. The Atlas exists because 98% of the genome is not coding and, as DeepMind’s page correctly notes, that non-coding fraction β€œhouses most trait-associated variants.”

One slice of the non-coding genome does have its own guidance. In July 2023 the same ClinGen working group’s splicing subgroup published β€œUsing the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing” (Walker and colleagues, American Journal of Human Genetics 110:1046, historical background here, and dated as such). So splicing is not a void. But that guidance was written for the splice predictors of its day, and AlphaGenome has not been put through it.

Expression has been touched once, and narrowly. A 2024 calibration against 5β€² cis-regulatory promoter variants found that β€œoptimized thresholds provided moderate evidence toward pathogenicity” for two tools (Villani and colleagues, American Journal of Human Genetics 111:1301, July 2024, also background and dated as such). The two tools were REMM and CADD. Hold that second name.

That is the whole of it: splicing, and promoters for two named scores. For chromatin accessibility there is no published calibration, and for AlphaGenome there is none anywhere. The strongest available computational evidence about most non-coding variants carries, in the framework a clinician is actually bound by, supporting weight.

Which is what the disclaimer is saying. It reads like legal boilerplate and it is legal boilerplate, but it is also an accurate description of where the science sits.

Genome-wide precomputed variant scores, 2014 and 2026 A near-flat line connecting two points twelve years apart. CADD version 1.0, published February 2014, precomputed scores for 8.6 billion possible single-nucleotide variants of the reference genome; the current CADD resource states approximately 9 billion. AlphaGenome Atlas, released September 2026, publishes predictions for 9 billion single-nucleotide variants. The count rose by roughly 5% across twelve years. What changed is the model class, from an ensemble over curated annotations to a deep sequence model, and the output, from one scalar score to a prediction decomposed by molecular process. GENOME-WIDE PRECOMPUTED SNVs Β· KIRCHER ET AL. 2014 Β· DEEPMIND, 8 SEP 2026 0 8.6 billion possible SNVs precomputed β€” Kircher et al., Nature Genetics, February 2014 9 billion SNVs predicted β€” AlphaGenome Atlas, September 2026 CADD v1.0 AlphaGenome Atlas 8.6 billion 9 billion February 2014 September 2026 twelve years, about 5% more variants What changed: an ensemble over curated annotations became a deep sequence model, and one scalar of predicted badness became a prediction decomposed by molecular process. The 2014 paper states 8.6 billion; CADD's current page states approximately 9 billion.
Exhaustive precomputation of every possible single-letter change has been a downloadable file since 2014. The count grew by about 5%; the advance is in what the number means.

The part that is twelve years old

CADD, then. The Combined Annotation Dependent Depletion score has, in its own words, β€œpre-computed CADD-based scores (C-scores) for all approximately 9 billion possible single nucleotide variants (SNVs) of the reference genome.” The founding paper, Kircher and colleagues in Nature Genetics on 2 February 2014 (historical background, and dated as such), put it at β€œall 8.6 billion possible human single-nucleotide variants.”

Twelve years and roughly 5% more variants. The ambition was identical on day one.

It is also one of the two tools in that promoter calibration, and the one that reached moderate evidence in both directions, toward pathogenicity and against it. The twelve-year-old annotation ensemble is calibrated further into the clinical framework than the petabyte released this morning.

Exhaustive precomputation of every possible single-letter change in the human genome is not the new thing here. It has been a downloadable file for twelve years, and it is the tool the Atlas is implicitly measured against. What changed is the model class: an ensemble over curated annotations has become a deep sequence model. What changed more usefully is the decomposition: the Atlas reports which molecular process a variant is predicted to disturb, splicing or expression or accessibility, rather than a single scalar of badness. The DNM1 case DeepMind highlights turns on exactly that. Researchers at the Broad found a variant linked to epileptic encephalopathy where the prediction β€œshowed exactly how the variant functioned: it created an incorrect splice site.”

Knowing the mechanism beats knowing the score. It is also a different advance from the one the headline number implies.

What I would want to see

What we have is a well-executed distribution of a model whose strongest verified claims are about recall against population genetics, whose one laboratory-grounded benchmark it loses by 0.03, and whose output cannot presently carry more than supporting weight in the only framework that decides whether a patient is told anything. None of that makes the Atlas a bad release. It makes it an instrument for generating hypotheses, which is what its authors say it is. The announcement’s last line of substantive prose, before its acknowledgements and the disclaimer, describes the tools as β€œdriving the next wave of targeted experimental validation.” That sentence is doing more work than the petabyte.

The number I would want, and which does not yet exist, is the conversion rate. Take the variants the Atlas scores in its top percentile, put a representative sample through an endogenous-context functional assay, and report what fraction move a measurable phenotype. Until someone runs that, β€œ22% more associations” and β€œ9 billion predictions” are measures of throughput. Throughput buys something. It is how you get to 19 candidate BMI regions instead of none, and to one solved epilepsy case. But it is the input to the expensive step, not a substitute for it, and the field has a habit of reporting the input as though it were the output.

A petabyte is a lot of hypotheses. The bottleneck was never hypotheses.

References

  1. Google DeepMind. (2026, September 8). AlphaGenome Atlas announcement.
  2. Avsec et al. (2026, January 28). Advancing regulatory variant effect prediction with AlphaGenome. Nature, 649, 1206–1218. DOI: 10.1038/s41586-025-10014-0. Read in the open-access version at PMC12851941. Benchmark figures here are quoted from this paper’s own text, not from coverage.
  3. Pejaver et al. (2022, December). American Journal of Human Genetics, 109, 2163–2177. DOI: 10.1016/j.ajhg.2022.10.013. Cited as the current ClinGen calibration standard, dated as such.
  4. Walker et al. (2023, July). ClinGen SVI Splicing Subgroup recommendations. American Journal of Human Genetics, 110, 1046–1067. DOI: 10.1016/j.ajhg.2023.06.002. Dated background.
  5. Villani et al. (2024, July). A cis-regulatory calibration. American Journal of Human Genetics, 111, 1301–1315. DOI: 10.1016/j.ajhg.2024.05.002. Dated background.
  6. Richards et al. (2015). Genetics in Medicine, 17, 405–424. Historical background for the ACMG/AMP framework.
  7. University of Washington. CADD project page.
  8. Kircher et al. (2014, February 2). Nature Genetics, 46, 310–315.