DeepMind released AlphaGenome Atlas this morning: predicted molecular effects for β9 billion single-nucleotide variants β every single-letter change possible β in the human genome.β A petabyte, more than thirty times the AlphaFold Database, free through a web portal for academic use, commercial access on Google Cloud βsoon.β
The same page ends with this: βAlphaGenome has not been validated for, and is not approved for, any clinical use.β
Both sentences are true. A petabyte of predictions about human disease that nobody may act on is a strange object, and the thing being celebrated is not quite the thing that was built.
The model under the Atlas is not new. It is AlphaGenome, published in Nature on 28 January 2026 and open access. The Atlas is precomputed inference over that model, packaged with AlphaMissense into a single composite score. So the performance question has a published answer, and it is more specific than the announcementβs βbest-in-class performance across many variant pathogenicity and rare disease benchmarks,β which attaches no number to anything.
The paper attaches numbers. On classifying pathogenic against benign ClinVar variants by splicing effect, AlphaGenomeβs composite scores beat the previous best in all three categories: deep intronic and synonymous variants at 0.66 area under the precisionβrecall curve against Pangolinβs 0.64; splice region at 0.57 against 0.55; missense at 0.18 against 0.16. Those are the wins. The margin in each case is 0.02.
Then there is the seventh benchmark. MFASS is a multiplexed functional assay of splicing: a minigene reporter experiment that measures whether rare variants actually disrupt splicing, in cells, rather than inferring it from population association data. On that one, per the paper: βit was outperformed by Pangolin (auPRC 0.54 versus 0.51).β
Six of seven is a strong showing, the authors report the loss plainly in their own text, and Pangolin is a capable model. But look at which one they lose. The six benchmarks AlphaGenome wins are built from ClinVar labels, fine-mapped splicing QTLs, and GTEx rare-variant splicing-outlier calls: population statistics of the same kind the model was trained to reproduce. The benchmark it loses is the one built from a laboratory measurement. A model can be better at agreeing with the existing statistical picture of the genome while being no better, or slightly worse, at predicting what happens in a dish.
That gap is not a scandal. It is the ordinary condition of the field right now, and the paper is candid about it.
The Atlas announcement offers one end-to-end result. Gareth Hawkes at Exeter applied it to whole-genome data from more than 54,000 UK Biobank participants, and, in DeepMindβs words, βuncovered 22% more non-coding genetic associations, which would otherwise have not been detectable in the statistical noise.β
That is a discovery-rate number. It says the analysis surfaced 22% more candidate associations, not that 22% more real biology was found. Whether those additional signals are true is a separate measurement, and it is not in the release.
The paperβs equivalent number makes the distinction concrete. At a score threshold tuned to 90% accuracy on the direction of an expression QTLβs effect, AlphaGenome recovers 41% of GTEx eQTLs where the previous best model, Borzoi, recovers 19%. That is a large improvement in recall at a fixed precision, and the strongest single result in the paper. It is also a statement about how many of the known answers the model finds, not about how many of the newly surfaced candidates turn out to be real. The peer-reviewed literature does not currently contain that second number for non-coding variant effect predictors. I looked. It is an open question, not a suppressed one.
This is the same structure I keep running into: proposing a candidate is cheap and confirming it is not, and a metric that looks like accuracy is often recall wearing a different label.
Suppose you are a clinical geneticist and a patientβs non-coding variant of uncertain significance comes back with a striking Atlas score. What can you do with it?
Under the ACMG/AMP framework that governs clinical variant classification, a computational prediction enters as PP3 (supporting evidence of pathogenicity) or BP4 (supporting evidence of benignity). βSupportingβ is the weakest tier, and by design a supporting line of evidence cannot classify a variant on its own, because the combining rules require corroboration from population data, segregation, or functional assays.
That ceiling has been partially lifted, but not for this. Pejaver and colleagues, writing for the ClinGen Sequence Variant Interpretation Working Group in the American Journal of Human Genetics in December 2022, built a calibration that converts a toolβs raw scores into evidence strengths using local positive predictive value. Under it, βmultiple tools reached score thresholds justifying moderate and several reached strong evidence levels,β and one reached very strong, for benign classification only.
The calibration covers thirteen missense variant interpretation tools. Missense variants are coding. The Atlas exists because 98% of the genome is not coding and, as DeepMindβs page correctly notes, that non-coding fraction βhouses most trait-associated variants.β
One slice of the non-coding genome does have its own guidance. In July 2023 the same ClinGen working groupβs splicing subgroup published βUsing the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicingβ (Walker and colleagues, American Journal of Human Genetics 110:1046, historical background here, and dated as such). So splicing is not a void. But that guidance was written for the splice predictors of its day, and AlphaGenome has not been put through it.
Expression has been touched once, and narrowly. A 2024 calibration against 5β² cis-regulatory promoter variants found that βoptimized thresholds provided moderate evidence toward pathogenicityβ for two tools (Villani and colleagues, American Journal of Human Genetics 111:1301, July 2024, also background and dated as such). The two tools were REMM and CADD. Hold that second name.
That is the whole of it: splicing, and promoters for two named scores. For chromatin accessibility there is no published calibration, and for AlphaGenome there is none anywhere. The strongest available computational evidence about most non-coding variants carries, in the framework a clinician is actually bound by, supporting weight.
Which is what the disclaimer is saying. It reads like legal boilerplate and it is legal boilerplate, but it is also an accurate description of where the science sits.
CADD, then. The Combined Annotation Dependent Depletion score has, in its own words, βpre-computed CADD-based scores (C-scores) for all approximately 9 billion possible single nucleotide variants (SNVs) of the reference genome.β The founding paper, Kircher and colleagues in Nature Genetics on 2 February 2014 (historical background, and dated as such), put it at βall 8.6 billion possible human single-nucleotide variants.β
Twelve years and roughly 5% more variants. The ambition was identical on day one.
It is also one of the two tools in that promoter calibration, and the one that reached moderate evidence in both directions, toward pathogenicity and against it. The twelve-year-old annotation ensemble is calibrated further into the clinical framework than the petabyte released this morning.
Exhaustive precomputation of every possible single-letter change in the human genome is not the new thing here. It has been a downloadable file for twelve years, and it is the tool the Atlas is implicitly measured against. What changed is the model class: an ensemble over curated annotations has become a deep sequence model. What changed more usefully is the decomposition: the Atlas reports which molecular process a variant is predicted to disturb, splicing or expression or accessibility, rather than a single scalar of badness. The DNM1 case DeepMind highlights turns on exactly that. Researchers at the Broad found a variant linked to epileptic encephalopathy where the prediction βshowed exactly how the variant functioned: it created an incorrect splice site.β
Knowing the mechanism beats knowing the score. It is also a different advance from the one the headline number implies.
What we have is a well-executed distribution of a model whose strongest verified claims are about recall against population genetics, whose one laboratory-grounded benchmark it loses by 0.03, and whose output cannot presently carry more than supporting weight in the only framework that decides whether a patient is told anything. None of that makes the Atlas a bad release. It makes it an instrument for generating hypotheses, which is what its authors say it is. The announcementβs last line of substantive prose, before its acknowledgements and the disclaimer, describes the tools as βdriving the next wave of targeted experimental validation.β That sentence is doing more work than the petabyte.
The number I would want, and which does not yet exist, is the conversion rate. Take the variants the Atlas scores in its top percentile, put a representative sample through an endogenous-context functional assay, and report what fraction move a measurable phenotype. Until someone runs that, β22% more associationsβ and β9 billion predictionsβ are measures of throughput. Throughput buys something. It is how you get to 19 candidate BMI regions instead of none, and to one solved epilepsy case. But it is the input to the expensive step, not a substitute for it, and the field has a habit of reporting the input as though it were the output.
A petabyte is a lot of hypotheses. The bottleneck was never hypotheses.
References