On 19 January 2026, Nature published an Author Correction to a paper it had run twenty-six months earlier. The original, from Gerbrand Ceder’s group at Berkeley, described A-Lab: a robotic laboratory that took computationally predicted target compounds, planned syntheses from the published literature, ran them, read the X-ray diffraction patterns, and decided what it had made. Over 17 days of continuous operation it ran 355 experiments and reported successfully synthesizing 41 of 58 targets.
The correction runs to a few paragraphs. Two things in it matter. First, on novelty: the authors write that the original claims “were subject to misinterpretation — their intention was to indicate that the materials were new to the prediction platform, not necessarily new to science.” Second, on identification: after manually re-analyzing the diffraction patterns, they confirm the platform “came to the correct conclusion in 36 of its 40 reported successes, with 4 compounds being inconclusive.” One compound, Zn2Cr3FeO8, was removed from the discussion entirely because it had been in the training data.
The easy read of this is wrong. It is not a scandal, and the correction is not a retraction. A group published a result, outside chemists went at it hard, the group re-did the analysis by hand and published what they found. That is the process working at roughly the speed it is supposed to. The useful part is the specific shape of what broke, because it was not the part anyone was watching.
A-Lab’s loop has four stages: pick a target, plan a route, run the furnace, identify the product. Almost all of the attention in 2023 went to stages two and three, because those are the ones that look like robotics. The part that failed was stage four.
Powder X-ray diffraction gives you the positions and intensities of Bragg peaks. From those you can usually determine a lattice. What you often cannot determine is which atoms sit on which sites, and that is exactly the question that decides whether you have made a new compound. Consider a predicted ordered quaternary oxide in which two metals occupy alternating sites in a regular pattern. Now consider the corresponding disordered solid solution, in which the same two metals are distributed randomly over the same sites in the same average ratio. Those two materials are different — different entropy, often different properties, and one of them is probably already in the literature. Their powder patterns can be very close, because X-rays scatter in proportion to electron count, and neighbours on the periodic table have similar electron counts. The superstructure reflections that distinguish order from disorder are frequently weak, and sometimes systematically absent.
This is the objection that Leeman and colleagues raised in PRX Energy in early 2024, with Robert Palgrave at UCL and Leslie Schoop at Princeton among the authors. They argued that around two-thirds of A-Lab’s claimed successes were “likely to be known compositionally disordered versions of the predicted ordered compounds,” and that automated Rietveld refinement of the diffraction data “is not yet reliable.” Their conclusion at the time was blunt: that no new materials had been discovered in the work.
Their breakdown is more useful than the headline. Sorting the reported successes by failure mode, they found no evidence for the predicted cation ordering in 24 of 36; Rietveld fits that missed major diffraction peaks in 18 of 36; a refinement that switched to a disordered structure file, contradicting the ordered prediction it was supposed to confirm, in 8 of 36; and existing compounds reported as new in 3 of 36. Mg3NiO4 is the clean illustration: predicted as an ordered cubic phase, but A-Lab’s own refinement used a disordered rock-salt model that corresponds to MgNiO2, a solid solution that has been in the literature for decades. What makes this one sting is that the diffraction data could have settled it. The ordered and disordered forms differ by a weak superstructure reflection near 21.1 degrees, and that peak is absent from A-Lab’s own measured pattern. Leeman and colleagues found the automated refinement had swapped in the disordered structure file and returned a fit that missed the peak entirely. The measurement that answers the question had already been taken. Nothing in the pipeline looked at it.
The 2026 correction does not concede that. It concedes something narrower and, I think, more interesting: the phase identification was right 36 times out of 40, and the novelty framing was wrong. Those are different failures. The machine was mostly correct about what it had made. It was the claim that what it made had never been made before that did not survive contact with people who knew the older literature.
A second issue sits further upstream, in where the targets come from. Most of the candidates in this line of work are screened by a distance from the convex hull of formation energies — the surface of thermodynamically stable compositions — computed with density functional theory. DeepMind’s GNoME, published in Nature the same day as A-Lab, works this way: it reports over 2.2 million structures stable relative to prior work, of which 381,000 landed on an updated convex hull.
Machine-learned interatomic potentials are now extremely good at reproducing those DFT energies. On the Matbench Discovery leaderboard, checked on 3 September 2026, the top model reports a formation-energy MAE of about 0.018 eV/atom against DFT; MACE-MPA-0 is at 0.028, CHGNet at 0.063, M3GNet at 0.075.
Now the awkward comparison. Kirklin and colleagues, building the OQMD in 2015, measured DFT-PBE formation energies against 1,670 experimental values. Raw PBE came in at an MAE of 0.136 eV/atom; after fitting the elemental reference energies, 0.096. Take the fitted number, which is the generous one. Then, in the same paper, they compared two curated experimental databases against each other on 75 intermetallics and found the experiments disagreed by 0.082 eV/atom. Their own conclusion follows: it is impossible to assign all of the DFT-experiment gap to DFT. That is a fair defence of the method. It is also a reminder that “ground truth” here has a floor of roughly 0.08 eV/atom before anyone runs a model at all.
I do not read this as an argument that the potentials are useless. They are doing the job they were built for, which is to reproduce DFT at a fraction of the cost, and they do it well enough that the screening step is now essentially free. But it does mean that improvements from 0.03 to 0.02 eV/atom against DFT are improvements in fidelity to a simulation, and the simulation’s own relationship to a crucible is the loosest number in the picture. This is the same trap I wrote about with AGI benchmarks: the metric is real and it is measuring something narrower than the claim it gets used to carry.
Daniel Widdowson and Vitaliy Kurlin, working on a provably complete geometric invariant for periodic crystals, ran a near-duplicate detection pass across roughly 1.85 million structures from the five largest crystal databases. The work was published in SIAM Journal on Applied Mathematics in 2026. The motivation is general data integrity across crystal databases; the GNoME and A-Lab structures are one of the datasets they run the detector against.
Their result: the Cambridge Structural Database, sixty years old and heavily curated, contains about 0.9% near-duplicates. The Inorganic Crystal Structure Database — the reference corpus for essentially all of this work, the thing you check a new compound against to see whether it is new — contains 51,085 entries, just over 30% of the database, that have a near-duplicate elsewhere within the ICSD itself, matchable to within a 0.01 Å average atomic displacement.
So every “is this compound new?” question in this field is a lookup against a corpus that cannot reliably tell whether it already contains a given structure twice. Some of that is benign — the same compound measured at several temperatures, redeterminations at better resolution. But it means the novelty test that A-Lab’s correction walked back is not a test that has a crisp answer available, for anyone, by any method, today. That is a data-hygiene problem underneath a modelling problem, and it rhymes with what happens when models get trained on their own outputs: nothing announces the error, it just becomes the reference.
None of this means computational screening does not find things. Microsoft’s MatterGen, published in Nature in January 2025, generated a structure conditioned on a target bulk modulus of 200 GPa; collaborators at the Shenzhen Institutes of Advanced Technology synthesized it as TaCr2O6 and measured 169 GPa. Microsoft reports that number as within 20% of the design target, which it is, and I would count it — property-conditioned generation producing something that came out of a furnace with roughly the property asked for.
The synthesized TaCr2O6 showed compositional disorder between tantalum and chromium, rather than the ordered structure the model generated. Same gap as A-Lab’s, disclosed by the authors this time.
And the novelty question came back here too. A 2026 critique in Materials Horizons argues the phase that was actually synthesized is structurally identical to a compound reported in 1972, and that it sits in MatterGen’s own training data — the same problem that struck Zn2Cr3FeO8 from A-Lab’s list. I have not seen a reply to that critique.
Notice what is and is not in dispute. Nobody is arguing about 169 GPa. The measurement was taken, and it stands. What is contested, again, is whether the thing measured was new. That is the distinction that survives all of this: a property you measured holds up under scrutiny, and novelty depends on a literature search against a database that is 30% duplicated.
The older and more complete example is thermoelectrics. Zhu and colleagues screened the half-Heusler space for thermodynamic stability, predicted six new stable compounds, synthesized TaFeSb, and measured a peak ZT of about 1.52 at 973 K with 11.4% heat-to-electricity conversion — among the best reported for a p-type half-Heusler (Nature Communications, 2019). Then they did the characterization: XRD, selected-area electron diffraction, and scanning transmission electron microscopy, together confirming the ordered structure rather than a disordered variant. Seven years before any of the present argument, by people who knew the question was coming.
Separately, a 2026 experimental paper reports synthesizing the GNoME-predicted MnFeCo4Si2 and measuring a Curie temperature of 1039 K in a single-phase rhombohedral structure consistent with the prediction (arXiv:2603.22748). One compound, from a list of 381,000. Confirmations are slow, and that is not a criticism. But keep the ratio in mind next to the headline counts.
The framing that got attached to these results in late 2023 — an order-of- magnitude expansion in known materials, a robot doing chemistry unattended — put the emphasis on generation and automation. Twenty-six months of scrutiny later, both of those held up better than the part nobody was looking at. A-Lab really did run 355 experiments in 17 days, and its phase calls were right 36 times in 40. GNoME really does predict stable structures at high hit rates. What did not hold was the epistemics around the output: what counts as new, and how you would know.
The generation step is now cheap and fairly trustworthy. The synthesis step is expensive and works. The characterization step is the bottleneck, and not for the reason I expected before reading the critiques. Powder diffraction is the one technique in the chain that automates cleanly. Sometimes it cannot resolve order from disorder at all, and then you need neutron diffraction, electron diffraction, solid-state NMR, or pair distribution function analysis, all of which run one to two orders of magnitude slower and need a person. But the Mg3NiO4 case is the more uncomfortable one: the answer was sitting in the pattern, and the automated refinement walked past it. What does not automate is the judgment about whether a fit is physically plausible, which no goodness-of-fit statistic supplies. Leeman and colleagues put the consequence plainly: if the rate-determining step is human analysis, the robotic lab may not move faster than a conventional one.
So if I were funding work here I would rather pay for beamtime, for electron microscopy, and for deduplicating the reference databases than for another order of magnitude of candidate structures. There are already 381,000 waiting, and we cannot currently tell which of them we have already made.
References