Why this page reads differently
Most of this manual is built around uncertainty, danger, and honest gaps in what can be verified. This page is different: it collects AI-driven scientific results that are dated, checkable, and in one case already carry a Nobel Prize — yet the same research program that produced that Nobel-winning method also produced a formally corrected Nature paper and a diagnostic tool that worked in the lab and stumbled in a real clinic. Holding both facts at once — genuine achievement and genuine overclaiming, from the same lab, sometimes the same year — is the point of this page, not a contradiction to resolve.
Everything below concerns narrow, task-specific machine learning applied to a scientific or medical problem, not AGI. That some future general system will accelerate all of science at once is a scenario, addressed on AI benefits and the opportunity costs of delay; this page stays with what has shipped, published, or been peer-reviewed.
Protein structures, solved: AlphaFold and the 2024 Nobel Prize
The clearest case is protein structure prediction. In October 2024 the Royal Swedish Academy of Sciences awarded the Nobel Prize in Chemistry jointly to Demis Hassabis and John Jumper of Google DeepMind, “for protein structure prediction,” and to David Baker of the University of Washington, for computational protein design (NobelPrize.org, October 9, 2024) — the field’s highest honor, four years after AlphaFold2’s original 2020 results and confirmed by a discipline that spent those years actually using the tool: reportedly more than two million researchers in 190 countries, predicting structures for virtually all 200 million known proteins at accuracy approaching experimental methods on hard targets (Nature News, October 9, 2024).
From structure to drug: Isomorphic Labs, still in the pipeline
DeepMind’s protein-folding work spun out Isomorphic Labs, built to carry structure prediction into drug discovery. In January 2024 it signed deals with Eli Lilly and Novartis worth a combined roughly $3 billion in potential milestones, aimed at “undruggable” small-molecule targets (TechCrunch, January 7, 2024). By early 2026, secondary reporting described the collaboration as having produced multiple preclinical candidates — still years from a clinical trial, let alone an approved medicine. No drug from this partnership has entered human trials as of this writing.
AlphaFold accelerated one step in a long pipeline; it has not been shown to accelerate the pipeline’s output, measured in approved therapies rather than milestone payments. AI benefits and the opportunity costs of delay draws the line between a measured benefit and a projected one — this is squarely a projected one.
Materials discovery — and a correction that proves the system works
The most instructive case here is materials science, precisely because it did not stay a clean success story. In November 2023 DeepMind published GNoME in Nature: a system that predicted 2.2 million candidate crystal structures, of which about 381,000 were predicted stable — roughly 45 times more stable-but-unrealized compounds than the entire prior historical record (Merchant et al., Nature, November 29, 2023). A companion paper described an autonomous robotic laboratory, the “A-Lab,” that used these predictions to synthesize dozens of new materials without human intervention, in seventeen days.
In 2024, three outside chemists — Robert Palgrave (UCL), Leslie Schoop (Princeton), and Susan Latturner (Florida State) — found that roughly two-thirds of the A-Lab’s “new” compounds were already-known structures, misidentified because its X-ray diffraction analysis used novice-level pattern matching rather than the careful Rietveld refinement a trained crystallographer would apply; many were already in the Inorganic Crystal Structure Database. On January 19, 2026, Nature published a formal author correction acknowledging the novelty claims had been “subject to misinterpretation” — new to the platform’s own catalog, not necessarily new to science.
This is not a story about AI failing. GNoME’s structure predictions were not the part that turned out wrong; the overclaim was in how “new” got defined downstream — and it was caught and formally corrected through ordinary peer review, two years later. That is the system working as intended: a Nature publication is strong evidence, but not immune to correction, and a claim’s real status is what survives scrutiny, not what the press release said the week it launched.
Mathematics: AlphaProof at the Olympiad, off the clock
In July 2024 DeepMind announced that AlphaProof and AlphaGeometry 2 had solved four of six problems at that year’s International Mathematical Olympiad, for a score of 28 out of 42 — silver-medal level, the first time any AI system had reached it (DeepMind, July 25, 2024). DeepMind itself disclosed that one problem was solved within minutes but others took up to three days of computation, worth noting since human competitors solve the same six problems across two 4.5-hour sessions. This was not entered or judged as an actual IMO submission and has not been independently peer-reviewed as a competition result; it is a company-run benchmark, honestly captioned, not a claim that AI competes with Olympiad mathematicians under the constraints they actually face.
Weather forecasting: a case peer review already settled
Weather forecasting is a useful contrast because ground truth arrives daily and cannot be gamed retroactively. DeepMind’s GraphCast, published in Science in November 2023, outperformed the European weather agency’s operational HRES model on the large majority of thousands of evaluated variable-and-lead-time combinations, and produces a ten-day global forecast in under a minute on a single TPU (DeepMind, November 14, 2023). Forecasts are scored against actual weather every day, so this is one of the harder domains to overclaim in, and years on, the result has held up.
Medicine, at regulatory scale: FDA clearances multiply
Clinical AI has crossed from novelty to routine regulatory business. In April 2018, IDx-DR (now LumineticsCore) became the first FDA-authorized fully autonomous AI diagnostic system in medicine, cleared to detect diabetic retinopathy without physician interpretation, on a 900-subject trial showing 87.2% sensitivity and 90.7% specificity against expert grading (npj Digital Medicine, 2018). Such clearances are now routine: FDA authorizations of AI-enabled devices rose from roughly 758 by end-2024 to 1,039 by December 2025, pace climbing from about 21 to about 30 a month, led by GE HealthCare, Siemens Healthineers, Philips, and Canon (The Imaging Wire, March 11, 2026). A clearance is a real, checkable fact — not the same claim as “this tool works well in your hospital,” which is where the next section picks up.
The lab-to-clinic gap: what happens when the model meets a real patient
The clearest documented case of the gap between validated accuracy and deployed reality is Google Health’s diabetic-retinopathy system, which exceeded 90% sensitivity and specificity in trials. Deployed across 11 clinics in Thailand between November 2018 and August 2019, it rejected roughly 21% of the roughly 1,940 images nurses submitted as “ungradable” — mostly lighting conditions a human reader would have accepted without a second thought — disrupting workflow and lengthening visits for patients who had often traveled hours to be screened. The finding was peer-reviewed at CHI 2020, a leading human-computer-interaction venue, and reported at the time by MIT Technology Review (MIT Technology Review, April 27, 2020). Later reviews of radiology AI describe this as a recurring pattern: real-world performance depends on local equipment and workflow, not a model’s accuracy figure alone.
A separate FDA-cleared stroke tool, Viz.ai, shows the split differently: a systematic review and meta-analysis pooling roughly a dozen studies and 15,000-plus patients sped up hospital workflow without a statistically significant improvement in the clinical outcomes those workflows exist to improve. Faster is not automatically better.
What this pattern says about the rest of the site
Every case above was produced by the same small set of frontier labs this manual discusses elsewhere — chiefly Google DeepMind — whose research culture is covered on how the frontier AI labs differ, and whose concentration of scientific capability in a handful of organizations is its own subject at power concentration: why a few labs matter. That concentration cuts both ways: it is why one program can accumulate a Nobel Prize, a silver-medal Olympiad run, and a corrected Nature paper within two years, and why the framing of these claims runs first through a lab’s own publicity and only later through peer review that does not answer to it.
None of this argues for dismissing the results. A Nobel Prize is not hype; a ten-day forecast produced in under a minute is not hype. But “AI is accelerating science” is true the way most durable claims about this technology are true: unevenly, with the clearest successes in problems that have clean, checkable ground truth — a folded protein, tomorrow’s weather, a crystal’s measured stability — and a persistent gap wherever the last mile runs through a messy clinic or an unchecked press release. The discipline that caught the GNoME error, two years later, is the discipline this page asks readers to apply everywhere else on this site.