1 · Concept overview

A machine that does science has to close a loop: form a hypothesis, design an experiment that could refute it, run that experiment against the physical world, read the outcome, and revise. Every stage of that loop has now been automated somewhere, and the whole loop has been closed end to end in yeast functional genomics, in organic synthesis, in inorganic materials synthesis, and in combinatorial mathematics. The engineering is real. What is not settled — and this brief argues it is not close to settled — is who or what decides whether the output is true.

That question turns out to be the whole subject. Sort the field’s results by whether they survived independent scrutiny and a single variable separates the two piles. Where the system paired a generator with a cheap, exact, automatic verifier — a program that runs, a proof that checks, a matrix identity that either holds or does not — the results held up and in a few cases improved on published human bounds. Where the verifier was a noisy physical instrument or another language model, the results did not hold up. The clearest case in the literature is an autonomous materials laboratory whose headline output was re-examined by an independent group and found to contain no new materials at all, for reasons that are entirely about measurement and not at all about reasoning.

Three findings carry what follows. The founding robot-scientist paper of 2004 does not contain the discovery claim it is universally cited for. The best-known autonomous-materials result exists in three mutually inconsistent versions ending in a refutation. And the most prominent industrial “AI co-scientist” changed both its title and the set of wet-lab validations its abstract carries between versions, so that two briefs written from two versions would describe two different systems. This brief names the version it used throughout, because in this field that is not pedantry.

2 · Current scientific position

Established The founding paper of the field claims efficiency, not discovery. King, Whelan, Jones, Reiser, Bryant, Muggleton, Kell and Oliver, “Functional genomic hypothesis generation and experimentation by a robot scientist,” Nature 427:247–252 (2004), describes the system later known as Adam. Established The abstract sets out a genuine closed loop: the system “automatically originates hypotheses to explain observations, devises experiments to test these hypotheses, physically runs the experiments using a laboratory robot, interprets the results to falsify hypotheses inconsistent with the data, and then repeats the cycle,” applied to aromatic amino acid synthesis in yeast. Established The result it claims is that “an intelligent experiment selection strategy is competitive with human performance and significantly outperforms, with a cost decrease of 3-fold and 100-fold (respectively), both cheapest and random-experiment selection.” Established That is a claim about the economics of experimental design, not a claim to have discovered anything.

Established The discovery claim attached to that paper belongs somewhere else. The English Wikipedia article on the robot scientist, fetched on 2 September 2026, states that Adam “became the first machine in history to have discovered new scientific knowledge independently of its human creators,” referenced to King et al. (2004). Established No such claim appears in that abstract. Frontier The likely correct provenance is King and colleagues’ later 2009 Science paper on the automation of science, which this brief could not obtain — the publisher domain refused to serve during the research pass — and which is therefore cited here as a lead rather than as a source for any specific claim. Established The correction matters beyond bookkeeping: the field’s founding date for machine discovery is five years earlier than its own evidence supports.

Frontier The robot scientists’ most valuable published result is a null result about human science. Roper and colleagues, “Testing the reproducibility and robustness of the cancer biology literature by robot,” J. R. Soc. Interface 19(189):20210821 (2022), used the successor system Eve for semi-automated reproducibility testing of published cancer biology. Frontier The figure widely quoted is that under a third of tested results reproduced; this brief reached that figure only through a secondary source, could not obtain the paper itself, and therefore does not state it as established. Speculative The shape of the finding matters more than its magnitude: the most useful thing a laboratory robot has done to date was to audit the literature rather than add to it, and that asymmetry recurs at every level of this subject.

Established Language-model agents can now run real chemistry, and none of the chemistry is new. Boiko, MacKnight, Kline and Gomes, “Autonomous chemical research with large language models,” Nature 624:570–578 (2023), describes Coscientist, a GPT-4-driven system that “autonomously designs, plans and performs complex experiments” using internet and documentation search, code execution and experimental automation. Established Six demonstrations are reported: synthesis planning for seven compounds, liquid-handler control under UV-Vis readout, execution of Suzuki–Miyaura and Sonogashira couplings with correct product formation, reaction optimization over Suzuki and Buchwald–Hartwig datasets, and HPLC code generation in a cloud laboratory. Established Every one of those reactions was known before the system ran it; the achievement is autonomous execution, not novel chemistry. Established ChemCrow (arXiv:2304.05376) reports the adjacent result from tool augmentation: eighteen expert-designed chemistry tools, with which the system “autonomously planned and executed the syntheses of an insect repellent, three organocatalysts, and guided the discovery of a novel chromophore” — where the verb in the last clause is doing real work.

Established Where an exact verifier exists, machines have produced results that improve on the published human record. FunSearch — Romera-Paredes and colleagues, “Mathematical discoveries from program search with large language models,” Nature 625:468–475 (2024) — pairs a pretrained language model with what the paper calls a systematic evaluator, and reports new constructions in the cap set problem: a cap set of size 512 in dimension 8, the cap set capacity lower bound improved from 2.2180 to 2.2202, a full-size admissible set I(12,7) giving capacity ≥ 2.219486, and a partial admissible set in A(24,17) of size 237,984. Established It also produced online bin-packing heuristics beating first-fit and best-fit, within 0.03% of the optimality lower bound on 100,000-item Weibull instances. Established AlphaGeometry (Nature 625:476–482, 2024) solved 25 of 30 problems on the IMO-AG-30 olympiad geometry set against a previous best of 10, using a language model trained from scratch on synthetic theorems to guide a symbolic deduction engine; AlphaTensor (Nature 610:47–53, 2022) found a 47-multiplication algorithm for 4×4 matrices in finite fields, improving Strassen’s two-level algorithm for the first time in fifty years, and a rank-76 decomposition for 4×5×5 against a previous state of the art of 80. Frontier AlphaEvolve (arXiv:2506.13131, 2025) multiplies 4×4 complex-valued matrices in 48 scalar multiplications, described by its authors as the first improvement on Strassen in that setting after 56 years. Established The structural point running through all four is the one the field under-states: the language model proposes, and something that is not a language model decides.

Established Remove the exact verifier and the record changes character completely. The clearest documented case is autonomous inorganic materials synthesis. Established GNoME (Merchant and colleagues, Nature 624:80–85, 2023) reports three distinct numbers that coverage routinely collapses into one: 2.2 million structures below the current convex hull, of which 381,000 are described as newly discovered stable materials, of which 736 “have already been independently experimentally realized.” Established Its companion paper, Szymanski and colleagues’ A-Lab (Nature 624:86–91, 2023), states that “over 17 days of continuous operation, the A-Lab realized 36 compounds from a set of 57 targets,” from recipes proposed by language models trained on the literature and refined by thermodynamically grounded active learning.

Established An independent re-analysis of that work concluded that none of it was new. Leeman, Liu, Stiles, Lee, Bhatt, Schoop and Palgrave, “Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis,” PRX Energy 3:011002 (2024), describes the same work as having reported “the autonomous discovery of 43 novel materials,” states “we discuss all 43 synthetic products,” identifies four shortfalls in the analysis, and concludes: “These errors unfortunately lead to the conclusion that no new materials have been discovered in that work.” Established Two root causes are named and both are instrumental: “Automated Rietveld analysis of powder x-ray diffraction data is not yet reliable,” and unmodelled compositional disorder, with the consequence that “two thirds of the claimed successful materials in Szymanski et al. are likely to be known compositionally disordered versions of the predicted ordered compounds.” Frontier Note that the count itself does not reconcile: the Nature abstract says 36 of 57 targets, the critique discusses 43 products, and neither abstract explains the gap, so no single figure for this experiment should be treated as settled until both full texts are read side by side. Frontier A companion critique of GNoME by Cheetham and Seshadri exists in Chemistry of Materials 36:3490–3495 (2024); its record was verified but the publisher gated the text, so no number from it appears here.

Established This is the brief’s thesis, and it is worth stating flatly: the failure was in measurement, not in reasoning. Nothing in Leeman et al.’s account faults the hypothesis generation, the recipe proposal, the active learning or the robotics. Established The system planned syntheses, ran them and produced powders; what it could not do was determine reliably what it had made. Frontier Generalised, that is the position this brief defends: for the large majority of science the automatable step is the experiment and the unautomated step is the adjudication of what the experiment showed, and no amount of improvement in the reasoning layer touches that constraint.

Established The paper-writing systems are best understood through the scope conditions their own abstracts carry. The AI Scientist (arXiv:2408.06292, Sakana AI with Oxford and UBC) generates ideas, writes code, runs experiments, writes the paper, then reviews its own output with an automated reviewer the same team built. Established Its cost claim is precise: “each idea is implemented and developed into a full paper at a cost of less than $15 per paper.” Established Its quality claim is circular in the abstract’s own words — papers that “exceed the acceptance threshold at a top machine learning conference as judged by our automated reviewer.” Established The AI Scientist-v2 (arXiv:2504.08066, 2025) drops the human code templates, adds agentic tree search, and reports that of three manuscripts submitted to an ICLR workshop, one exceeded the average human acceptance threshold — “the first instance of a fully AI-generated paper successfully navigating a peer review.” Frontier Three scope conditions must travel with that sentence or it becomes false: a workshop rather than the main conference, one submission of three, and a score comparison rather than an acceptance decision standing in a published record.

Frontier Google’s co-scientist is the field’s cleanest case of version drift, and this brief used version 1. Gottweis and colleagues, arXiv:2502.18864v1 (26 February 2025), titled “Towards an AI co-scientist,” describes a Gemini-based multi-agent system with an asynchronous task-execution framework and a tournament evolution process for self-improving hypothesis generation. Frontier That version’s abstract carries three validations: drug-repurposing candidates for acute myeloid leukemia showing tumour inhibition in vitro at clinically applicable concentrations; new epigenetic targets for liver fibrosis “validated by anti-fibrotic activity and liver cell regeneration in human hepatic organoids”; and recapitulation of “unpublished experimental results via a parallel in silico discovery of a novel gene transfer mechanism in bacterial evolution.” Established A later version is retitled “Accelerating scientific discovery with Co-Scientist,” and its abstract retains only the AML repurposing and combination-therapy result: the liver-fibrosis organoid validation and the bacterial gene-transfer recapitulation are gone from it. Frontier The title moved from a hedge to an achievement while the abstract’s evidence base narrowed, and both facts belong with any citation of this system. Established On the third validation as originally stated, recapitulating a result the collaborating human team already possessed is not an independent discovery, and the v1 abstract does not claim it is.

Established Independently constructed benchmarks agree on a ceiling of roughly twenty to thirty-five percent, and that agreement is the most robust quantitative fact in the subject. DiscoveryBench, 264 real tasks from six domains plus 903 synthetic, puts the best system at 25%; CORE-Bench, 270 tasks drawn from 90 papers, puts the best agent at 21% on its hardest tier; ScienceAgentBench, 102 tasks from 44 peer-reviewed papers validated by nine experts, reports 32.4% independently, 34.3% with expert-supplied knowledge and 42.2% for o1-preview at substantially higher cost. Established MLE-bench, across 75 Kaggle competitions, records a bronze-medal-equivalent result in 16.9%; BixBench, on 50-plus real bioinformatics scenarios, records around 17% on open-answer questions and no better than random on multiple choice; PaperBench, replicating 20 ICML 2024 Spotlight and Oral papers across 8,316 gradable tasks, records 21.0% for the best model tested, with machine-learning PhDs still ahead. Frontier Six benchmarks built by different groups for different fields converge on the same band, which is a stronger signal than any one of them, and it is a band nobody designed for.

Established A frontier industrial lab has published the negative result explicitly. MLGym (Nathani and colleagues at Meta; the identifier returned during research was arXiv:2502.16111 and should be verified before citation) puts frontier models on thirteen open-ended AI research tasks across computer vision, natural language processing, reinforcement learning and game theory. Established The finding, verbatim: models “can improve on given baselines” but “do not generate novel hypotheses, algorithms, architectures, or substantial improvements.” Established Against that sits the most-quoted positive result in the field, Si, Yang and Hashimoto’s blinded study with over a hundred NLP researchers (arXiv:2409.04109): “LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility.” Frontier The same abstract reports failures of language-model self-evaluation, a lack of diversity in generation, and that human novelty judgement is itself difficult even for experts — which is why those authors propose a follow-up in which researchers actually execute the ideas. Frontier Until that reports, the result is a fact about a rating scale, not about the world.

3 · Frontier questions

Frontier Kosmos is the largest autonomy claim to date and it arrives with an accuracy audit attached, which is progress. Mitchener and colleagues, “Kosmos: An AI Scientist for Autonomous Discovery” (arXiv:2511.02824, November 2025, FutureHouse and collaborators, 38 authors), reports a structured world model maintaining coherence across more than 200 agent actions, roughly 42,000 lines of code executed and around 1,500 papers read per run. Frontier It claims seven discoveries across metabolomics, materials science, neuroscience and statistical genetics, of which three independently reproduced findings from preprinted or unpublished manuscripts not accessed during the run, and four are described as novel contributions. Frontier Independent evaluators judged 79.4% of statements in Kosmos reports to be accurate, and collaborators equated a single 20-cycle run to roughly six months of their own research time. Established Read the audit the other way: approximately one statement in five in an autonomously produced scientific report was judged inaccurate. Speculative That is a remarkable number to publish and a disqualifying number to build on, and both readings are correct at once.

Frontier Two systems have taken a machine-generated hypothesis into a wet laboratory and out again. Robin (Ghareeb and colleagues, arXiv:2505.13400, FutureHouse) claims to be “the first multi-agent system capable of fully automating the key intellectual steps of the scientific process,” and its concrete output is ripasudil — a clinically used ROCK inhibitor never previously proposed for dry age-related macular degeneration — hypothesised through enhanced retinal-pigment-epithelium phagocytosis, with follow-up RNA sequencing showing upregulation of ABCA1. Frontier The Virtual Lab (Swanson, Wu, Bulaong, Pak and Zou, Nature 646:716–723, 2025) puts a language-model principal-investigator agent in charge of language-model scientist agents running research meetings, “with a human researcher providing high-level feedback.” Frontier The agents assembled a pipeline of ESM, AlphaFold-Multimer and Rosetta that designed 92 new nanobodies; experimental validation found a range of functional nanobodies, with two showing improved binding to the JN.1 or KP.3 SARS-CoV-2 variants while retaining ancestral-spike binding. Frontier As of this writing that is the strongest published instance of an AI-agent system producing a wet-lab-validated novel molecule. Frontier Both results come from the laboratories that built the systems, and in both the human contribution is described qualitatively and quantified nowhere.

Frontier The economics are moving faster than the capability. Agent Laboratory (arXiv:2501.04227) reports an 84% decrease in research expenses relative to prior autonomous methods across a three-phase pipeline of literature review, experimentation and report writing — and reports, as its own headline finding, that “human involvement, providing feedback at each stage, significantly improves the overall quality of research.” Frontier Aviary (arXiv:2412.21154) claims that agents built on open-source, non-frontier language models can match and exceed both frontier agents and human experts on several scientific tasks at up to 100× lower inference cost, across five language-grounded environments including molecular cloning, literature question answering and protein stability engineering. Speculative If that holds, the constraint on running these loops stops being money within a few years, which moves the entire bottleneck to physical experiment throughput and to verification.

Speculative A small strand of work targets the generation of novelty rather than the solution of posed problems, and it is the part of the field closest to the exotic question. Automated Search for Artificial Life (arXiv:2412.17799) uses foundation models to search substrate spaces including Lenia and Boids for lifelike and open-ended dynamics. Speculative It is not doing science, but it is the only line of work in this pack that treats open-endedness itself as the objective rather than as a hoped-for side effect of scale. Handwave If a machine ever produces a result nobody prompted it toward, the lineage will plausibly run through open-endedness research rather than through agent scaffolding, because agent scaffolding is defined by the goal it is handed.

Frontier The open question the field has not framed as a question: nothing here automates refutation. Every system surveyed generates and tests. Established None takes a published claim and designs the experiment that would falsify it. Frontier That is the task Eve was pointed at, it is the task with the clearest social value given the state of the biomedical literature, and it is the only task in this subject where a machine’s known weaknesses — tirelessness, indifference to reputation, no stake in the outcome — are advantages rather than liabilities.

4 · Technological bottlenecks

Established The workback target is a system that chooses its own question, produces a result, and has that result treated as a contribution by a field that did not commission it. The chain below is ordered, and the binding link is the third.

Established L0, reproduction. Saturate CORE-Bench, currently 21% on its hardest tier, and PaperBench, currently 21.0%, on the principle that an agent which cannot reproduce a known result has no standing to claim a new one. Established L1, verified execution in a closed formal domain: done. FunSearch, AlphaTensor, AlphaGeometry and AlphaEvolve all clear this bar, and all four require a cheap exact verifier to do it. Frontier L2, verified execution in an open wet domain against a human-specified target: done once, with a human in the loop. The Virtual Lab nanobodies, and arguably Robin’s ripasudil. Frontier The outstanding measurement at this link is not another demonstration; it is the fraction of the intellectual work that was human, currently reported only as “high-level feedback.”

Frontier L3 is the binding link: calibrated open-ended verification. The requirement is an instrument that decides whether a non-formal claim is true, at expert reliability, cheaply. Established This is exactly where A-Lab failed: automated Rietveld refinement judged its products as new, and hand re-analysis by a different group concluded that none were. Frontier Every link downstream depends on L3, and no result in this brief clears it. Established Kosmos’s 79.4% statement accuracy is the closest existing measurement and it is a single self-reported point with no specificity attached. Frontier No receiver-operating-characteristic curve exists for any AI verifier in this literature, which is a striking absence for a field this well funded.

Speculative L4, autonomous problem selection under blinded evaluation; L5, field uptake as the criterion; L6, automated refutation. L4 has had its ideation half run once, by Si et al.; the execution half is proposed and unreported. Speculative L5 has not been attempted for any AI-originated result, including FunSearch’s cap sets, which are the easiest case because the constructions were released. Speculative L6 has not been attempted at all.

Frontier The concrete, fundable, currently unsolved instrument problem sits inside L3 and its authors have already named it. Leeman et al. explicitly request “development of a reliable artificial-intelligence-based tool for Rietveld fitting.” Established That is a specific piece of software, in a specific measurement modality, whose absence demonstrably invalidated a Nature paper’s discovery claim. Speculative It costs a fraction of what a single agent-scaffolding effort costs and nobody has built it.

5 · Research dependencies

Established This brief depends on Artificial General Intelligence for the question of what the reasoning layer can be expected to do. If the generality claims made for frontier models are right, the hypothesis-generation half of this subject is close to solved and the whole weight falls on verification. Frontier If they are wrong in the way the sceptics argue — retrieval and recombination over a training corpus rather than novel inference — then a second, harder problem sits under the first, and no verifier fixes it.

Frontier It depends second on Intelligence Measurement, and the dependency is unusually literal. The question of whether a machine-generated hypothesis was derived or retrieved is the same question as benchmark contamination, wearing a different hat: in both cases what must be shown is that the supporting evidence was absent from the training corpus. Frontier The instruments that would settle contamination would also settle novelty, and the field is currently building neither.

Established Three non-brief dependencies are load-bearing. Automated structure determination — Rietveld refinement above all, but also automated spectroscopy interpretation — must reach expert reliability before autonomous synthesis claims mean anything. Established Cheap, high-throughput physical experimentation must keep falling in cost, because the loop runs at the speed of its slowest physical step. Frontier And machine-readable access to the full published literature, including negative results, corrections and retractions, is a precondition for any system whose method is recombination over the corpus. Speculative That third dependency is contractual and legal rather than technical, which makes it the one most likely to bind unexpectedly: a system that reads 1,500 papers per run is reading whatever its licences permit, and the retraction notices are frequently on the far side of a paywall from the papers they retract.

6 · Required experiments

Established Experiment one, and the one that unblocks the field: an ROC curve for an automated verifier. Take a corpus of roughly 200 published claims with known replication status — the large reproducibility-project corpora, or the cancer-biology set Eve was run against — and measure an automated verifier’s sensitivity and specificity against ground truth. Established Report the full curve, not a single accuracy point. Frontier Nothing of this kind exists for any system in this literature, and until it does, every autonomous-discovery claim rests on an uncharacterised instrument. Speculative The expected result is that verifiers are far more sensitive than specific — good at endorsing true claims and bad at rejecting plausible false ones — which is precisely the failure profile that produced 43 confidently certified non-discoveries.

Speculative Experiment two: reliable automated Rietveld fitting, benchmarked against expert crystallographers on a blinded set. Include compositionally disordered phases deliberately, since unmodelled disorder was one of the two named causes of the A-Lab failure. Speculative Success criterion: agreement with expert consensus on phase identity at a rate the experts themselves achieve with each other.

Speculative Experiment three: the execution half of the Si et al. design. An agent proposes N research questions in a field; a blinded panel of that field’s researchers rates them against N human-proposed questions; the top-rated from both arms are executed by independent laboratories that do not know the provenance. Speculative The measured outcome is not rated novelty but research yield, which is the thing rated novelty is supposed to predict and has never been shown to.

Speculative Experiment four: quantify the human share. Re-run a Virtual Lab-style loop with the human feedback logged, timestamped and independently classified as directive, corrective or absent. Frontier The claim under test is whether the human contribution is high-level steering, as reported, or hypothesis selection under another name.

Speculative Experiment five: uptake at 24 months. Track citations, replication attempts and derivative work for AI-originated results, with a matched human-originated control set. Speculative Experiment six: an automated refuter. Given a published claim, design and run the experiment that would falsify it. Handwave That last is the missing symmetric half of the entire enterprise and no group has announced an attempt.

7 · Engineering requirements

Established The hardware end is the mature end. Robotic liquid handlers, automated powder synthesis lines, cloud laboratories with programmable HPLC, and the plate-reader and diffractometer stack behind them are commercial products, and Coscientist and A-Lab both demonstrate that a language model can drive them through an API without a human in the room, the latter for 17 days continuously. Frontier Throughput is no longer the binding constraint in inorganic synthesis; characterisation is.

Frontier The software requirement that matters most is unglamorous. Automated Rietveld refinement, automated phase identification under disorder, automated spectral assignment — the instrument-interpretation layer — is what converts a physical outcome into a proposition a reasoning system can act on. Established It is currently the weakest link in the only fully autonomous discovery pipeline that has been independently audited. Speculative An engineering programme aimed squarely at that layer would do more for autonomous discovery than another generation of agent orchestration.

Frontier Inference economics are already tractable and getting better. The AI Scientist reports under $15 per generated paper; Agent Laboratory reports an 84% cost reduction over prior autonomous methods; Aviary reports open non-frontier models matching frontier agents at up to 100× lower inference cost. Speculative Taken together these imply that within a few years the marginal cost of a machine-generated, machine-executed, machine-written study approaches the cost of the physical reagents. Frontier That is an engineering success with an unattractive corollary, treated in section 10.

Established One piece of measurement hygiene should be a build requirement rather than a virtue. Evaluations in this field are routinely reported as bare percentages on small task sets; error bars on evaluation results are computable and are usually omitted. Frontier A discovery-agent scoreboard without confidence intervals cannot distinguish a real improvement from resampling noise, and most of the numbers in section 2 are quoted from abstracts that give none.

8 · Adjacent technologies

Established Structure prediction and protein design are the adjacent technologies doing the most work here. The Virtual Lab’s nanobody result is a pipeline of ESM, AlphaFold-Multimer and Rosetta assembled by agents; the agents contributed orchestration and selection, and the predictive power came from tools built by people. Frontier That pattern — agent as conductor, specialised model as instrument — is the one that has actually produced validated molecules, and it is a weaker claim than autonomous discovery.

Frontier Automated theorem proving is the domain where the verification problem is already solved, which is why it is both the field’s best evidence and its least generalisable. A proof checked by a proof assistant is true in a way no wet-lab result ever is, and AlphaGeometry’s 25 of 30 on IMO-AG-30 is a real result precisely because the symbolic deduction engine, not the language model, certifies each step. Frontier This brief flags a limit on its own coverage: the proof-assistant literature proper — Lean, its mathematical library, and the recent competition systems built on it — could not be sourced during the research pass, and nothing about it is asserted here beyond what the AlphaGeometry and FunSearch abstracts support.

Frontier Three further adjacencies are structural rather than incidental. Multi-Agent Intelligence Systems supplies the architecture every recent scientific agent uses, and inherits its unsolved question of whether the ensemble beats its best member under matched compute. Frontier Artificial Creativity supplies the measurement problem: whether machine output is novel, and whether human raters can tell. Established Synthetic Biology supplies both the substrate for the most valuable wet-lab loops and the sharpest dual-use exposure. Speculative The adjacency that would matter most is the one nobody is building: an automated instrument-interpretation layer shared across diffraction, spectroscopy and imaging, standing between the robots and the reasoning systems in the way that a sequence database stands between a sequencer and a biologist. Frontier Its absence is why every autonomous laboratory currently ships its own bespoke, unvalidated adjudication of what its own experiments showed.

9 · Institutional requirements

Established The adjudication of discovery is an institutional function, and no institution has been assigned it here. The A-Lab episode is the worked example: a discovery claim published in Nature, a refutation published in PRX Energy, and a literature in which both now stand. Frontier There is no mechanism that puts the second in front of a reader who arrives at the first. Speculative Whatever else automated discovery requires, it requires a mechanism by which an automated claim can be automatically un-made, and journals do not have one.

Established The propagation of the 2004 attribution shows how long a bad citation survives when nothing is designed to kill it. A claim absent from a paper’s abstract has been attached to that paper in the most-read reference source in the world for years, and is repeated in introductions, funding cases and popular coverage. Frontier The cost of that error is not the error; it is that the field’s founding date for “machine discovery” is five years earlier than the evidence supports.

Frontier Peer review will be passed before it is deserved, and the AI Scientist-v2 workshop result is the first data point. The criterion that would actually settle the question is not acceptance but uptake: AI-generated papers accepted at venues blind to their provenance, at rates matching human submissions, and subsequently cited and built upon. Established The building-upon criterion is the one nobody measures, for machine or human work.

Frontier The asymmetry between production and refutation is the institutional problem in one sentence. A paper generated for under $15 was refuted by seven named researchers hand-examining 43 synthetic products. Speculative No funding line exists for that kind of work, no journal prioritises it, and no career is built on it. Speculative Institutions that want autonomous discovery to be worth having should fund the refutation side first, on the grounds that it is currently the cheaper of the two to make ten times better, and that the value of a discovery claim is bounded above by the credibility of the audit behind it. Frontier A national instrument programme for automated characterisation, jointly owned by crystallographers and machine-learning groups, would be a more useful public investment here than another scientific-agent laboratory. Frontier Scientific Governance Models and AI Governance both bear on who could plausibly do that.

10 · Ethical & societal considerations

Established The cost collapse is the ethical event, not the capability. At under $15 per generated paper, and with roughly one statement in five in an autonomously produced scientific report judged inaccurate by independent evaluators, the marginal cost of a plausible false claim has fallen by orders of magnitude while the cost of refuting one has not moved. Frontier A literature already struggling with reproducibility is about to receive a large volume of superficially competent, cheaply produced, unevenly true material. Speculative The failure mode is not fraud; it is volume.

Frontier Dual use is concrete in this subject rather than hypothetical. Systems that plan and execute syntheses autonomously, and systems that propose molecular, protein and genomic designs, are the same systems whether the target is an organocatalyst or something else, and safety evaluation for scientific agents is an active area with dedicated benchmarks. Frontier The relevant asymmetry is that a chemistry agent connected to a cloud laboratory has a physical effector, which distinguishes it from almost every other class of language-model application.

Speculative Credit and authorship are unresolved in a way that will be decided by practice before it is decided by principle. If a system proposes ripasudil for dry age-related macular degeneration and a human runs the RNA sequencing, the intellectual contribution is genuinely joint and the incentive is to describe it either way depending on audience. Frontier Every wet-lab result in this brief comes from the laboratory that built the system, which is the ordinary situation early in a technology and is also exactly the situation in which self-report is least reliable.

Speculative The distributional question deserves naming. Autonomous laboratories concentrate discovery capacity in the small number of institutions that can afford the robotics and the inference, at the moment when the reasoning layer is becoming cheap. Speculative If Aviary’s 100× cost claim generalises, the reasoning becomes commodity and the physical instruments become the moat, which inverts the usual assumption about where the inequality will sit.

11 · Civilizational implications

Speculative If the verification link is solved, the rate of science becomes a function of capital rather than of trained attention, and that is a genuine discontinuity. The constraint on scientific output has always been the number of people capable of judging whether a result is real. Speculative A calibrated automatic verifier removes that constraint in whichever domains it covers, and the domains it covers first will be the ones with cheap ground truth — mathematics, algorithms, materials with clean characterisation, structural biology.

Frontier If it is not solved, the ceiling is set by instruments and the shape of science barely changes. This brief’s reading of the evidence favours the second outcome for at least the next decade, on the grounds that the one fully autonomous discovery pipeline subjected to independent audit failed on measurement, and that the measurement problem has attracted a small fraction of the investment that the reasoning problem has. Speculative On that reading the striking civilizational fact of the coming decades is not accelerated discovery but a large increase in the volume of scientific-looking output at roughly constant epistemic yield, which is a worse outcome than no change.

Speculative The most valuable civilizational outcome may be the unglamorous one. A machine that reliably audits the existing literature — the task Eve was pointed at — would be worth more than any single discovery claim made to date, because the stock of published results is large, the fraction that replicates is unknown, and every downstream decision built on it inherits that uncertainty. Speculative Automating refutation is the intervention with the best ratio of civilizational value to technical difficulty in this entire subject, and it is the one nobody is pursuing. Handwave A civilization that could cheaply establish which of its recorded beliefs are true would have acquired something more consequential than a faster discovery engine, and it would have acquired it from the same machinery.

12 · Timelines

These horizons track the verification problem, because it is the link every other outcome waits on. They do not track agent capability, which is improving on a faster and less informative schedule.

  • 10 yr: Frontier Reproduction benchmarks are saturated and the 20-to-35% band is a historical curiosity; agents reproduce published computational results routinely. Frontier A reliable automated Rietveld tool exists, because the problem is well posed and someone eventually funds it. Speculative The first ROC curve for an automated claim verifier is published and is worse than its authors expected. Speculative Several more wet-lab-validated molecules arrive from agent pipelines, all from the laboratories that built the pipelines, and the human share remains unquantified. Speculative AI-generated papers are accepted at mainstream venues without disclosure, and at least one high-profile retraction follows.
  • 25 yr: Speculative Automated verification is solved in domains with cheap ground truth and unsolved everywhere else, and the map of which sciences accelerate follows that line rather than any line drawn by cognitive capability. Speculative Autonomous laboratories are standard instrumentation in materials and synthetic chemistry, and their outputs are trusted at roughly the level a competent postdoc’s are. Speculative A machine-originated result has been independently replicated by a laboratory with no relationship to its authors, and this is treated as the field’s real milestone in retrospect. Frontier Automated refutation exists as a small funded programme, probably attached to a reproducibility initiative rather than to an AI laboratory.
  • 50 yr: Speculative Question selection is partially automated in well-mapped fields, and the resulting portfolio is measurably different from the human one — less fashionable, more exhaustive, worse at recognising when a field is exhausted. Speculative The authorship convention has settled, and it settled by practice rather than by argument. Handwave A machine has produced a result in a domain it selected, which a field that did not commission it treats as a contribution; if this happens, the domain will have been one with an exact verifier, which will be held against it.
  • 100 / 250+ yr: Handwave Either open-ended machine discovery is ordinary and the interesting historical question is why it took so long to admit that instruments, not ideas, were the constraint — or the whole enterprise is understood as having been an efficiency technology all along, exactly as the 2004 paper actually claimed, and the discovery framing is remembered as a category error the field committed on its founding day and never corrected.

13 · Technology tree & dependencies

  • Depends on This brief depends on Artificial General Intelligence for its reasoning layer, and the dependency is sharper than it looks. Every system surveyed here uses a frontier language model as its hypothesis generator, so any claim about what an artificial scientist can originate is downstream of the unsettled question of whether such models infer or retrieve. If the sceptical reading is right, an autonomous discovery system is a recombination engine over the published corpus and its outputs are bounded by that corpus; if the generality claims are right, the generator is adequate and the entire remaining problem is verification. This brief argues the second regardless of how the first resolves, because the one autonomous pipeline that has been independently audited failed at measurement rather than at reasoning — but the reading matters for what happens after verification is solved.
  • Enables What this brief would enable, if its binding link were cleared, is a general instrument rather than a set of results. A calibrated verifier for non-formal claims — something that decides at expert reliability and low cost whether a stated scientific proposition is supported — is simultaneously the missing component of autonomous discovery, the missing component of automated literature audit, and the missing test for whether a machine-generated hypothesis was derived or retrieved. Those three problems are usually treated as belonging to different fields. They are the same problem, and a solution to any one of them would be recognised within months as a solution to the other two.
  • Adjacent The adjacent work sits in three directions. Multi-agent systems supply the architecture every recent scientific agent is built on, and hand this brief their unresolved question about whether ensembles beat their best member under matched compute. Artificial creativity supplies the novelty-measurement problem, including the uncomfortable finding that expert judgements of novelty are themselves unreliable instruments. Automated theorem proving supplies the existence proof: one domain where machine-generated results are checkable by construction, which is why the field’s hardest results live there and why they generalise least. Structure prediction and protein design supply the tools that produced the only wet-lab-validated molecules in this brief, with the agents contributing orchestration rather than predictive power. The fourth direction is the one this brief keeps returning to and which has no obvious institutional home: automated instrument interpretation, of which reliable Rietveld fitting is the worked example. It belongs to crystallography rather than to artificial intelligence, it is the named cause of the field’s most public failure, and its absence is the reason a robot that can synthesise a compound cannot yet tell anyone what it synthesised.

14 · Common misconceptions & speculative claims

Established “A robot scientist discovered new knowledge independently in 2004.” The 2004 Nature abstract claims a closed hypothesis-experiment-falsification loop and cost-efficient experiment selection — competitive with human performance, with 3-fold and 100-fold cost decreases against cheapest and random selection — and contains no discovery claim. Frontier The discovery claim most plausibly belongs to King and colleagues’ 2009 Science paper, which this brief could not obtain and does not cite for any specific result. Established Anyone repeating the 2004 attribution is repeating an encyclopaedia entry, not a paper.

Established “DeepMind discovered 2.2 million new materials.” The GNoME abstract gives three different numbers for three different things: 2.2 million structures below the current convex hull, 381,000 newly discovered stable materials, and 736 that “have already been independently experimentally realized.” Frontier The first is a computational screen, the second a subset of it under a stability criterion, and only the third has been made and measured by anyone; coverage that quotes the first number is quoting a search result.

Established “An autonomous laboratory synthesised 41, or 43, new compounds.” The Nature abstract says 36 compounds from 57 targets over 17 days. Established The independent re-analysis in PRX Energy describes the same work as reporting the autonomous discovery of 43 novel materials, examines all 43 synthetic products, and concludes “no new materials have been discovered in that work,” with two thirds of the claimed successes “likely to be known compositionally disordered versions of the predicted ordered compounds.” Established The named causes are unreliable automated Rietveld analysis of powder x-ray diffraction data and neglected disorder. Frontier The 36-versus-43 discrepancy is itself unexplained by either abstract, so no single count for this experiment should be quoted as settled; what should be quoted is that the failure was instrumental.

Established “An AI wrote a paper that passed peer review.” One of three manuscripts, at an ICLR workshop rather than the main conference, exceeded the average human acceptance threshold — a score comparison. Established The authors’ own phrasing is “successfully navigating a peer review,” not publication, and nothing in the abstract claims the paper entered the published record. Frontier Strip any of those conditions and the sentence becomes false.

Established “The AI Scientist produces conference-quality papers for $15.” The $15 is a compute cost and is probably accurate. Established The quality judgement in the v1 abstract is made by “our automated reviewer” — a reviewer built by the same team, evaluating the same team’s system. Established That is not a quality claim; it is a self-consistency claim.

Established “LLMs generate more novel research ideas than expert researchers.” True as stated in Si et al., at p < 0.05, in a design that controlled the obvious confounders and blinded its reviewers. Established The same abstract reports the ideas were judged weaker on feasibility, that language-model self-evaluation failed, that generation lacked diversity, and that expert novelty judgement is itself unreliable — which is why the authors propose an execution study. Speculative Until it reports, the finding is about a rating scale.

Established “AlphaEvolve did worse than AlphaTensor — 48 multiplications against 47.” Different arithmetic. Established AlphaTensor’s 47 is for 4×4 matrices in finite fields; AlphaEvolve’s 48 is for 4×4 complex-valued matrices, described as the first improvement on Strassen in that setting after 56 years. Established Both papers state their setting in their own abstracts, and the comparison is meaningless without it.

Established “Google’s co-scientist independently discovered a bacterial gene-transfer mechanism.” The v1 abstract says it “recapitulated unpublished experimental results via a parallel in silico discovery” — the collaborating human team already had the result. Established And the claim is version-dependent: the later version, retitled “Accelerating scientific discovery with Co-Scientist,” drops both that recapitulation and the liver-fibrosis organoid validation from its abstract, retaining the AML repurposing result. Frontier A citation to this system without a version number does not identify a claim.

Established “AI systems now do science end to end.” The finding from a frontier industrial laboratory’s own thirteen-task open-ended research benchmark is that models “can improve on given baselines” but “do not generate novel hypotheses, algorithms, architectures, or substantial improvements.” Established Agent Laboratory’s own headline finding is that stage-wise human feedback significantly improves research quality. Established Six independently constructed benchmarks put the best systems between roughly 17% and 42%.

Established The sceptic-side error, stated fairly: “no machine has ever produced a genuinely new result.” FunSearch improved a published cap set capacity lower bound from 2.2180 to 2.2202 and found a 512-element cap set in dimension 8; AlphaEvolve improved on Strassen in the complex setting; the Virtual Lab’s agent-designed nanobodies bound JN.1 and KP.3 variants in a wet laboratory. Frontier The honest sceptical position is narrower and stronger: in every one of those cases a human specified the objective, and none of the systems chose its own question.

Handwave The exotic version, stated as strongly as it can honestly be put. Suppose scientific taste is nothing more than a learned evaluator. Speculative FunSearch and AlphaEvolve show that with a scalar evaluator an LLM’s proposal distribution can be steered into regions containing genuinely new mathematics; if taste is an evaluator, it is trainable, and the barrier between closed formal domains and open empirical ones is quantitative rather than categorical. Handwave What would have to be true: an evaluator learned from human scientific judgement in one domain would have to transfer to a domain it was never trained on. Speculative Nobody has tried it, which makes it the most interesting unattempted experiment in the subject. Frontier What the mainstream says back is the A-Lab result: an evaluator learned in one place transferred to powder diffraction and confidently certified 43 products as novel, and hand re-analysis found none were. Speculative Transfer of a verifier is not a lesser problem than transfer of a generator; it is the same problem, and it is the one where being wrong is invisible.