1 · Concept overview
The science of science is the proposal that discovery is a production process that can be measured like any other: inputs of people, money and instruments; stages with clocks on them; outputs whose quality can be scored. Three briefs on this site each hold a piece of it. Artificial Scientists owns the automation of the discovery loop and its verification problem; Scientific Funding Models owns the design of the funder; Innovation History owns the long-run patterns and the measurement-validity traps under them. None of them owns the production function itself — which tools shorten which research stages, whether output quality changes when they do, where tacit knowledge binds, and how funding and infrastructure alter discovery yield. That joint question is this brief’s.
Established The field’s condition can be stated in one sentence: its inputs are superbly measured, its outputs are measured by instruments under active dispute, and its interventions are almost never evaluated at all. Publication counts, team sizes, career ages and grant dollars are among the best-instrumented quantities in social science. The headline output measures — citations, novelty scores, the disruption index — are each the subject of a live methodological fight declared in section 2 rather than smoothed over. And the natural experiments that could settle what actually raises discovery yield — funding lotteries, people-not-projects contracts, AI tool rollouts — are mostly running unanalysed, their assignment and outcome data sitting in administrative records nobody has published.
Frontier The organising finding of the last two years is a cautionary symmetry. The single most influential empirical estimate of AI’s effect on discovery yield was withdrawn when its host institution said it had no confidence in the data; the one randomised field measurement of an AI research tool found its users slower while believing themselves faster. The field that studies how science knows things was itself, on its most policy-relevant question, running on an unverified number and an unmeasured intuition. Section 2 carries both cases with dates.
Established A note on sourcing. This brief was commissioned in September 2026 from the Institute’s research base. Reading-list entries without links are cited from the bibliographic record rather than re-fetched, and claims are dated no later than early 2026 unless carried by a linked source.
2 · Current scientific position
Established Start with the inputs, which are not in dispute. The scientific literature has grown at 8–9% a year on Bornmann and Mutz’s bibliometric series — a doubling time near nine years. Teams have displaced solo authors across sciences, engineering and patenting, and the citation advantage of teams grew for five decades (Wuchty, Jones and Uzzi). The age at which scientists make first contributions has risen and specialisation has narrowed — Benjamin Jones’s “burden of knowledge”: each generation must climb a longer ladder before reaching the frontier. Established Bloom, Jones, Van Reenen and Webb assembled the aggregate consequence: across the economy and within case studies — transistor density, crop yields, drug discovery — measured research effort has risen enormously while output per researcher has fallen steadily; holding Moore’s law on its exponential now takes many times the researchers it took in the early 1970s. Frontier What that shows is contested at the interpretive layer only: ideas may be getting harder to find, or the maintained exponential may be the wrong null. The input-per-output arithmetic itself has not been overturned.
Frontier The output side is a declared dispute, and this brief declares it. Park, Leahey and Funk (Nature, 2023) computed the CD disruption index over 45 million papers and 3.9 million patents and reported declines on the order of 90% for papers since 1945: science, on this instrument, increasingly consolidates rather than disrupts. Frontier Petersen, Arroyave and Pammolli replied that the index is mechanically biased by citation inflation: mean reference lists grew from about 9 in the 1960s to about 23 in the 2000s, the citation network grows 5.1% a year, and correcting for reference-list growth shrinks the decline and reverses the sign of the team-size coefficient. Established Innovation History types the recomputed series as a missing scientific result — arithmetic on public data that nobody has published — and until it exists the most cited quantitative claim about the trajectory of science can be neither used nor discarded. This brief adopts that position and adds one consequence: Wu, Wang and Evans’s celebrated finding that small teams disrupt and large teams develop rests on the same disputed instrument and inherits the same suspension.
Established Quality has one direct measurement tradition: replication. The Reproducibility Project in psychology (2015) got significant results in roughly a third of about a hundred replications, with effect sizes near half the originals; the Social Sciences Replication Project (2018) replicated 13 of 21 Science and Nature papers, again at about half the original effect size; the cancer-biology counterpart found median effect sizes 85% smaller than originally reported. Established The startling half of this record is that failure is predictable: prediction markets anticipated psychology replication outcomes with about 70% accuracy in Dreber and colleagues’ PNAS study, performed comparably in the 2018 project, and structured expert elicitation has since matched them — the community holds reliable collective knowledge about which findings are false, and the publication system does not use it.
Frontier Funding-side evidence is thin, quasi-experimental, and consistent. The one outside-economist comparison of a people-not-projects contract — Azoulay, Graff Zivin and Manso on HHMI investigators against matched NIH-funded peers — found more top-percentile papers, more new research directions, and more flops: higher variance, exactly as the long-leash contract design predicts. Wang, Lee and Walsh find stable non-competitive funding associated with more novel output than competitive project grants in Japanese data. Established Funding lotteries have run at real funders since New Zealand’s Health Research Council began randomising in 2013, and Scientific Funding Models records the punchline this brief inherits: more than a decade on, no funder anywhere has published an outcome comparison between lottery- and panel-allocated grants, though assignment and outcomes sit in their own records. The experiments exist; the analyses are withheld.
Frontier The AI-for-science record as of early 2026 divides cleanly by verification, and its two poles are instructive. Where the output is checked by an exact or near-exact verifier, gains are real and large: protein structure prediction reached accuracy competitive with experimental structures for a large class of targets, shrank a months-long experimental stage to minutes of inference for that class, and earned the 2024 Nobel Prize in Chemistry. Frontier Where verification is weak the record collapses on inspection. Google DeepMind’s GNoME predicted 2.2 million crystal structures with 381,000 flagged stable — a generator’s claim awaiting synthesis; the companion autonomous lab claimed 41 new materials, and Artificial Scientists carries the independent re-examination that found no new materials at all, for reasons of measurement rather than reasoning. Established The MIT materials-economics preprint — which claimed a corporate AI tool raised materials discovered by 44% and patents by 39%, and was widely quoted as the first credible productivity estimate — was withdrawn in May 2025 after MIT stated it had no confidence in the provenance, reliability or validity of the data. It was retracted; every downstream citation of its numbers is now unsupported, and this brief treats the episode as a finding about the field’s hunger for the number rather than as an isolated fraud. Frontier And the one randomised field trial of an AI research tool on experienced practitioners — METR’s early-2025 study of open-source developers on mature repositories — measured a 19% slowdown against a believed 20% speedup. Software is not bench science, but it is the best-instrumented stage we have, and the perception–measurement gap is the single most transferable result on this page.
3 · Frontier questions
Frontier Which tools shorten which stages is answerable stage by stage, and mostly unanswered. The discovery pipeline has perhaps six clocks: literature synthesis, hypothesis generation, experiment design, execution, analysis, and writing-and-review. Execution has the one clean win (structure prediction, for the subset of questions where a predicted structure suffices). Writing and coding show large self-reported gains and the METR result stands as the warning that self-report can carry the wrong sign. Frontier Hypothesis generation is genuinely contested: Si, Yang and Hashimoto found LLM-generated research ideas judged more novel than those of over a hundred NLP researchers — while Artificial Scientists adds that expert novelty judgement is itself an unreliable instrument, which converts the finding from an answer into a measurement problem. Established Literature synthesis, the stage AI should most obviously compress, has no published trial with quality endpoints at all.
Frontier Where tacit knowledge binds is the frontier’s second question, and the evidence points at the bench. The autonomous-lab failure was an XRD-interpretation failure — the machine could pipette but not judge a powder pattern; Harry Collins’s classic studies of laser builders found that no written protocol sufficed to transmit a working device without person-to-person contact. Speculative The live hypothesis is that AI compresses exactly the stages whose knowledge is already codified and leaves the tacit stages as the new binding constraint — which would predict rising returns to experimental craft and no aggregate acceleration until robotic execution closes the gap. No study yet tests it.
Frontier The third frontier question is whether quality moves when quantity does. Every tool that cheapens production shifts the quantity–quality frontier somewhere; the replication-market result says the community can score quality cheaply, so the design that would settle it — tool-assisted versus unassisted output, scored blind by replication markets — is buildable with existing parts. Nobody has run it.
4 · Technological bottlenecks
Established The instrument bottleneck comes first. Citations measure attention, not truth; the disruption index is under the dispute of section 2; novelty scores disagree with each other; and replication — the one direct quality probe — costs so much that the flagship projects managed a few hundred studies in a decade against a literature growing by millions of papers a year. A validated, affordable, non-citation outcome measure of discovery yield is the field’s missing voltmeter.
Established The telemetry bottleneck is second. Nobody logs the pipeline. Stage-level timings of actual research — how long the literature took, how long the failed syntheses took — exist nowhere outside a handful of instrumented software studies, which is why METR’s design was possible in code and has no bench-science counterpart. Electronic lab notebooks capture some of it and are neither standardised nor shared. Frontier The UMETRICS tradition — linking grant transactions to personnel and outputs in administrative data — shows the plumbing can be built when a consortium wants it built.
Established The evaluation bottleneck is institutional and inherited. Funders run the natural experiments and do not publish them; the assignment data that would identify funding effects is theirs alone. This brief takes the finding from Scientific Funding Models without restating its casework: what is missing is not a method but a decision to publish.
5 · Research dependencies
Established This brief consumes the three siblings’ results as inputs: the verification dichotomy from Artificial Scientists (cheap exact verifiers separate the AI results that survive scrutiny from those that do not), the funder-design record from Scientific Funding Models (every alternative running, none compared), and the instrument-validity discipline from Innovation History (any pattern found in patents is first a pattern in patenting — a rule this brief extends to citations and disruption scores). The peer-review and lottery machinery beneath the funding layer belongs to Scientific Governance Models and is used, not re-derived.
Frontier Outside the corpus, two dependencies bind. Causal inference on observational science-of-science data — the field’s identification revolution is one generation old, and most of its celebrated correlations (team size, interdisciplinarity, mobility) still lack designs that would survive an economics seminar. Frontier And access: the administrative, notebook and telemetry data the next results need is held by funders, publishers and labs with no obligation to share, so the field’s progress is rate-limited by data politics rather than by ideas — an irony it is well equipped to appreciate.
6 · Required experiments
Frontier The decisive test is a stage-instrumented randomised rollout of an AI research tool across working laboratories, with quality endpoints scored by replication rather than by citation. Randomise access at the lab level; log stage clocks (literature, design, execution, analysis, writing); pre-register the quality metric; score a sample of outputs by independent replication or replication market two years on. It is the METR design scaled from code to bench, and every component — the randomisation, the telemetry, the market scoring — has been demonstrated separately. Established No funder has yet commissioned it, and until one does, every claim that AI is accelerating science rests on self-report, vendor benchmarks, or the withdrawn preprint of section 2.
Frontier Three inherited experiments would each settle a standing dispute. The disruption series recomputed with the published citation-inflation correction, across the same six datasets — typed by Innovation History as arithmetic awaiting an author — would settle whether the slowdown-of-disruption is a fact about science or about reference lists. Frontier A funder publishing lottery-versus-panel outcomes — the analysis Scientific Funding Models shows is a decision, not a discovery — would put the first causal number on peer review’s selection value. Frontier And a blinded replication-market score of AI-assisted versus unassisted papers in one field would put the first number on whether tool-compressed production moves quality, in either direction.
Speculative The long experiment is institutional: a funder that randomises at the portfolio level — contract length, failure tolerance, person-versus-project — and holds an evaluation window longer than a programme authorisation. The corpus’s funding brief types that window as a missing institutional constraint; this brief notes only that semi-endogenous growth arithmetic makes the value of information from such a funder enormous relative to its cost.
7 · Engineering requirements
Established The engineering is data plumbing, and it is modest by physical-science standards. Stage telemetry means electronic lab notebooks with standardised event schemas and privacy-preserving export — components that exist commercially and have never been federated for research-on-research. The UMETRICS precedent shows universities will pool transaction-level administrative data under a consortium agreement; extending the linkage from payroll to instruments and outputs is schema work, not invention. Frontier Replication capacity is the expensive component: scoring even 1% of a national portfolio by direct replication would require dedicated replication labs funded as infrastructure — the Institute for Replication currently re-analyses economics and political-science papers on a budget that would not commission a single wet-lab replication programme.
Frontier The market machinery is the cheap component. Replication markets and repliCATS-style structured elicitation run on software and small honoraria; their demonstrated accuracy near 70% is already useful as a triage instrument, and the engineering task is integration — wiring the score into editorial and funding decisions — rather than construction. Speculative A fully instrumented “glass pipeline” laboratory, logging every stage of real projects for research-on-research, has been proposed informally many times and built never; nothing in it exceeds current engineering.
8 · Adjacent technologies
Established The three parents first: Artificial Scientists (the loop and its verifier), Scientific Funding Models (the funder), Innovation History (the long series and their artefacts). Scientific Governance Models holds the peer-review and integrity machinery this brief’s quality measures would plug into; Human-AI Integration holds the perception–performance gap literature the METR result belongs to; Artificial General Intelligence holds the question of whether the hypothesis generators are inferring or retrieving, on which the ceiling of AI-for-science depends. Frontier A useful boundary: this brief measures the production of knowledge; it does not adjudicate any particular scientific claim, which keeps its quarrels methodological.
9 · Institutional requirements
Established Metascience has begun acquiring state organs, which is new. The US NSF has run a science-of-science programme for two decades; the UK stood up a dedicated government metascience unit in 2023–24 whose early grant rounds were built around AI-in-science questions; DARPA’s SCORE programme demonstrated confidence-scoring of social-science claims at scale. Frontier None of these yet holds the two levers that matter: none can compel a funder to publish its natural experiments, and none funds replication as standing infrastructure rather than as projects.
Established The incentive structure is the standing obstacle, and it is measured. Competitive grant systems consume large fractions of the capacity they allocate — surveyed federally funded US faculty report over 40% of research time going to administrative and grant tasks, and contest models show honest conditions under which application costs consume the value distributed. A field whose own findings indict its host institutions’ allocation machinery faces the awkward fact that those institutions are also its funders. Frontier The publication system has the same reflexive problem: journals profit from the quantity growth whose quality consequences metascience is trying to measure, and no journal yet prices a replication-market score into acceptance. Speculative The plausible institutional home for the decisive experiment of section 6 is therefore not a legacy funder but one of the new philanthropies or state units with no portfolio to defend — a prediction this brief states so it can be scored.
10 · Ethical & societal considerations
Established Telemetry on discovery is surveillance of discoverers. Stage clocks and notebook logs are exactly the data an employer could repurpose for individual performance management, and the history of bibliometrics — invented as a mapping tool, deployed as an evaluation weapon — is the precedent every consent framework should be designed against. Established Goodhart’s law is not hypothetical here: citation metrics measurably reshaped publication behaviour, and any new yield metric this field validates will be gamed in proportion to its adoption. Frontier Credit assignment under AI tools is unsettled in both directions: tools erase the contribution trails of the people whose work trained them, while authorship conventions have not decided what a scientist must have done to sign a machine-drafted finding. Speculative And the field’s own results can harm: a published finding that a demographic or institutional group’s output scores lower on some yield metric will be read as an allocation instruction long before its confounds are audited.
11 · Civilizational implications
Frontier The stakes are growth-theoretic. In semi-endogenous growth accounting, the Bloom–Jones arithmetic — flat output per researcher maintained only by exponential growth in researchers — implies that when the researcher population stops growing, frontier growth stalls. Rich-country research workforces are close to that regime, which is why the question of whether AI genuinely raises discovery yield per person is not a productivity curiosity but the hinge variable of long-run growth. Frontier The honest current answer assembled in section 2 — real gains where verification is exact, no credible aggregate estimate elsewhere, and the flagship estimate withdrawn — means civilisation is currently making trillion-scale capital allocations on that hinge without a measurement. Speculative If the decisive experiment ran and found large stage-compression with quality held, the case for treating AI-for-science as core infrastructure would be made; if it found METR-pattern illusory gains at the bench, the correction would be worth more than the tools. Handwave The strongest claim in circulation — that AI will compress a century of science into a decade — is not supported by any measured stage clock on this page and works entirely by assertion from benchmark performance to discovery yield.
12 · Timelines
These horizons track when the measurement deficits above get instruments, not when science speeds up.
- 10 yr: Frontier The disruption dispute is settled by recomputation; at least one funder publishes a lottery-versus-panel outcome comparison; at least one stage-instrumented randomised AI-tool trial with replication endpoints reports; replication-market scores are used in triage by some journals and funders.
- 25 yr: Speculative Discovery yield has a validated non-citation outcome measure in at least some fields; funder portfolio experiments with decade evaluation windows exist; the tacit-knowledge boundary has been mapped by which stages robotic execution actually absorbed.
- 50 yr: Speculative Research-on-research operates the way clinical epidemiology operates on medicine — routine, regulated, and consulted before major allocation decisions; whether AI moved aggregate discovery yield is by then an answered historical question rather than a live one.
- 100 / 250+ yr: Handwave Deliberate civilizational steering of research portfolios against measured yield curves; no current evidence grounds a forecast at this range.
13 · Technology tree & dependencies
- Depends on Three sibling results this brief builds on rather than re-derives: the verification dichotomy of Artificial Scientists, which predicts where AI gains will and will not survive audit; the none-compared finding of Scientific Funding Models, which locates the funding evidence gap in publication decisions rather than missing experiments; and the instrument-validity rule of Innovation History, which governs every output series used here.
- Requires (not on this map) Five constraints. A disruption series recomputed with the published correction for reference-list growth, shared with Innovation History and repeated here because two briefs now wait on the same arithmetic; a published outcome comparison against the mechanism replaced, shared with Scientific Funding Models for the same reason; a stage-instrumented randomised trial of an AI research tool with replication endpoints, this brief’s own decisive test, buildable from demonstrated parts; routine third-party replication capacity funded as infrastructure rather than as heroic projects, without which no quality endpoint scales; and a validated outcome measure of discovery yield that is not citation-based, without which every acceleration claim reduces to counting attention.
- Enables An evidence base for the largest discretionary capital allocation of the decade — AI-for-science investment — currently made without one; funder design chosen on outcomes rather than tradition; and a quality-scored literature in which the replication-market signal the community demonstrably possesses is finally wired into the machinery that decides what gets published and funded.
- Adjacent Scientific Governance Models for the review and integrity layer; Human-AI Integration for the perception-performance gap; Artificial General Intelligence for the generator ceiling.
14 · Common misconceptions & speculative claims
Frontier “Science is measurably slowing down; the disruption index proved it.” The decline is on the instrument; whether it is in the world is exactly what the citation-inflation critique contests, with the correction shrinking the effect and flipping the team-size sign. The claim is suspended, not settled — and the patent-count precedent from Innovation History, where a celebrated 1970s innovation slowdown turned out to be a Patent Office budget artefact, is the base rate to hold in mind.
Established “AI has already been shown to accelerate science across the board.” The record as of early 2026: one exact-verifier triumph (structure prediction), one flagship autonomous-lab claim refuted on independent re-examination, the leading economics estimate withdrawn by MIT with no confidence in its data, and the one randomised trial of practitioner use measuring a slowdown its subjects experienced as a speedup. Real gains exist; the across-the-board claim is not supported by any study that survives the verification test.
Frontier “The replication crisis shows most published research is false.” Measured replication rates run from roughly a third (psychology, 2015) to nearly two-thirds (top-journal social science, 2018), with large effect-size shrinkage everywhere — bad, but field- and design-specific, and coexisting with the finding that experts and markets can predict which results will fail. The literature is not uniformly rotten; it is unevenly rotten in ways the community can already detect and the institutions do not act on.
Frontier “Fund people, not projects — it is proven.” One quasi-experimental comparison at one elite funder, showing higher variance and more top-percentile hits among hand-picked investigators, is the entire causal evidence base. It is genuinely encouraging and genuinely singular, and the selection into HHMI is nothing like the marginal case a policy would govern.
Handwave “Double the science budget, double the science.” The system’s own overheads are measured: over 40% of funded faculty research time goes to administration, contest models show application costs consuming allocated value under honest assumptions, and the Bloom–Jones arithmetic says even proportional output growth buys diminishing idea output per dollar. Scaling the current machine scales its losses with it; the whole case for this field is that the machine’s design, not its throughput, is the free variable nobody has tuned.