1 · Concept overview

Science governs itself through four machines: it allocates money by peer-reviewed grant competition, it gatekeeps publication by editorial and referee judgement, it sets the rules under which a result counts as reproducible, and it polices misconduct through institutional and statutory integrity offices. Each machine keeps administrative records, and that is why this subject can be argued from measurement rather than from anecdote — 102,740 scored grants with their downstream output attached, a randomised trial of referee error detection, a hundred-study replication project, retraction counts by country, a federal integrity office publishing its own allegation totals, and roughly a dozen funders who have allocated real money by drawing lots.

The framing under test is that this machinery works, and that where it fails the fix is more of the same machinery. It comes back wrong, and it is wrong in a specific and useful direction rather than a boring one. The interventions in the record with demonstrated effects are all checkpoints applying a rule to every item — screen every accepted paper's images, hide the author's name, fix the publication decision before the data exist, draw lots among everything above a threshold. The interventions that fail are all attempts to improve a fine-grained judgement — better review criteria, referee training, shorter application forms, more review stages. More of the same machinery means more judgement. What is measured favours less judgement and more rule.

2 · Current scientific position

Frontier Start with the central dispute, because it is usually reported as settled and it is not. Fang, Bowen and Casadevall analysed 102,740 funded NIH R01 grants from 1980–2008 and found the percentile score explained r² = 0.0078 of variation in both publications and citations — a slope of −0.132 publications and −9.6 citations per percentile point, a random-forest model attributing about 1% of productivity variance to score, and discrimination of AUC 0.54 against a chance baseline of 0.50. Inside the 3rd–20th percentile band that actually decides funding, adjacent cohorts were statistically indistinguishable; only proposals at the 2nd percentile separated out on citations. Frontier Li and Agha, in Science, on more than 130,000 NIH R01 grants — essentially the same population — report the opposite. A one-standard-deviation worse peer-review score is associated with 15% fewer citations, 7% fewer publications, 19% fewer high-impact publications and 14% fewer patents. The disagreement is specification, not data. Fang and colleagues regress raw percentile on raw output; Li and Agha condition on applicant and field characteristics. The first specification asks whether score predicts output among the grants that were funded; the second asks whether score carries information once you hold the investigator's track record and field constant. Read together they say the score contains real signal that is small relative to who the applicant already is, and any brief printing one result without the other is printing half a literature.

Established The mechanism underneath the sceptical result is reviewer disagreement, and it has been measured directly. Pier and colleagues recreated four NIH study sections with 43 oncology researchers rating 25 real R01 applications, yielding 83 primary-reviewer ratings. The intraclass correlation for overall rating was ICC = 0, 95% CI [0, 0.14]; Krippendorff's alpha was 0.024. Ratings for the same application were no more similar than ratings for different applications, and reviewers scored previously unfunded applications as favourably as funded ones (p = 0.58). Their conclusion was that the outcome “depended more on the reviewer to whom the grant was assigned than the research proposed.” Established That result is substantially rehabilitated by a statistical correction, and the correction is serious enough that the zero must not be quoted bare. Erosheva, Martinková and Lee showed in the Journal of the Royal Statistical Society Series A that near-zero agreement estimates are largely an artefact of range restriction: computing agreement only among top-quality proposals mechanically collapses the between-proposal variance component. Across the full range of NIH and AIBS applications, single-rater reliability was 0.34 (95% CI 0.31–0.37) for NIH and 0.37 (0.22–0.52) for AIBS, rising to 0.61 and 0.64 for three-rater panels. Zero estimates appeared once fewer than about 70% of top-quality and 45% of bottom-quality proposals were retained.

Established Both are true at once, and read together they give the sharpest statement available about grant peer review: it reliably separates good proposals from bad ones and cannot separate good proposals from each other. Funders reject the bad ones long before the funding line and then make fine-grained cuts inside the surviving pool. The reliability that exists is spent where it is not needed and is absent where the decision is made. Established The distributional record points the same way. Taffe and Gilpin, on NIH award data, report applications with white principal investigators 1.7 times more likely to be funded than those with African-American or Black principal investigators — 29.3% against 17.1% in FY2000–2006 and 17.7% against 10.7% in FY2011–2015 — with topic choice accounting for just over 20% of the gap. A later bibliometric analysis of 2,397 NIH biosketches from FY2003–2006 found a raw gap of −13.4 percentage points, of which a full model including publication record explained 52%, leaving −6.9 points unexplained.

Frontier The largest documented effect of any governance rule in this subject comes from changing when the publication decision is made, not who makes it. Scheel, Schijen and Lakens compared 71 psychology registered reports — papers accepted on the strength of the protocol, before the data exist — against 152 standard reports sampled from 633 psychiatry and psychology journals over 2013–2018. Hypotheses were supported in 96.05% (95% CI 91.61–98.54) of standard papers and 43.66% (31.91–55.95) of registered reports: a gap of 52.4 percentage points, or 95.95% against 50.00% restricting to original rather than replication studies. Frontier The authors are explicit that this “was not an experiment” — hypotheses, authors and editors were not randomly assigned to format — and name differing hypothesis priors, author conscientiousness and editorial policy as confounds. Take that caveat at full weight and the discontinuity is still the largest in the published record attributable to a governance rule, because the rule changes only the timing of one decision.

Established Adoption has not followed the evidence, and the ceiling is measured. Lin, Cheng, Cheng and Hung's July 2023 census found 278 journals offering registered reports, 186 of them indexed in the Journal Citation Reports. Uptake ran 0–7% across major research fields and 0–34% across subfields: 7% in the psychiatry and psychology group, 34% in experimental psychology, and at or below 1% in most fields. Frontier The format also degrades in use. Claesen and colleagues audited the first generation of preregistered psychology studies: of 23 articles carrying 38 plans, only 16 articles with 27 plans were accessible and detailed enough to audit at all — 7 plans lacked a statistical analysis plan, a variable list or a sample size. Of those 27, 25 (93%) deviated from the plan, 9 (36%) disclosed none of their deviations and 1 (4%) disclosed all. The sample is small and drawn from the earliest cohort and should be quoted with that attached, but it is the direct measurement of whether the rule binds, and it mostly does not.

Established Replication rates cluster in a narrow band across fields that share almost nothing else. Psychology (Open Science Collaboration, 100 studies): 97% of originals significant, 36% of replications (35 of 97), 39% on the subjective “did it replicate?” judgement, mean effect size 0.403 falling to 0.197, cognitive psychology 50% against social psychology 25%, main effects 47% against interactions 22%. Economics (18 laboratory experiments from AER and QJE, 2011–2014): 61.1% replicated in the same direction against 92% expected if all originals were true, standardised effect 0.474 to 0.279. Social science in Nature and Science, 2010–2015 (21 studies): 62% significant in the same direction on replication samples about five times larger. Preclinical cancer biology: 43% on direction-plus-significance, 58% on prediction interval, 62% meta-analytically, 40% on a combined three-of-five rule, with a median replication effect 85% smaller than the original. Established Many Labs 2 removes the standard defence. Across 28 effects, 125 samples, 36 countries and 15,305 participants54% significant in the same direction at p<.05, median Cohen's d 0.60 falling to 0.15, 32% in the opposite direction — variability in observed effect sizes was “attributable more to the effect being studied than to the sample or setting in which it was studied”. Hidden moderators are not the explanation for the effects tested.

Established The binding constraint on repeating a result is not statistics. It is that the reporting rules journals already have do not make work repeatable. The Reproducibility Project: Cancer Biology set out to replicate 193 experiments from 53 papers published 2010–2012. It published 87 experiments as registered reports, started 76, and completed 50 from 23 papers — roughly 74% attrition over seven years and over budget. None of the 193 experiments was described in sufficient detail in the original paper to design a protocol without contacting the authors: zero required no clarification. Original authors were “extremely or very helpful” for 41% of experiments and not at all helpful or non-responsive for 32%. Raw data were obtained for 16%, summary data for 14%, and nothing at all for 68%. The project's own summary adds 2% of experiments with open data, 70% requiring a request for key reagents, 0% with completely described protocols and 67% of the experiments actually run requiring protocol modification. Every one of those figures describes a rule that journals and funders formally have — methods sections, data availability statements, materials sharing — and do not enforce. It is the sharpest available test of “the fix is more of the same machinery”: the machinery already existed on paper.

Established Integrity enforcement runs two to four orders of magnitude behind the volume it is supposed to police, and the arithmetic has to be stated carefully because the quantities are routinely merged. More than 10,000 papers were retracted in 2023, a record, of which over 8,000 came from Hindawi journals after Wiley's cleanup; the overall retraction rate passed 0.2% in 2022 and has “more than trebled over the past decade”. Per 10,000 articles the national rates run Saudi Arabia about 30, Pakistan 28.1, China 24.9, Russia 23.5, against roughly 50,000 retractions on record in total. Frontier Estimated prevalence is a different quantity from the retraction rate and the two must not be added or compared. Bik, Casadevall and Fang screened 20,621 papers from 40 journals (1995–2014) and found 3.8% containing problematic image duplication, at least half with features suggesting deliberate manipulation. Speculative A preprint by Sabel and colleagues applied a red-flag rule to 15,120 PubMed records and estimated 11.0% of 2020 biomedical publications — about 150,000 papers a year — as potentially fake, at 90% sensitivity and a 37% false-alarm rate, and said plainly that red-flagging is not proof. Speculative A 2026 machine-learning screen of 2.6 million cancer papers (1999–2024), reported in the trade press, flagged 261,245 papers (about 10%) at a claimed accuracy near 90%, with the lead author saying the figure “could actually be more” and a commenter objecting that the model may be learning the writing patterns of Chinese authors rather than fakery.

Established Against that volume, the statutory enforcement number is small, and it must be stated in the units it was measured in. The US Office of Research Integrity published two Federal Register findings of research misconduct in 2025 — the lowest count since at least 2006, against a typical run rate of about ten a year (three in 2021, seven in 2016), reported in December 2025. That is a count of published findings and it is not the same quantity as cases closed. ORI's own annual reports record 446 allegations and 177 closures in 2025 (against 713 allegations and 119 closures in 2024), and the closure charts in those reports carry a different and larger, two-digit findings count. The two figures measure different things and no ratio should be computed across them — conflating them is the exact error this brief warns about two paragraphs above. Established Institutional self-investigation is the weakest link on the National Academies' own account. Fostering Integrity in Research records ORI allegations rising to 423 in 2012 against a 1992–2007 average of 198, NSF Inspector-General investigations from 45 in 2004 to 75 in 2014, and a survey by Bonito and colleagues in which 97% of research integrity officers identified fewer than half the appropriate actions in misconduct scenarios — alongside inconsistent internal definitions, faculty committees without relevant expertise and administrators handling investigations as a side duty. That survey is from 2009. It is seventeen years old and it is still the best national evidence there is.

Established And journal peer review does not catch what it is assumed to catch, on the only randomised test of it. Schroter and colleagues ran a single-blind trial with 607 BMJ reviewers (522, 440 and 418 completing three test papers), each paper seeded with 9 major and 5 minor deliberate methodological errors. Reviewers found a mean of 2.58, 2.71 and 3.05 of the 9 major errors and about 1 of the 5 minor ones. Training, whether face-to-face or self-taught, produced only slight improvement. The authors' conclusion is the load-bearing sentence for the whole framing: “Journal editors should not assume that their reviewers will detect most major flaws in manuscripts.”

Established The integrity half of this machinery now has a volume problem that is measured rather than asserted, and the measurement is dominated by a single publisher’s cleanup. The 2023 record of more than 10,000 retractions in one year was not a broad deterioration spread evenly across the literature: on the order of 8,000 of those retractions came from Wiley’s Hindawi imprint, concentrated in guest-edited special issues, and the imprint was closed with 19 journals discontinued. Frontier The continuous census of the quantity is a privately maintained database rather than a public register. The Retraction Watch database now holds tens of thousands of records and has grown by roughly an order of magnitude across two decades, but no published decomposition separates growth in misconduct from growth in detection. That undecomposed ratio is the weakest link in everything below it, because every claim below it is a claim about a rate whose numerator is detections and whose denominator is submissions, and only the numerator is published.

Frontier The supply side has been traced to named commercial operations rather than inferred from anomalies. Abalkina reconstructed a Russian-language brokerage that advertised authorship positions on manuscripts before submission, with prices quoted per slot and by position in the author list, and traced on the order of a thousand advertised manuscripts into print. Frontier A 2024 Scientific Reports analysis of authorship-for-sale networks pushes the same finding from one broker to a network structure, which is the more useful object: if authorship slots clear at a price, the thing being governed is a market and not a deviance, and an intervention designed against individual bad actors is addressed to the wrong unit. Established The joint COPE and STM work on paper mills reported editor estimates of suspect submissions in the low single-digit percentages on average, with far higher concentrations in particular journals and fields. Handwave Editor estimates are belief, not measurement, and the circulating 2% figure should never be quoted as a prevalence rate.

Established What screening actually catches is exactly what the checkpoint reading predicts — signals a rule can apply to every item. The Problematic Paper Screener flags tortured phrases, the synonym-substituted renderings of standard technical terms that no specialist would write, together with duplicated and non-existent references, across a whole corpus rather than by sampling. Established Undeleted chatbot boilerplate has appeared verbatim in published papers, and that is a detection signal of the same kind: a fixed string, checkable by rule, on every submission. Frontier What screening does not catch is the part that matters. Text that is fluent, novel and false carries no fixed string, and general-purpose detectors for machine-written prose have repeatedly failed to hold calibration outside their validation sets, at false-positive rates high enough that no major publisher has been willing to make a detector score a ground for rejection. Speculative The plausible near-term equilibrium is that each fixed-string signal is engineered out by the mills within a year of being named, leaving screening effective against last year’s fabrication and blind to this year’s.

3 · Frontier questions

Frontier The organising question is whether the checkpoint-versus-judgement split generalises, or whether it is a pattern read backwards out of nine studies. The split is stated here as a hypothesis with a clean falsification condition: find a governance intervention that improves a fine-grained ordinal judgement and shows a replicated effect on a research outcome. Referee training is the best-designed attempt and it moved detection from 2.58 to about 3 of 9 seeded errors. Shortening a grant form is the second and it moved burden the wrong way. Nobody has run the third. Speculative The rival reading is that the checkpoint interventions are simply the ones that are easy to measure — a binary rule produces a binary outcome and a clean comparison — and that judgement improvements are real but invisible to the designs used. That is not refuted by anything in this brief.

Established The largest open question has a documented answer of “nobody has looked”, and the absence is itself the finding. Lottery allocation is the one governance mechanism that funders have actually trialled on live money. The Health Research Council of New Zealand has allocated its Explorer Grants (NZ$150,000, up to two years, about 2% of the Council's annual expenditure) by modified lottery since 2013: panel members mark each application transformative and viable, two or more yes votes puts it in a pool, and the pool is drawn by a spreadsheet random function until the budget is exhausted. The Volkswagen Foundation ran partial randomisation in its Experiment! initiative from 2017 to 2021, roughly half by jury and half by lot after a quality check; the programme is closed. The Swiss National Science Foundation adopted random tie-breaking across a portfolio of about CHF 880 million from late 2021. The British Academy has randomised part of its Small Research Grants since 2022, extended to 2028. Germany has trialled drawing lots for the right to submit a full proposal. A 2026 Nature feature counts more than a dozen funders in the Research on Research Institute's catalogue, 11 of them added since 2022. Established Thirteen years after the first randomised round, no funder anywhere has published a comparison of research outcomes between lottery-allocated and panel-allocated grants. Not a weak one. None.

Established What has been measured about lotteries is acceptability and application behaviour. Pomeroy and colleagues surveyed 325 Explorer Grant applicants from 2013–2019 with 126 responding (39%): 63% supported randomisation for that scheme against 25% opposed, but support fell to 40% for and 37% against for other grant types, and split sharply by outcome — 78% among funded applicants against 44% among declined. Median preparation time was about 10 days and 75% said the lottery had not changed how much time they invested. Frontier The British Academy reports applications rising from roughly 600 a round to over 1,100, which the Academy puts at about 70% growth and the Nature feature renders as 643 in 2022 to 1,257 in 2025–26; the two agree in direction and the discrepancy in the base is not resolved by either source. About 60% of applicants clear the quality threshold each round, success stays at 20–30%, and roughly 900 awards were in progress in January 2025 across 112 institutions plus more than 40 independent researchers. These are behavioural and distributional findings reported by the funder running the trial. They are not evidence about the science produced.

Frontier Is the registered-report effect a rule effect or a selection effect? Scheel and colleagues name three confounds and cannot rule any of them out: authors who choose the format may hold different hypothesis priors, may be more conscientious, and may face different editorial policy. Speculative The competing hypotheses are worth stating separately because they imply opposite policies. If the gap is caused by the rule, mandating registered reports would move the literature by tens of percentage points. If it is caused by who volunteers, mandating them would move nothing and would add a protocol-review stage to every submission. A trial randomising submissions to format at a single journal would separate these and has not been run at any journal in the census of 278.

Frontier How much of the published literature is fabricated is genuinely open, and the three instruments in circulation measure three different things on three different populations. Image screening on 20,621 papers gives 3.8% for one specific manipulation in one literature. A red-flag rule on 15,120 records gives 11% at a 37% false-alarm rate, in a preprint. A machine classifier on 2.6 million cancer papers gives about 10% with an unresolved objection that it may be detecting authorship language rather than fraud. Speculative The hypothesis that the true rate is an order of magnitude above the retraction rate is consistent with all three and demonstrated by none. The hypothesis that it is concentrated in a small number of paper mills and journals is consistent with the 40-fold spread between journals in Bik's screen and with the 8,000 Hindawi retractions, and is likewise not established.

Speculative Four further positions are live in this field, are held by serious people, and are not established. They are set out here rather than dropped because the disagreement is the subject. (a) The apparatus is net-negative on cost grounds — 130 million reviewer-hours a year and 550 working years in a single national grant round buy an allocation whose in-pool discrimination is close to nil, so radical thinning would lose little and save much. Speculative (b) Inconsistency is a feature. The National Academies' 2019 study on reproducibility and replicability declines the “crisis” framing and treats non-replication as potentially a precursor to discovery rather than a defect. Speculative (c) Enforcement will never scale and detection must move to automated screening at the acceptance checkpoint, which is the checkpoint hypothesis applied to fraud — and which inherits the false-alarm rates above. Speculative (d) The real bottleneck is documentation rather than judgement: on the cancer-biology evidence, zero of 193 experiments were reproducible from the paper alone, so an enforced protocol-and-materials rule would do more than any change to who reviews what. (d) is the position this brief finds best supported and it is still not established, because nobody has run the enforcement.

Frontier The open question the checkpoint-versus-judgement split now has to survive is automated review. Machine assistance at the referee stage is a judgement intervention by the classification used here — it tries to improve a fine-grained ordinal assessment — and the split therefore predicts it will not move a research outcome. The prediction is falsifiable and cheap to test: randomise submissions at one journal to machine-assisted against unassisted review, scored against seeded errors at the standard Schroter and colleagues used. Frontier Nobody has published that trial. What exists instead is the weaker evidence that machine-generated review text is rated about as useful as human review text by the authors who receive it, which measures the referee report and not the paper. Speculative The more interesting possibility is that automation converts review from judgement into checkpoint rather than improving the judgement — a model that reliably checks statistical reporting, reference existence, image reuse and protocol completeness on every submission has built a rule applied to every item, and the split predicts that does work. Neither reframing has been tested.

Speculative The frontier risk specific to this decade is contamination of the evidence base rather than pollution of any one journal. Retracted papers keep being cited, usually with no acknowledgement of the retraction, and citation counts to retracted work accumulate for years after the notice appears. Frontier A fabricated result that enters a systematic review, a clinical guideline or a model training corpus stops being a bad paper and becomes an input, and the correction machinery described in this brief operates on papers rather than on inputs. Handwave How much of the current literature is synthetic is unknown and probably unknowable with present instruments; every number in circulation for it extrapolates from detected cases to undetected ones with no measured detection rate to anchor the extrapolation, and this brief declines to print one.

4 · Technological bottlenecks

Established The first bottleneck is that the machinery's running cost is one of the best-measured facts in the field, and the one attempt to reduce it by redesigning the machinery made it worse. Herbert and colleagues observed the 2012 Australian NHMRC round: 3,727 proposals, 38 days of average preparation for new proposals and 28 days for resubmissions, 20.5% success — roughly 550 working years and AU$66 million in salary, with the equivalent of “some four centuries of effort” going into unsuccessful applications. Established NHMRC then streamlined, cutting online form fields from 180 to 68 and applications from about 100 pages to about 50. Barnett and colleagues surveyed before (2012: 446 researchers, 685 applications) and after (2014: 236 researchers, 440 applications). The all-application mean rose from 34 working days (95% CI 33–35) to 38 (37–39), p<0.001 — a different quantity from Herbert's 38-day figure, which is new proposals only — and total researcher time rose from 547 working years (535–559) to 614 (600–629), an increase of 67 working years. Success was 21% in both periods. The authors' explanation is that effort tracks the expected value of the prize (about A$130,000) and the level of competition, not the length of the form. This is the cleanest before-and-after test of the framing's own prescription in the record, and it went the wrong way.

Frontier The second bottleneck is that the total cost of reviewing cannot be stated as a single number, and the reason is a denominator dispute rather than a measurement error. Aczél, Szaszi and Holcombe estimate reviewers spent more than 130 million hours on peer review in 2020 — about 14,932 years — worth US$1.51 billion in the United States, US$626 million in China and US$391 million in the United Kingdom, on assumptions of six hours per review and 4.7 million articles, and they call these under-estimates. A separate survey by LeBlanc and colleagues (354 respondents, 33 countries, 44% Canadian) found a median of 4 hours for an initial review and 2 for a re-review, about 30 hours a year per reviewer, US$1,272 per reviewer per year, and a global cost of US$1.1–1.7 billion on Scopus article counts or US$6 billion on Dimensions counts. Established Do not average those. The spread from $1.1bn to $6bn is entirely a function of which publication database is treated as the world's literature, and naming that is the finding.

Established The third bottleneck is that national research assessment costs a measurable fraction of what it distributes and demonstrably induces gaming. The Stern review put REF 2014 at £246 million total — £212m borne by the higher education sector including £55m for the impact element, £19m in panellists' time and £14m to the funding bodies — 133% more than RAE 2008, to distribute about £2 billion a year of quality-related funding across roughly 190,000 outputs and 6,975 impact case studies. Stern also documents what the exercise induced: fractional contracts signed shortly before the census date, deliberate exclusion of staff and outputs, and submission sizes managed to stay below impact-case-study thresholds. An assessment machine large enough to be worth gaming is gamed, and the gaming is documented by the government's own reviewer.

Established The fourth bottleneck is conceptual and it defeats most public argument in this area: four different quantities are routinely merged into one. Retractions are about 0.2% of papers and roughly 10,000 a year. Image-manipulation prevalence is 3.8% in a 20,621-paper screen. Red-flagged potential fakes are about 11% in a preprint with a 37% false-alarm rate. Machine-flagged paper-mill candidates are about 10% of 2.6 million cancer papers in a news report of a 2026 study. These are four measurements of four things on four populations by four instruments. Cite them separately or not at all. Established The same failure mode applies inside enforcement statistics: published Federal Register findings and cases closed are not the same count, and the ratio between them is not computable from the published categories.

Established The fifth bottleneck is that “the replication rate” is criterion-dependent by a factor of two inside a single project. In preclinical cancer biology the success rate moves from 40% on a combined three-of-five rule to 62% meta-analytically to 80% for null effects, and the measured band across fields runs 36% to 62%. Any figure quoted without the criterion attached is overclaiming. Frontier That is not a reason to dismiss the finding: 92% of replication effects in cancer biology were smaller than the original and the median was 85% smaller, and effect-size shrinkage is consistent across every project in this brief regardless of which significance criterion is applied.

Established A further bottleneck is that the reproducibility rules already in force specify deposit and not executability, and that difference is the whole of the effect. The National Academies separated the two quantities the field had been conflating: reproducibility, the same result from the same data and code, and replicability, a consistent result from new data. The report declined to state a single replication rate for science, on the ground that the fields sampled are not a sample of science and that non-replication has legitimate sources other than error. Established That refusal is more useful than any number it could have printed, because it moves the burden onto per-field measurement. Frontier Where an executability condition has been tested against real deposits it fails at the first step: none of the 193 cancer-biology experiments in the Reproducibility Project could be run from the published description alone, and 67% of those eventually run needed protocol modification after contact with the original authors. A data-availability statement is not an executable protocol, and current mandates require the former.

Frontier The last bottleneck is capacity: replication is a service nobody is funded to supply at scale. Every large replication effort in the record was a one-off grant-funded consortium rather than a standing capability, each running for years against a fixed and pre-chosen set of studies, and each ending when its grant did. Speculative A standing replication service is the obvious institutional answer and has never been costed in public — not by a funder, not by a national academy, not by a publisher — so the argument against it is conducted entirely on assumed prices. Handwave Whether it would cost less than the cleanups it displaces is the comparison nobody has published.

5 · Research dependencies

Established Nothing on this map produces a result this brief waits on, and no typed depends-on edge is claimed. The evidence base here is administrative rather than experimental: the grant scores, the retraction counts, the allegation totals and the replication outcomes all already exist in institutional records. What is missing is not a discovery but a decision by the institutions holding the data to link two columns and publish the comparison. Established Three such decisions are recorded as typed institutional requirements in section 13: a funder publishing outcomes for lottery-allocated against panel-allocated grants, a national register of misconduct investigations with durations and outcomes, and an enforced condition that a deposited protocol be executable by a third party. None is a research result. Each is something a funder, a ministry or a journal could choose to supply and has so far chosen not to.

Frontier What this brief waits on from research is narrow and specific. First, a specification-robust settlement of the Fang-versus-Li-and-Agha dispute — the same 100,000-grant population run under both specifications with the conditioning set stated, which is an afternoon's work for anyone holding the NIH administrative file and has not been published. Second, a design that separates the rule effect of registered reports from the selection effect, which requires random assignment of format rather than a further observational census. Third, a current national measurement of institutional misconduct-investigation capacity to replace a survey from 2009. Speculative Fourth, and least tractable: an outcome measure for research allocation that is not a citation count. Every allocation result in this brief — the r² of 0.0078, the Li and Agha elasticities, the HHMI comparison — is denominated in publications, citations or patents, which are themselves outputs of the machinery under evaluation. That circularity is unresolved and nobody in this literature claims to have resolved it.

6 · Required experiments

Established The highest-value experiment in this subject has already been run and simply not analysed. At every funder using partial randomisation, assignment above the quality threshold is random by construction. That is a genuine randomised controlled trial of allocation mechanism, executing on live money at a dozen institutions since 2013, with the treatment assignment recorded and the outcome data — publications, citations, follow-on funding, personnel trained — accumulating in the funders' own reporting systems. The comparison requires no new grant, no new consent and no new instrument. It requires a funder to publish it. The British Academy states that it will “look more closely at any changes in the outcomes from the awards” only during the extended 2025–2028 period; the Volkswagen Foundation's accompanying research was a qualitative triangulation of document analysis, surveys and interviews and reported “high level of acceptance” rather than an outcome differential. Frontier The design constraint worth stating is power: Explorer Grants are about 2% of one national funder's spend, so a single scheme will not reach it. Pooling the randomised arms across the Research on Research Institute's catalogue would.

Frontier Second: randomise submission format at a journal. Assign eligible submissions to registered-report or standard track and compare the proportion of supported hypotheses, effect sizes and subsequent replication. This is the experiment Scheel and colleagues explicitly say they did not have, it separates the rule from the volunteers, and it is the only way to know whether mandating the format would move a literature or merely add a review stage. None of the 278 journals offering the format has run it.

Frontier Third: convert Bik's between-journal comparison into a before-and-after. The 40-fold spread in image-duplication prevalence between a journal that screens at acceptance and one that does not is cross-sectional, so journals that screen may differ in every other way. Journals adopt screening on datable schedules; comparing the same journal's duplication rate in the years either side of adoption, with a matched non-adopting control, would turn the strongest pro-checkpoint result in this brief into a causal one. Established The same design applies to double-blinding, which several journals have adopted on announced dates.

Established Fourth: replicate the seeded-error trial on modern manuscripts, and extend it to the checkpoint alternative. The 607-reviewer BMJ trial is from 2008. Rerunning it establishes whether detection has moved in eighteen years; running a third arm in which a statistical checklist is applied mechanically to every manuscript, against a fourth in which reviewers are trained, would test the brief's structural claim directly on a single outcome. Frontier Fifth, and cheapest: audit adherence at scale. Claesen and colleagues audited 27 analysis plans and found 93% deviation and 36% non-disclosure. Repeating that on a thousand plans from the current cohort would establish whether the format's binding has improved as it has spread, which is the single number that decides whether adoption is worth pushing.

7 · Engineering requirements

Established Three findings in the record cut for the framing, they are the strongest evidence the machinery has, and they share a structure that is the whole argument of this brief. Each is a rule applied mechanically to every item passing a point, and none is an improvement to anyone's judgement.

Established First, screening images at acceptance produces a 40-fold difference between journals. Bik, Casadevall and Fang's census of 20,621 papers found inappropriate image duplication in 0.3% of papers at the Journal of Cell Biology, which screens images at acceptance, against 12.4% at the International Journal of Oncology, which does not — with an overall negative correlation between duplication rate and impact factor (p = 0.019) over 2005–2014. A mechanical check applied to every accepted paper produces an order-of-magnitude difference in a quantity that referee judgement had been failing to catch for a decade.

Established Second, blinding demonstrably removes identifiable biases, and the result is more interesting than the version usually quoted. Giannakakos, Karanfilian, Dimopoulos and Barmettler's systematic review of 29 comparative studies of single- against double-blind review found, in the highest-quality tier, that single-blind review produced more favourable outcomes for authors at top institutions, for famous authors and for papers from high-HDI and English-speaking countries, and that double-blind review equalised these. Established Gender is where the clean story fails: five high-quality studies found no significant improvement under double-blinding. And the review records one study in which established authors benefited more from double-blinding than junior ones — which cuts against the assumption that hiding names is uniformly levelling. Blinding removes an identifiable bias. It does not demonstrably help disadvantaged groups, and it is not guaranteed to move in the direction its advocates expect.

Established Third, panels do agree about proposal quality when the range of quality is real, and the correction that establishes this is the most important statistical result in the subject. Across the full range of applications, single-rater reliability is 0.34–0.37 and three-rater reliability is 0.61–0.64. That is a working instrument. The engineering conclusion follows directly: the reliability is a property of the range, so a system that spends panel time discriminating inside an already-selected pool is operating its instrument outside the range in which the instrument works. A threshold-plus-lottery design is the mechanical response to that fact, which is why funders reached for it.

Established A fourth result belongs with these, and it is the one that most embarrasses the formal machinery: the research community can forecast replication, and the published machinery did not act on the forecast. In the economics replications, survey beliefs correlated with replication outcome at r = 0.52 (p = 0.028). Frontier The prediction markets run alongside them did not reach significance — r = 0.30, p = 0.232 — so the informative instrument was the simple survey, not the market, and the frequent claim that markets outperformed surveys here is backwards. Market prices averaged 75.2% and surveys 71.1% against an observed 61.1%: both over-optimistic in level, with the survey carrying the rank information. In the Nature and Science social-science set, peer beliefs were likewise “strongly related to replicability”. Aggregate peer judgement contained real signal about results that the same peer community had already published.

Established On the funding side, the best-identified evidence that a governance design changes scientific behaviour is a comparison of two funders, and what it changed was the shape of the output distribution rather than its rank. Azoulay, Graff Zivin and Manso matched 73 HHMI investigators appointed 1993–1995 against 348 NIH-funded early-career prize winners using propensity weighting and semi-parametric difference-in-differences. HHMI's long horizon, tolerance of failure and rich feedback produced 39% more publications overall, 55% more in the top 5% of citations and 97% more in the top 1% — while also producing 35% more papers falling below the investigator's own pre-appointment citation baseline and a 10% reduction in keyword overlap between pre- and post-appointment work. The design bought more hits and more misses together. That is the point, and it is the rare quantified case of a governance choice changing what gets attempted.

Frontier The ARPA model is thinner than its reputation, and the document usually cited for it says so itself. The National Academies' 2017 assessment of ARPA-E is the completed statutory six-year review, not an interim one, and its finding is affirmative: ARPA-E “is making progress toward its statutory mission” and “is not failing, or on a path to failing”. It reports that roughly half of supported teams had published peer-reviewed results, about 13% had obtained patents and one quarter had received follow-on funding. Established What is interim is the impact horizon, not the document. The committee stated that “few data were available for this study regarding ARPA-E's impact on energy technologies or the sector as a whole”, that six years captures only intermediate impacts against a multi-decade technology cycle, that a three-year project term “is too short to expect a technology to move from concept to market”, and recommended that ARPA-E build a measurement framework. There is no counterfactual and no comparison group, and the committee says so.

Established Priority-setting completes the picture: a funder can show its allocation tracks burden and cannot show the allocation changed the burden. NIH publishes its own funding-versus-burden comparison across 74 disease categories against Global Burden of Disease data, with the explicit caution that there is no comprehensive standard approach to measuring burden across diseases. Nimgaonkar and colleagues, across 38 matched categories, found 2017 funding tracked 2017 disability-adjusted life years reasonably (R² = 0.309, p = 0.0003) and 2007 funding tracked 2017 burden about as well (R² = 0.304) — but cumulative 2006–2017 spending had no significant relationship with the change in disease burden over the same period (F(1,36) = 0.199, p = 0.65). Frontier Allocation is proportionate; the return on allocation is not demonstrated. That is a null on a hard question with heavy confounding, not evidence that funding does not work.

8 · Adjacent technologies

The boundary that matters most runs against Scientific Advisory Institutions, and it is worth stating in both directions because the two slots are otherwise easy to merge. That brief is about advice reaching a decision-maker outside science: a chief scientific adviser briefing a minister, an emergency advisory group responding to a government, an assessment summary written for policymakers, a regulator's expert committee. Its unit of analysis is the advisory relationship — mandate, independence, the transmission of uncertainty, and what happens when the advice is ignored. This brief is about how science governs itself, with no external decision-maker in the loop: a study section scoring a grant, an editor deciding what to publish, a funder choosing a lottery over a panel, a journal setting a registered-reports policy, an integrity office deciding whether to open an investigation, a university adjudicating an allegation, an assessment exercise distributing block funding. Its unit of analysis is the allocation or gatekeeping rule and its measured consequences. Two cases look like they cross the seam and do not. The National Academies reports used here — on research integrity, on reproducibility, on ARPA-E — are academies reporting on research governance, and belong here as evidence about the machinery; they belong to the other slot only if the question is how an academy report reaches a legislature. And a funder auditing its own allocation against disease burden is a funder auditing itself, not an institution briefing a health ministry.

Against Failed Technologies the relationship is complementary and the correction runs one way. That page identifies a documentation asymmetry — successes are written up, failures are dispersed, and nobody's career is advanced by a thorough account of a project that did not work — and names registered reports once, in a list of remedies alongside published negative results and programme post-mortems, with the diagnosis that the obstacle is incentives rather than difficulty. It cites no empirical registered-reports literature. This brief treats registered reports instead as a measured intervention with a known effect size and a known adoption ceiling: a 52-percentage-point drop in supported hypotheses, an explicit non-experimental caveat, 278 journals and at or below 1% uptake in most fields, and 93% deviation from plan in the earliest audited cohort. That partly answers the other page: the incentives diagnosis is right, and the empirical record shows that naming the remedy has not moved adoption in thirteen years.

Elsewhere on this map: Institutional Design, where the incentive-compatibility problem measured here as reviewer effort tracking prize size is stated generally; Artificial Scientists, which inherits both the replication band and the screening-at-a-checkpoint result as the quality floor any automated pipeline has to clear; Technocracy and Democracy, on the legitimacy of expert self-regulation; Future Public Administration, whose finding that evaluation is discretionary explains why funders trial mechanisms and do not evaluate them; and Long-Term Institutions, which faces the same problem of an institution that cannot be scored on outcomes within its own tenure.

Outside it: scientometrics and the economics of science, which supply the allocation estimates; metascience and the reproducibility literature, which supply the replication projects; statistical methodology, which supplied the range-restriction correction that changed the reliability answer; and forensic image and text analysis, which now supplies most of the fraud detection that works.

9 · Institutional requirements

Established The institution that most obviously constrains this subject is a statutory integrity office operating at a scale two to four orders of magnitude below the problem it names. Against more than 10,000 retractions in a single year and an image-manipulation prevalence of 3.8% in a 20,621-paper screen, the United States integrity office published two Federal Register findings of research misconduct in 2025 against a typical run rate of about ten, while recording 446 allegations and 177 closures in the same year and a two-digit findings count in its own closure charts — a different quantity from the published findings, and one no ratio should be computed across. Frontier A departmental spokesperson said the investigative division “continues to carry out its oversight responsibilities”, and a former official has noted that a published finding is not the only possible outcome of a case. Take the most generous reading available and a statutory federal body publishing single-digit findings still sits against a retraction record of ten thousand papers a year.

Established The weakest institutional link is the one that does most of the work: the employing institution investigating its own staff. The National Academies' account records inconsistent internal definitions, faculty committees without relevant expertise, administrators handling investigations as a side duty, and a survey in which 97% of research integrity officers identified fewer than half the appropriate actions in misconduct scenarios. That survey is seventeen years old and there is no current national dataset on investigation duration or the fraction of allegations resulting in any action. The gap is recorded here rather than filled.

Established Meanwhile the function that actually detects large-scale fraud has no governance status and is largely unpaid. The Retraction Watch leaderboard is topped by Joachim Boldt (240 retractions), Yoshitaka Fujii (172), Yoshihiro Sato and Hironobu Ueshima (124 each) and Ali Nazari (104) — cases broken by outside statistical and forensic scrutiny, not by referee reports or institutional process. The database behind it has since been acquired by Crossref and made openly available. Post-publication scrutiny by outsiders has become load-bearing infrastructure with no formal standing, no funding line and no mandate.

Established The strongest against-interest evidence in the subject is a publisher documenting the failure of its own review process and pricing it. Wiley's interim chief executive put the Hindawi cleanup at US$35–40 million in lost revenue for the fiscal year. The publisher paused special issues from October 2022, built a paper-mill indicator checklist, brought in outside vendors and at least two evaluators per paper, and banned several hundred guest editors. A firm paying tens of millions to say that its own peer review did not work is worth more than any number of external critiques, and it is again a checkpoint — a checklist applied to every submission — that replaced the judgement that failed.

Established Several of the best sources here are interested parties and it matters which way. The funders running randomisation trials report their own acceptability and diversity figures and have published no outcome comparison; the organisation that ran the cancer-biology replication project also advocates the reforms the project recommends; the integrity office reports its own caseload. Established Two bodies of evidence run against their producers' interest and are weighted more heavily for it — Wiley's cleanup accounting, and the National Academies stating that institutional misconduct investigation is inadequate on evidence the Academies themselves compiled.

Established The constraint that binds is therefore not knowledge but publication of what institutions already hold. Three specific decisions would change what can be said in this brief, none of them a research result: a funder publishing outcomes for lottery-allocated against panel-allocated grants; a national register of misconduct investigations with durations and outcomes; and an enforced condition that a deposited protocol be executable by a third party, which the cancer-biology attrition shows no journal currently enforces. All three are recorded as typed requirements in section 13.

10 · Ethical & societal considerations

Established The allocation machinery has a measured distributional record and it is not neutral. Applications with white principal investigators were 1.7 times more likely to be funded than those with African-American or Black principal investigators across two separate NIH periods, and after a full model including publication record, 6.9 percentage points of a 13.4-point raw gap remained unexplained. Frontier Topic choice accounted for just over 20% of the gap, which is itself a governance fact rather than an applicant fact: the topics that attract lower scores are chosen inside a system that signals what will be funded.

Established The obvious remedy does part of the job and not the part usually claimed for it. Double-blinding demonstrably removes institution, fame and country effects. Five high-quality studies found no significant improvement on gender, and one found established authors benefiting more from blinding than junior ones. A rule that removes an identifiable bias is not the same as a rule that helps the people the bias was assumed to harm, and conflating the two oversells the only intervention in this area that clearly works.

Frontier Randomisation has a distributional record too, and it is the funder's own. At the British Academy, applicants of Asian or Asian British background rose from 13% to 19% of applications and 11% to 17% of awards; Black or Black British applicants from 3% to 4% and 2% to 3%; awards outside the concentration of established research universities moved from 85% to 87%. These are association figures from a funder describing its own trial, over a period in which application volume roughly doubled. They are the best distributional evidence lotteries have and they are not causal estimates.

Established The cost of the machinery falls on the people it evaluates, and most of it is destroyed. A single national grant round consumed roughly 550 working years and AU$66 million of salaried time, with the equivalent of four centuries of effort going into applications that failed — and the reform intended to reduce that burden increased it by 67 working years. Established Globally, more than 130 million hours a year go into reviewing, unpaid and largely unrecorded, on top of it.

Frontier And the automation of fraud detection creates an accusation problem that the technical papers acknowledge and the reporting rarely does. A red-flag rule with a 37% false-alarm rate is a screening instrument, not a finding about any individual paper or author; a machine classifier flagging 261,245 papers carries an unresolved objection that it may be detecting the writing patterns of non-native English speakers rather than fabrication. Speculative Applied as a checkpoint to a submission the error is a cost; applied as a public list of names it is a defamation with a stated false-positive rate. The checkpoint architecture this brief argues for is defensible precisely because it acts on items rather than on people, and that distinction is doing ethical work that should be stated rather than assumed.

11 · Civilizational implications

Established The terminal position is that the framing is wrong about the direction of the fix, not about whether a fix exists. Every intervention in this record with a demonstrated effect is a checkpoint applying a rule to every item that passes it: screen every accepted paper's images (0.3% against 12.4%), hide the author's name (institution, fame and country effects removed), fix the publication decision before the data exist (96.05% falling to 43.66%), draw lots among everything above a threshold. Every intervention that fails is an attempt to improve a fine-grained judgement: better review criteria, referee training (2.58 of 9 seeded errors, barely improved), shorter application forms (burden up 67 working years), more review stages. More of the same machinery means more judgement. The measured record favours less judgement and more rule.

Established The reason this is a structural claim rather than a list is that the reliability evidence explains it. Panel assessment is a working instrument across the full range of proposal quality — 0.34–0.37 single-rater, 0.61–0.64 for three raters — and stops working inside a pre-selected pool. A rule is a judgement made once, in advance, by whoever designs it, and then applied without further discrimination. That is exactly the operation an instrument with range-dependent reliability can perform well, and exactly what asking twelve reviewers to rank twelve good proposals is not.

Frontier The uncomfortable implication is that the scalable response to fraud is automated screening, and automated screening is the checkpoint architecture with a false-positive rate. Enforcement will not scale: a statutory office publishing two findings a year cannot process ten thousand retractions, and institutional self-investigation has been documented as inadequate for seventeen years without being fixed. Machine screening of millions of papers is the only mechanism in the record operating at the volume of the problem, and its published accuracy claims come with objections about what the classifier has actually learned. A civilisation that governs its knowledge production by rule rather than by judgement is delegating to whoever writes the rule and whoever trains the classifier.

Speculative The counterweight belongs in the terminal position rather than in a footnote. The National Academies declined the “crisis” framing entirely, and treats inconsistency between studies as potentially a precursor to discovery rather than a defect to be engineered out. A literature in which 96% of papers support their hypothesis is not a healthy literature; a literature in which 44% do is not obviously an unhealthy one. Handwave Whether the correct long-run rate is 44%, 60% or something else is not a measurable quantity and nobody in this field has proposed a way to make it one. That is the step where any argument for a target replication rate does its work by assertion.

Frontier What is not in doubt is that the community can see further than its machinery acts. Peer surveys forecast which results would replicate at r = 0.52 in the economics set, and peer beliefs tracked replicability in the social-science set — on results the same community had already published. The information was inside the system and the system had no rule that used it. Every proposal in this brief is, at bottom, a proposal to convert judgement the community already holds into a rule the machinery can execute.

12 · Timelines

These horizons track funder trial completion dates, adoption curves and enforcement capacity rather than technology:

  • 10 yr: Frontier Three dated things resolve. The British Academy's extension runs to 2028 and commits to examining award outcomes in that period, so either the first lottery-versus-panel outcome comparison exists by then or the documented absence passes eighteen years. The Swiss National Science Foundation's portfolio-wide random tie-breaking, running since late 2021 across about CHF 880 million, accumulates the largest randomised arm in the record whether or not anyone analyses it. Speculative Registered-report adoption is the slowest-moving number here: from 278 journals and at or below 1% uptake in most fields, the plausible range at ten years runs from unchanged to a doubling, and nothing in the thirteen-year record supports more. Frontier Expect machine screening at submission to become standard at large publishers within this window, because the Hindawi cleanup priced the alternative at $35–40 million for one imprint. Frontier Expect the fixed-string era of fabrication detection to close inside this window, because a named signal is cheap for a mill to remove: detection counts will keep rising while saying less about prevalence.
  • 25 yr: Speculative The plausible split is that checkpoint interventions keep accumulating evidence because they produce clean comparisons, and judgement interventions stay unevaluated because they do not. Speculative A second possibility, held by serious people and not established, is that the allocation problem is dissolved rather than solved — if research costs fall enough that threshold-plus-lottery becomes the default above a low bar, the ranking problem that consumes 130 million hours a year stops being posed. Speculative A third is that automated screening reaches accuracy high enough to make the integrity office a rule-setter rather than an investigator. Handwave Choosing between these is forecasting institutional adoption, which nothing in this brief measures.
  • 50 yr: Speculative If the checkpoint-versus-judgement split is real, the long-run shape of scientific governance is a small number of enforced rules at a few defined points — protocol deposit, materials availability, image and statistical screening, blinded assessment, randomised allocation above a threshold — with expert judgement concentrated in writing the rules rather than in applying them case by case. Speculative The alternative is that the circularity named in section 5 binds first: every outcome measure in this literature is an output of the machinery being evaluated, and a governance system optimised against its own outputs is a system with no external check. Handwave Both are extrapolations from a pattern across nine studies.
  • 100 / 250+ yr: Handwave Beyond useful forecasting. The only long-baseline data point available is that organised peer review of journal submissions is a twentieth-century administrative arrangement rather than a feature of science, and that it has already been redesigned several times within living memory. Handwave One institution's history in one century is a story, not a base rate.

13 · Technology tree & dependencies

  • Depends on Nothing on this map. This brief waits on no result another brief produces: the evidence base is administrative rather than experimental, and every missing comparison could be published tomorrow from records institutions already hold. No typed depends-on edge is claimed.
  • Requires (not on this map) Three institutional decisions, none of them a research result, and all three currently withheld. First, a funder publishing outcomes for lottery-allocated against panel-allocated grants. Lottery allocation is the one governance mechanism actually trialled on live money — the Health Research Council of New Zealand since 2013, the Volkswagen Foundation's Experiment! from 2017 to 2021, the Swiss National Science Foundation across about CHF 880 million from late 2021, the British Academy since 2022 and extended to 2028, and more than a dozen funders in the Research on Research Institute's catalogue with 11 added since 2022. Assignment above the quality threshold is random by construction, so the trial has already run. Thirteen years after the first randomised round, no funder anywhere has published an outcome comparison: the New Zealand funder's own page on its lottery carries none, the Volkswagen accompanying research was a qualitative triangulation reporting acceptance rather than an outcome differential, and the British Academy states it will examine award outcomes only during 2025–2028. What exists is acceptability (63% for and 25% against among applicants to the randomised scheme, falling to 40% and 37% for other schemes, and splitting 78% against 44% by whether the respondent was funded) and the funder's own distributional figures. That documented absence is the single largest gap in the subject. Second, a national register of misconduct investigations and their outcomes. The best evidence on institutional investigative capacity is a 2009 survey in which 97% of research integrity officers identified fewer than half the appropriate actions in misconduct scenarios; it is seventeen years old, and there is no current national dataset on investigation duration or the fraction of allegations resulting in action anywhere. Published Federal Register findings and cases closed are separately reported and are not the same quantity, so nothing can be computed from what is published. Third, an enforced condition that a deposited protocol be executable by a third party. Journals and funders already require methods sections, data availability statements and materials sharing; none of 193 cancer-biology experiments was described in sufficient detail for a protocol to be written without contacting the authors, 68% yielded no data at all, 0% had completely described protocols, and 67% of the experiments actually run required protocol modification. The rule exists and the enforcement does not. All three are things a funder, a ministry or a journal could choose to supply and has so far chosen not to, and so is the fourth. Fourth, a publisher disclosing integrity-screening hit rates and false positives. Submission screening now runs at scale inside large publishers and through shared industry infrastructure, and it is the checkpoint intervention with the most money behind it; what is published about it is the existence of the tool, not its operating characteristics. Without a hit rate, a false-positive rate and a denominator of submissions screened, no editor can say what share of rejected submissions was rejected wrongly. That is a disclosure decision rather than a research problem, and it sits with the same handful of firms that already hold the data.
  • Enables In principle any brief on this map that assumes a functioning research base inherits these findings, and Artificial Scientists inherits them most directly — the replication band of 36–62%, the finding that zero of 193 cancer-biology experiments were reproducible from the paper alone, and the screening-at-a-checkpoint result are the quality floor any automated research pipeline has to clear and the training corpus it would learn from. No typed enabling edge is claimed, because the relationship has never been measured in either direction.
  • Adjacent Scientometrics and the economics of science, which supply the allocation estimates; metascience and the reproducibility projects; statistical methodology, which supplied the range-restriction correction that changed the reliability answer; forensic image and text analysis, which now supplies most of the fraud detection that works; and within this map Institutional Design, Scientific Advisory Institutions and Artificial Scientists.

14 · Common misconceptions & speculative claims

Frontier “Grant peer review scores do not predict anything.” This is the most-quoted claim in the subject and it is one specification of a two-specification dispute. Regressing raw percentile on raw output across 102,740 funded R01 grants gives r² = 0.0078 and AUC 0.54. Conditioning on applicant and field characteristics across more than 130,000 R01 grants gives a one-standard-deviation worse score associated with 15% fewer citations, 7% fewer publications, 19% fewer high-impact publications and 14% fewer patents. Established The honest statement is that the score carries real information which is small relative to what is already known about the applicant, and that it does not discriminate inside the band where funding is actually decided. Quote either result alone and you have misdescribed the literature.

Established “Grant peer review is arbitrary.” Not supported. The zero-agreement result is real and it is range-restricted: across the full quality range, single-rater reliability is 0.34–0.37 and three-rater reliability 0.61–0.64, and zero estimates appear only once fewer than about 70% of top-quality and 45% of bottom-quality proposals are retained. “Peer review cannot tell good from bad” is false. “Peer review cannot rank within the fundable pool” is the supported claim, and it is the one that matters, because that is where the money is decided.

Established “Lotteries produce better research.” There is no evidence for this in either direction, and the absence is documented rather than inferred. Thirteen years after the first randomised round, no funder has published a comparison of research outcomes between lottery-allocated and panel-allocated grants. What is established is that lotteries are acceptable to applicants who applied to them — 63% for and 25% against, collapsing to 40% and 37% for other schemes and splitting 78% against 44% by whether the respondent was funded — and that partial randomisation coincided with more applications and a more diverse applicant pool at one funder reporting its own trial. Speculative The correct statement is that the mechanism has been trialled and not yet evaluated. The case for it rests on the reliability argument, not on an outcome measurement.

Established “Replication failures are explained by context and hidden moderators.” Empirically closed for the effects tested. Many Labs 2 ran 28 effects across 125 samples, 36 countries and 15,305 participants and found variability in observed effect sizes attributable more to the effect being studied than to the sample or setting. Sample and setting explained little.

Established “The retraction rate tells you the fraud rate.” Four different quantities, routinely merged: retractions (above 0.2% of papers, about 10,000 in 2023), image-manipulation prevalence (3.8% in a 20,621-paper screen), red-flagged potential fakes (11% in a preprint with a 37% false-alarm rate), and machine-flagged paper-mill candidates (about 10% of 2.6 million cancer papers in a news report of a 2026 study, with the lead author calling it an underestimate and a commenter raising a language-bias objection). Cite them separately or not at all. Established National retraction rates per 10,000 articles are Saudi Arabia about 30, Pakistan 28.1, China 24.9, Russia 23.5 — frequently reported with the last two transposed.

Established “Research-integrity enforcement is scaling.” The published annual reports give allegation counts (713 in 2024, 446 in 2025) and closure counts (119 and 177), but the closure categories do not resolve into a clean allegations-to-action ratio and none should be computed from them. The one clean number is two Federal Register findings published in 2025 against a typical ten a year — and that is a count of published findings, not of cases closed, and not the two-digit findings quantity recorded in the annual report's own closure charts. Established The strongest documented weakness is that 97% of research integrity officers identified fewer than half the appropriate actions in misconduct scenarios, from a survey published in 2009. There is no current national dataset on investigation duration or outcome.

Frontier “The ARPA model is validated” — and its mirror image, “the assessment declined to say so.” Both are wrong. The National Academies' 2017 review is the completed statutory six-year assessment and it found that ARPA-E “is making progress toward its statutory mission” and “is not failing, or on a path to failing”. Established What it does not contain is a counterfactual. It reports intermediate outputs — about half of teams publishing, about 13% patenting, one quarter with follow-on funding — states that “few data were available” on impact on energy technologies or the sector, notes that a three-year project term is too short for concept-to-market, and recommends that the agency build a measurement framework. An affirmative finding on mission progress is not a demonstration that the model outperforms the alternative, and the committee did not claim it was.

Frontier “Prediction markets showed that scientists knew which results would replicate.” Backwards. In the economics replications the surveys were the informative instrument — r = 0.52, p = 0.028 — while the prediction markets did not reach significance at r = 0.30, p = 0.232. Both were over-optimistic in level, forecasting 75.2% and 71.1% against an observed 61.1%. Established The finding survives the correction and is stronger for being stated properly: aggregate peer judgement carries real rank information about replicability, and the cheap instrument carried it.

Established “Peer review costs about X billion dollars a year.” No single figure is supportable. The published range runs US$1.1–1.7 billion on a Scopus denominator to US$6 billion on a Dimensions denominator, against a separate estimate of 130 million hours and roughly US$2.5 billion for the United States, China and the United Kingdom combined. The disagreement is about which publication database counts as the world's literature. Report the range and name the cause; do not average. Established The same rule applies to “the replication rate”: the band across fields is 36–62%, and within a single project the success rate moves from 40% to 80% depending on which of five criteria is applied.

Frontier “Registered reports are the fix, and adoption is just a matter of time.” The effect size is the largest in the subject and the adoption record is thirteen years of near-flat uptake: 278 journals, 7% in psychiatry and psychology, 34% in experimental psychology, at or below 1% in most fields. Frontier And where the format is used, 93% of audited analysis plans deviated and 36% disclosed no deviation at all. The remedy has been named repeatedly and naming it has not moved adoption, which is the point at which the incentives diagnosis in Failed Technologies stops being a hypothesis and starts being a measurement.

Speculative And the framing itself: “more of the same machinery” is a proposal to add judgement to a system whose judgement is the part that fails. Referee training moved detection from 2.58 to about 3 of 9 seeded errors. Shortening the application form raised total burden by 67 working years. More review stages have never been shown to help. Established Meanwhile every rule applied mechanically to every item has produced an effect large enough to see: a 40-fold spread on image screening, the removal of institution and fame effects on blinding, a 52-point drop in positive results on fixing the publication decision early. The framing is not simply wrong. It is wrong about the direction of the fix.