The Institute's method is the workback plan: start from what is currently considered impossible and walk back to the first checkable step. This page collects those first checkable steps. For each brief it records the experiment, observation, demonstration, measurement, natural experiment or policy result the brief itself treats as decisive, when the brief says it could happen, and the flag the brief attaches to the claim it would test.
The headline is the horizon column. Of 295 decisive results, 43 are already running and 21 are tied to a stated near-term window — but 214 are unscheduled: tests the briefs say could be run, that nobody has scheduled or funded. The frontier's bottleneck, on this evidence, is less often that the decisive test is impossible than that no one has bought it.
- Horizons
- 43 running · 21 near · 11 mid · 6 long · 214 unscheduled
- Types
- 131 experiment · 83 measurement · 22 observation · 25 natural experiment · 24 demonstration · 10 policy result
I — Physics & Propulsion
Gravity Modification
Whether any laboratory configuration modifies gravity, or whether the running precision tests keep returning nulls.
Take a claimed gravity-modification effect and vary the parameter the theory says it scales with to see whether the signal follows, as Tajmar's 2011 rerun did at 50 times the rotation rate; generalised by the 2024 three-balance vacuum search across the whole hypothesis class.
“The single most informative experiment anyone could run has already been run twice. Tajmar's 2011 rerun at 50 times the rotation rate is the model: take the claimed effect, vary the parameter the theory says it scales with, and see whether the signal follows. It did not.”
Artificial Gravity
Whether intermittent or continuous partial gravity in orbit prevents the physiological deterioration of weightlessness in humans.
A human-rated centrifuge flown in orbit: a 2.5-metre instrument spanning 0.01 to 2 g was built and never flown, and a station-integrated ergometric version was cancelled on structural grounds.
“Second: a human-rated centrifuge in orbit. This is the experiment that answers the question the field is actually asking, and it is the one that has been cancelled twice.”
Warp Drives
Whether any proposed warp metric can be sourced without violating the energy conditions.
Run an energy-condition solver over the field's own catalogue of metrics, computing the Einstein tensor and evaluating the conditions pointwise across observers rather than in a single frame; Warp Factory and, independently, warpax keep returning the same answer.
“The first is numerical relativity applied to the metrics themselves, and it has already produced this brief's decisive result. Warp Factory takes an arbitrary metric, computes the Einstein tensor, and evaluates the energy conditions pointwise across observers rather than in a single frame. Running it on the field's own catalogue found that every metric tested violates conditions its authors had believed it satisfied.”
Wormholes
Whether negative-mass compact objects or traversable wormholes exist at any detectable abundance.
A microlensing search for the negative-mass lensing signature - total or partial eclipse of the source in the umbra region and a Shapiro time gain rather than a delay - run against a quasar-lens survey to set abundance upper limits.
“Microlensing, which is the mature line. Safonova, Torres and Romero's 2001 analysis gives the signature precisely: an effective negative-mass lens produces a total or partial eclipse of the source in the umbra region, and a Shapiro time gain where an ordinary lens produces a delay.”
Reactionless Propulsion
Whether any claimed reactionless thruster produces a thrust that survives the standard controls.
Apply the five controls, cheapest first: orientation reversal, a ten-thousand-fold power attenuator, magnetic shielding of the feed lines with a documented residual, thermal isolation of the feedthrough with the surface-tension term measured, and publication of the full systematic budget by a group with no stake.
“The experiment that matters is already specified, because it is the one that has already killed every signal.”
Inertial Manipulation
Whether the claimed inertial-manipulation devices produce a real force or a measurement artefact.
Run the remaining cheap discriminators on the claimed devices: reverse the orientation and see whether the thrust reverses, and attenuate the drive power by four orders of magnitude and see whether the thrust falls.
“What would actually be discriminating is cheap and has been specified by a sympathetic assessor. Reverse the device's orientation and see whether the thrust reverses; attenuate the drive power by four orders of magnitude and see whether the thrust falls.”
Vacuum Energy Engineering
Whether the quantum inequalities bounding negative energy density survive a direct, purpose-built test against squeezed light.
A dedicated test of the Ford-Roman inequality against squeezed light: squeezing generated and characterised specifically to evaluate the inequality with a pre-registered sampling function, rather than a meta-analysis of other people's data with sampling functions chosen after the fact.
“A purpose-built experiment — squeezing generated and characterised specifically to evaluate the inequality with a pre-registered sampling function — would either substantiate the most important claim in this cluster or dispose of it. Nobody appears to have run one, and the omission is strange given what turns on it.”
Quantum Gravity
Whether the gravitational field is quantum - whether gravity can mediate entanglement between two masses.
A gravitationally induced entanglement experiment: two masses in adjacent interferometers, each in spatial superposition, interacting only gravitationally with a conducting plate screening the Casimir-Polder interaction, read out through spin correlations, with independent replication.
“Tier one, decisive-ish: a gravitationally induced entanglement signal. Two masses in adjacent interferometers, each in spatial superposition, interacting only gravitationally, with a conducting plate screening the Casimir–Polder interaction, read out through spin correlations, with independent replication.”
Negative Mass
Whether antimatter falls upward - whether any substance has negative gravitational mass.
ALPHA-g's direct free-fall measurement on magnetically confined antihydrogen, with a laser-cooled successor aimed at 1% precision against a theorists' target of one part in ten million.
“Direct measurement, and it is done. ALPHA-g released roughly 100 magnetically confined antihydrogen atoms over 20 seconds and recorded annihilation positions relative to the trap opening”
Space-Time Metric Engineering
Whether any engineered metric perturbation can be produced or detected.
Numerical relativity on a collapsing null-energy-condition-violating spacetime, producing computed gravitational waveforms that can be searched for in existing detector data - a search rather than a construction.
“This is the only experimental handle anyone has proposed, and it is a search rather than a construction — the authors themselves label the search application “rather speculative” and frame the value as understanding the stability of such spacetimes.”
Advanced Nuclear Propulsion
Whether low-enriched HALEU fuel can match HEU performance in a nuclear thermal rocket.
A comprehensive assessment of HEU against HALEU for nuclear thermal propulsion, backed by irradiation data on candidate low-enriched fuel forms, as the National Academies formally recommended.
“The single most consequential missing experiment is the fuel comparison the National Academies formally recommended. A comprehensive assessment of HEU against HALEU for NTP, backed by irradiation data on candidate low-enriched fuel forms.”
Fusion Spacecraft
Whether any fusion device can reach the specific power a spacecraft propulsion system requires.
Take a fusion device that produces net energy, weigh it including magnets, cryoplant, shielding, power conversion and radiators, and report kilowatts per kilogram.
“The experiment that would move this brief is a specific-power measurement, and nobody is running it. Take any fusion device that produces net energy, weigh it including magnets, cryoplant, shielding, power conversion and radiators, and report kilowatts per kilogram.”
Antimatter Propulsion
Whether a small number of antiprotons can catalyse a target burn releasing more energy than the antiprotons cost to make.
An antiproton-catalysed micro-fission or micro-fusion ignition at any scale, showing that a small number of antiprotons initiates a burn releasing more energy than they cost to produce.
“The decisive concept-level experiment is an antiproton-catalysed micro-fission or micro-fusion ignition at any scale. Demonstrate that a small number of antiprotons initiates a target burn that releases more energy than the antiprotons cost to make, and the catalysed branch becomes a research programme rather than a design study.”
Solar Sail Systems
Whether a full-scale solar sail can be deployed reliably enough to become an operational vehicle.
The test the ACS3 failure report asks for: high-fidelity full-scale deployment testing without gravity compensation, with full-visibility cameras, and off-nominal simulation done early rather than late.
“The single most productive thing anyone could do next is the test that report asks for: high-fidelity full-scale deployment testing without gravity compensation, with full-visibility cameras, and off-nominal simulation done early rather than late.”
Beam-Powered Propulsion
Whether high-power laser cost per watt is actually falling fast enough to make beamed propulsion economics credible.
Publish a defensible time series of cost per watt for high-power fibre lasers over the last two decades, which would either restore the projected eighteen-month halving or retire it.
“The experiment that would settle the most is not an experiment at all: publish a laser cost series. The concept's economics rest on a projected eighteen-month halving; the only contemporary spot price found for this brief is $100 per watt against a $0.01–0.05 requirement.”
Interstellar Probes
Whether an institution can fund, launch and then operate a spacecraft for fifty years.
Fly a precursor mission to several hundred astronomical units - an experiment on an institution rather than on hardware, since nothing about the physics of getting there is unknown.
“The decisive mission-level experiment is a precursor itself. Nothing about the physics of reaching several hundred astronomical units is unknown; what is unknown is whether an institution can fund, launch and then operate a spacecraft for fifty years.”
Alcubierre Metrics
Whether any Alcubierre-type metric satisfies the energy conditions when evaluated across all observers rather than in a single frame.
Run an all-observer energy-condition solver over the published back catalogue of warp metrics; it has been run twice by different people, with violations found in every metric tested.
“The first computational programme is all-observer energy-condition evaluation, and it settled the 2020–21 episode. Warp Factory computes the Einstein tensor for an arbitrary metric and evaluates the energy conditions pointwise across observer classes rather than in a single frame. Running it across the family found violations in every metric tested, including one its own authors had published as satisfying the conditions.”
Gravitational Wave Engineering
Whether the operating gravitational-wave detectors can serve as anything other than passive astrophysical instruments.
Search the operating detection network's existing data for the computed gravitational-wave signatures of exotic spacetime collapse - using instruments that already exist to look for somebody else's metric engineering.
“The one genuinely novel experimental idea in the subject uses instruments that already exist and looks for somebody else's device. If exotic spacetime collapse produces computable gravitational-wave signatures, then the operating detection network is incidentally an instrument for finding artificial metric engineering.”
Exotic Materials for Propulsion
Whether candidate nuclear-thermal fuel forms survive flowing hydrogen above 2700 K for a mission duration.
Hold uranium nitride kernels in a molybdenum-tungsten matrix, or in zirconium carbide, above 2700 K in flowing hydrogen for a mission duration, with carbon interaction with the kernel and thermal-expansion mismatch at the cladding instrumented rather than inferred.
“The decisive experiment in this brief is a hot-hydrogen fuel test at the mission condition, and it has not been run. Uranium nitride kernels in a molybdenum–tungsten matrix, or in zirconium carbide, held above 2700 K in flowing hydrogen for a mission duration, with the two named failure mechanisms — carbon interaction with the kernel and thermal-expansion mismatch at the cladding — instrumented rather than inferred.”
Electrodynamic Propulsion Concepts
Whether an electrodynamic tether can produce useful thrust in flight rather than only measured current.
Fly a tether that drives current against the motional electromotive force with onboard power in low Earth orbit, and measure the resulting orbit raising against the current and the field.
“The decisive experiment for the underused member is a thrust-producing tether mission, and it has never been flown. Drive current against the motional electromotive force with onboard power, in low Earth orbit, and measure the resulting orbit raising against the current and the field.”
Electrogravitics
Whether high-voltage electrode and field geometries produce any force beyond ion wind.
Sweep the geometry rather than testing one famous device, go to vacuum and vary the pressure, instrument to well below the claimed effect, and test the theory families' distinctive predictions - the template the 2024 campaign completed at up to 40 kV on balances resolving below a nano-newton.
“The experiment that settled the gravitic question has been run four times and the design is worth stating as a template. Sweep the geometry rather than testing one famous device; go to vacuum and vary the pressure; instrument to well below the claimed effect; and test the theory families' distinctive predictions rather than only the headline claim.”
Mach Effect Thrusters
Whether the Mach-effect thruster produces thrust at its claimed operating point under a protocol both sides agreed in advance.
A seven-item joint test: a pre-registered drive configuration, proponent-supplied tuned hardware, voltage at the claimed operating point, vibration isolation with an independent accelerometer channel, a same-session symmetry null, orientation reversal and a power-attenuator control, and resolution sufficient for a claim the proponents must first name.
“The outstanding experiment can be specified completely, from the dispute itself, which is why this brief is prescriptive rather than plaintive. Seven items. A pre-registered drive configuration agreed by both parties before the run: frequency, voltage amplitude, transformer, piezoelectric stack construction, reaction-mass geometry and settling protocol.”
Planetary Scale Energy Systems
Whether coherent phased-array power beaming works at the element count and aperture a planetary-scale system assumes.
Coherent phase control across of order 100,000 amplifiers on a kilometre-class transmitting aperture, holding a beam onto a multi-kilometre rectenna.
“The experiment that has not been run is the one at scale. Coherent phase control across of order 100,000 amplifiers on a kilometre-class transmitting aperture, holding a beam onto a multi-kilometre rectenna, is the untested step, and it is an array-engineering experiment rather than a power-conversion one.”
Space-Based Manufacturing
Whether the seed-crystal effect generalises, and so whether orbital manufacturing is a production business or a seed business.
Test whether flight-grown polymorphs are retained through terrestrial generations across a wide panel of pharmaceutically relevant molecules, as two of five did through ten generations and a third through nine.
“Second, and the highest-value experiment in the whole subject: does the seed-crystal effect generalise? Two of five molecules retained their flight polymorph through ten terrestrial generations and a third through nine.”
Deep Space Infrastructure
Whether the deep space network can carry a crewed lunar programme and flagship science at the same time.
Artemis I's unintended capacity test of the network, which could not serve everything at once and measurably degraded flagship science; the scheduled follow-on is Artemis III flying two crewed vehicles in cislunar space simultaneously on the same network.
“The decisive experiment in this brief has already run and it was not designed as one. Artemis I flew, the network could not serve everything at once, and flagship science was measurably degraded. That is a capacity experiment with a published result, and it is worth more than any modelling study of network loading.”
Precision Quantum Sensing
Whether thorium-229 can be operated as a true metrological clock at or below one part in ten to the eighteenth, settling its value for timekeeping and for fundamental-constant drift searches.
Operate thorium-229 as a closed-loop clock — nucleus excited on demand, laser locked to the transition, complete systematic error budget published — and compare it against strontium and aluminium-ion clocks for a year at or below one part in ten to the eighteenth.
“The decisive experiment is to turn thorium-229 from a measured transition into a working clock: excite the nucleus on demand, lock a laser to it, publish a complete systematic error budget, and hold a year-long comparison against strontium and aluminium-ion clocks at or below one part in ten to the eighteenth.”
Laboratory Astrophysics
Is the measured iron opacity enhancement at solar convection-zone conditions a property of iron, or an artefact of the platform that measured it?
An independent replication of the Sandia Z iron-opacity enhancement on a physically different platform, using a different driver and a different diagnostic chain.
“The decisive experiment is an independent replication of the iron opacity enhancement on a physically different platform. If a second facility, using a different driver and a different diagnostic chain, reproduces the measured opacity at solar convection-zone conditions, then laboratory astrophysics has demonstrated its strongest possible claim: that a terrestrial measurement corrected a stellar model.”
Quantum Materials
Is the parity measurement reported in industrial semiconductor-superconductor devices evidence of a topologically protected qubit, or of a non-topological state?
An independent laboratory with no commercial stake reproducing the topological parity measurement in a device it fabricated itself.
“The decisive experiment is an independent laboratory, with no commercial stake in the outcome, reproducing the topological parity measurement in a device it fabricated itself. Confirmation would establish that a protected qubit degree of freedom exists in a real material and would justify the largest single bet in quantum materials.”
Post-LHC Colliders
Can muon beams be cooled in all six dimensions, with reacceleration, in a repeating cell at the rate the design studies assume?
Build and operate a full six-dimensional muon ionization-cooling cell with radio-frequency reacceleration in a real beam, and measure the cooling factor per unit length against the design assumption.
“The decisive experiment for this whole field is a full six-dimensional muon ionization-cooling cell, with radio-frequency reacceleration, operated in a beam and measured.”
Hidden-Sector Searches
Does the QCD axion exist in the tens-to-hundreds of micro-electronvolt mass window that post-inflationary string-network simulations point to?
Scan the tens-to-hundreds of micro-electronvolt axion mass window at a sensitivity reaching the pessimistic QCD axion band, using a dielectric or plasma haloscope in a large-bore dipole above roughly nine tesla, or a next-generation helioscope.
“The decisive experiment is a scan of the tens-to-hundreds of micro-electronvolt axion mass window at a sensitivity that reaches the pessimistic QCD axion band.”
Programmable Metasurfaces
Does a reconfigurable metasurface deliver better coverage economics than an amplifying network-controlled repeater on the same site, once control overhead and calibration drift are counted?
An independent, co-sited, multi-month field trial of a reconfigurable surface against a network-controlled repeater and against no intervention, publishing per-user throughput distributions rather than spot measurements.
“The decisive experiment is an independent, co-sited, multi-month field trial comparing a reconfigurable surface against a network-controlled repeater and against doing nothing, on the same site, with published per-user throughput distributions.”
II — Space & Astronomy
Lunar Industry
Does a permanently shadowed region hold accessible volatiles in the quantity, depth and form that lunar ISRU economics assume?
A metre-class drill and a volatile-sensitive instrument suite operating inside a permanently shadowed region for at least a hundred days — VIPER's specification — answering where volatiles are, how easy they are to access, and how much is ice crystals versus mineral-bound.
“The single most informative experiment available is a metre-class drill and a volatile-sensitive instrument suite inside a permanently shadowed region, operating for at least a hundred days.”
Mars Colonization
Can a closed ecology sustain humans, waste processing included, for the full duration of a Mars mission?
A closed ecology run with humans for the mission duration including waste — regenerative operation, feedstock and nutrient recycling and human waste processing in one integrated closed system, which the literature prices at four to eight years of continuous operational experience.
“The single most informative experiment is one nobody is currently funding: a closed ecology run with humans for the mission duration, including waste.”
Orbital Rings
Does active support — a circulating mass loop holding up a stationary sheath and a static load — work at any real scale?
A terrestrial active-support demonstrator: a closed loop of circulating mass supporting a stationary sheath and a static load at a scale of tens of metres, testing the one mechanism the whole orbital-ring family shares.
“A terrestrial active-support demonstrator — a closed loop of circulating mass supporting a stationary sheath and a static load, at a scale of tens of metres — would test the one mechanism the whole family shares.”
Space Elevators
Does the tensile strength of record carbon-nanotube fibre fall with sample length, as the defect argument predicts, or hold flat as the doubling trend assumes?
A length-scaling series on record CNT fibre: measure tensile strength at 10 mm, 100 mm, 1 m, 10 m and 100 m, and publish the curve Pugno's argument says should fall.
“The most valuable unrun experiment is a length-scaling series, and it is cheap. Take the record fibre and measure its tensile strength at 10 mm, 100 mm, 1 m, 10 m and 100 m.”
Dyson Swarms
Are the Hephaistos infrared-excess candidates Dyson-swarm signatures or background galaxies?
The publish-then-resolve candidate cycle already run twice on the same objects: publish a pipeline and its output as candidates rather than detections, then let independent groups and better instruments close them — radio interferometry closed G in 2025, JWST/MIRI closed D and E in July 2026.
“The experiments in this subject are observations, and the most informative one has already been run twice on the same objects. The Hephaistos candidate list is the model: publish a pipeline, publish its output as candidates rather than detections, and then let independent groups and better instruments resolve them.”
O'Neill Cylinders
What chronic low-dose-rate, high-LET radiation limit is defensible for a settlement population including pregnancy and children — the number that sets habitat shielding mass?
Chronic low-dose-rate high-LET exposure studied in the populations a settlement actually contains, including pregnancy, children and testicular effects, because the assumed 20 mSv/yr general and 6.6 mGy/yr pregnancy limits are the terms of the whole mass calculation.
“The experiment that would decide the most is biological rather than structural: chronic low-dose-rate high-LET exposure in the populations a settlement actually contains.”
Space Habitats
What is ageing the ISS pressure boundary, and what does that say about designing a hull meant to last decades?
Find the root cause of the ISS leak: seven years, two agencies and a focus narrowed to internal and external welds with no identification yet, on the only long-duration pressure-boundary ageing dataset that exists.
“The most valuable experiment available is one already running badly: find the root cause of the ISS leak. Seven years, two agencies, a focus narrowed to internal and external welds, and no identification.”
Asteroid Mining
Can a commercial prospector reach a near-Earth asteroid and return the first commercial composition data taken there?
A commercial prospector that actually reaches an asteroid — DeepSpace-2, about 200 kg with electric propulsion, landing legs and platinum-group-metal instrumentation, manifested on IM-3 — measuring composition rather than mining.
“The decisive near-term experiment is a commercial prospector that actually reaches an asteroid. DeepSpace-2 — about 200 kg, 1.7 kW, electric propulsion, landing legs, platinum-group-metal instrumentation, manifested on IM-3 — is built.”
Planetary Defense
What are Dimorphos's bulk density and internal structure, and therefore how far does momentum enhancement really vary from the assumed value?
Hera's arrival at Didymos in November 2026 to determine both bodies' masses, measure the DART crater and characterise Dimorphos's internal structure and porosity — above all a bulk density, which moves beta between about 2.2 and 4.9.
“The next decisive experiment has a date. Hera arrives at Didymos in November 2026 and will determine the masses of both bodies, measure the DART crater, and characterise Dimorphos's internal structure and porosity.”
Space Weather Engineering
Does a spacecraft at L5 convert tens of minutes of solar-storm warning into days of useful notice?
Vigil at L5, contracted to Airbus in May 2024 for a 2031 launch, as the test of whether a different vantage on the Sun converts tens of minutes of warning into days.
“The experiment that would settle the lead-time question is a spacecraft, and it has a launch date. Vigil at L5, contracted to Airbus in May 2024 for a 2031 launch, is the test of whether a different vantage converts tens of minutes into days of useful notice.”
Stellar Engineering
What does a star actually do when an enclosing structure returns energy to it?
A stellar-structure calculation with existing codes modelling a star with an enclosing structure returning energy to it, where the physics — negative heat capacity, so heating causes expansion and cooling — is standard.
“The fourth is theoretical and is the one that would matter most: model a star with an enclosing structure returning energy to it. This is a stellar-structure calculation with existing codes, not a new instrument, and the physics — negative heat capacity, so heating causes expansion and cooling — is standard.”
Artificial Satellites for Climate Control
Can a gram-scale free-flyer hold a commanded orientation against solar radiation pressure for a year — the assumption both primary shade architectures rest on?
A gram-scale free-flyer with working attitude control: one flyer, 1.2 g, 0.6 m across, holding a commanded orientation against solar radiation pressure for a year.
“The decisive experiment nobody has proposed is a gram-scale free-flyer with working attitude control. One flyer, 1.2 g, 0.6 m across, holding a commanded orientation against solar radiation pressure for a year.”
Interstellar Civilization Models
Did life arise more than once, so that the abiogenesis rate gains a floor instead of a distribution spanning 200 orders of magnitude?
Find a second origin of life: a confirmed biosignature on an outer-system ocean world, or an independent abiogenesis event in Earth's own record.
“Experiment one, and it is the only one that touches the binding parameter: find a second origin. Not a second civilisation — a second instance of life arising.”
Exoplanet Colonization
Does TRAPPIST-1 e hold an atmosphere?
Consecutive transits of TRAPPIST-1 e paired with TRAPPIST-1 b as a bare-rock reference in the same system at the same epoch, using b as a contamination standard to subtract the star.
“Experiment one is already designed and is the highest-value observation in the subject: consecutive transits of TRAPPIST-1 e paired with TRAPPIST-1 b as a bare-rock reference. The point is to subtract the star.”
Deep Space Communications
Can store-and-forward networking carry traffic over multiple hops across interplanetary distance with real relay spacecraft?
A multi-hop store-and-forward network across interplanetary distance with real relays — the experiment the brief says matters most and has never been run.
“What has never been run is the experiment that matters most: a multi-hop store-and-forward network across interplanetary distance with real relays. That requires relay spacecraft, and the relays are proposals.”
Artificial Magnetospheres
Can a deployed superconducting loop produce a measurable magnetopause standoff against the real solar wind, and cut the particle flux inside it?
An orbital mini-magnetosphere: a deployed superconducting loop generating a measurable standoff against the real solar wind, instrumented with particle counts inside and outside, moving the field from simulation to measurement and testing the plasma-assistance argument that makes the power budget tractable.
“The decisive missing experiment is an orbital mini-magnetosphere. Nothing has flown. A deployed superconducting loop generating a measurable standoff against the real solar wind, with an instrumented particle count inside and outside, would move the whole field from simulation to measurement and would test the plasma-assistance argument that makes the power budget tractable.”
Terraforming
How long do 9-micrometre conductive rods stay suspended in a CO2 atmosphere at Martian pressure and temperature?
A chamber measurement of effective particle lifetime for the nanoparticle warming route: suspension behaviour, coagulation and settling for 9-micrometre conductive rods in a CO2 atmosphere at Martian pressure and temperature.
“The decisive next experiment for the nanoparticle route is a chamber measurement of effective particle lifetime. Suspension behaviour, coagulation and settling for 9-micrometre conductive rods in a CO 2 atmosphere at Martian pressure and temperature.”
Mega-Telescopes
Does a flown space coronagraph deliver the performance the next flagship's technical case depends on?
Roman's coronagraph in flight — launched 30 August 2026 with first images expected in early 2027 — as the flight test of the technology the next flagship's case depends on.
“Roman lifted off at 07:26 EDT on 30 August 2026 with a field of view at least 100 times Hubble's and a coronagraph aboard, on a three-month cruise with first images expected in early 2027.”
Black Hole Physics Applications
Is Hawking radiation an observed phenomenon rather than a theoretical expectation?
Detect Hawking radiation, astrophysically as a gamma-ray flash from an evaporating primordial black hole; searches were null as of January 2024.
“Experiment one, and it is the only one that bears on the assumption everything else rests on: detect Hawking radiation. The astrophysical route is a gamma-ray flash from an evaporating primordial black hole, and those searches were null as of January 2024.”
Interstellar Archaeology
Do Earth's co-orbital objects, the quasi-satellite 2016 HO3 above all, carry any artificial signature?
Survey Earth's co-orbitals by radar and SETI — objects Benford says have never been examined at all — with 2016 HO3 the strongest named target, already characterised as an astronomical body.
“Experiment one is the cheapest unexploited option in the subject: survey Earth's co-orbitals. Benford's checkable claim is that these objects have not been examined by SETI or by planetary radar at all, and the strongest named target — the quasi-satellite 2016 HO3 — is stable, close and already characterised as an astronomical body.”
Space Resource Economies
Can an in-situ resource plant run autonomously for more than five years, the lifetime the published break-even analysis turns on?
A lifetime test flown as a surface mission: an in-situ plant operating autonomously for more than five years, because the difference between four years and six years of plant life reverses the published break-even conclusion.
“The decisive economic experiment is a lifetime test, and it is a surface mission rather than a study. An in-situ plant operating autonomously for more than five years is what NASA's own break-even analysis turns on.”
Orbital Shipyards
Can two vendors' servicers dock with the same client fitting — does a standardised servicing interface exist in flight?
A standardised servicing interface flown by more than one operator: a demonstration in which two vendors' servicers dock with the same client fitting.
“A demonstration in which two vendors' servicers dock with the same client fitting would be a smaller flight than any of the above and would matter more than all of them.”
Moon-Based Manufacturing
Can a usable article be fabricated from real returned lunar material, turning process plausibility into performance data?
Fabricate one object from real lunar material — returned Apollo, Luna or Chang'e sample, beyond the microgram-scale laboratory work that is all that exists.
“The experiment that would change this subject most is also the smallest: fabricate one object from real lunar material. Nothing has ever been sintered, printed, cast or formed from returned Apollo, Luna or Chang’e material beyond microgram-scale laboratory work.”
Space Law and Governance
Is operator-defined, self-notified, temporary exclusivity around a lunar operation compatible with the Outer Space Treaty's non-appropriation rule?
The first declaration of a safety zone around a real lunar operation, which would put practice behind — or against — the claim that operator-defined temporary exclusivity is compatible with Article II.
“The experiment nobody has run is the safety zone. None has ever been declared around a real operation, so the claim that operator-defined, self-notified, temporary exclusivity is compatible with Article II has no practice behind it.”
Multi-Planetary Civilization
Can a human crew be closed — air, water, food and waste — for a conjunction-class Mars duration with no resupply?
Close a human crew for a conjunction-class Mars duration with no resupply: air, water, food and waste, at a crew size in double figures, for the length of a real mission.
“Experiment one, and everything else is secondary to it: close a human crew for a conjunction-class Mars duration with no resupply. Air, water, food and waste, at a crew size in double figures, for the length of a real mission.”
Orbital Debris and Space Traffic
Whether a large, uncooperative piece of orbital debris can be captured and deorbited reliably and cheaply enough for remediation to become an operational service rather than a demonstration.
JAXA's CRD2 Phase II mission: Astroscale's ADRAS-J2 grapples the H-2A upper stage inspected by ADRAS-J with a robotic arm and drags it into a destructive reentry.
“The decisive demonstration is JAXA’s CRD2 Phase II: Astroscale’s ADRAS-J2 is contracted to grapple the same H-2A upper stage with a robotic arm and drag it into a destructive reentry — a capture of a large, uncooperative, slowly tumbling object that nobody has ever performed.”
In-Space Servicing and Depots
Whether cryogenic propellant can be transferred between two spacecraft in orbit at architecture-relevant scale.
The Starship-to-Starship cryogenic propellant transfer demonstration under NASA's Human Landing System programme: two vehicles docking in low Earth orbit and moving liquid oxygen and methane between them at architecture-relevant scale.
“The decisive test is spacecraft-to-spacecraft cryogenic propellant transfer, and it is the single result that would most change the assessment this brief makes.”
Cislunar Navigation and Lunar Time
Whether a purpose-built lunar navigation broadcast can deliver position and time to an independent user on the Moon at its specified accuracy.
An Augmented Forward Signal broadcast from a lunar orbiter, used by an independent receiver to fix position and time on the surface and checked against laser-ranged ground truth.
“The decisive demonstration is an Augmented Forward Signal broadcast from lunar orbit that an independent receiver uses to fix position and time on the surface, checked against laser-ranged ground truth; until a beacon that no simulation controls closes that loop, every navigation architecture in this brief is a paper system.”
Multi-Messenger Astronomy
Whether the alert-to-counterpart chain works under routine conditions, or whether GW170817 was a one-off produced by an emergency mobilisation that does not scale.
A second binary neutron star merger localised tightly enough that a kilonova is found and a host redshift is measured, converting the counterpart rate from a quantity estimated off one detection into a measured number and supplying a second independent standard siren.
“The single result that would most change this brief is a second binary neutron star merger localised tightly enough that a kilonova is found and a host redshift measured.”
Neutrino Astronomy
Whether the neutrino mass ordering is normal or inverted, which is the hinge the cosmological mass tension, the reach of tonne-scale double beta decay, and the extraction of the CP phase all hang on.
A 20-kilotonne liquid scintillator detector 53 kilometres from two reactor complexes reads the mass ordering out of the reactor antineutrino energy spectrum from vacuum oscillation alone, independent of matter effects and of the CP phase, using energy resolution near 3% at one megaelectronvolt.
“The decisive experiment in this subject is already running, and it is not an astronomical one: a 20-kilotonne liquid scintillator detector 53 kilometres from two reactor complexes determines the neutrino mass ordering from vacuum oscillation alone, independent of matter effects and independent of the CP phase.”
Megaconstellation Externalities
Whether the optical cost of satellite constellations to wide-field survey astronomy is the modelled tens of per cent of contaminated images, or small enough that scheduling and masking absorb it.
The Vera C. Rubin Observatory's ten-year survey returns a measured satellite-trail contamination rate, and a residual-artefact rate after masking, against the contamination predictions published before the survey began.
“The decisive test is already running: the Rubin Observatory’s survey statistics against its own pre-survey predictions. Predictions for the fraction of images contaminated by trails, and for residual damage after masking, were published before the ten-year survey began, at stated assumptions about fleet size and brightness.”
Planetary Protection
Whether the spore-count assay that defines planetary-protection compliance bears a known relation to the total viable microbial burden actually carried on flight hardware.
Measure the bioburden of one spacecraft in final integration twice on the same surfaces at the same time - once by the standard heat-shock culture assay that defines legal compliance, once by a validated culture-independent method counting total viable organisms - and publish the conversion factor with error bars.
“The decisive experiment is a paired assay of the same flight hardware. Take a spacecraft in final integration and measure its bioburden twice: once by the standard culture assay that defines legal compliance, once by a validated culture-independent method that counts total viable organisms, on the same surfaces, at the same time, with published protocols and error bars.”
Biosignature Standards
Whether the best temperate rocky exoplanet accessible to current instruments has a secondary atmosphere at all, which sets whether remote biosignature work has a near-term target list.
The large dedicated JWST observing programme on TRAPPIST-1 e either detects a secondary atmosphere - the first for a temperate rocky exoplanet - or returns a null at full depth, establishing that the nearest and most favourable M-dwarf rocky planets are airless.
“The decisive observation is whether TRAPPIST-1 e has a secondary atmosphere. It is the best-placed temperate rocky planet accessible to the most capable telescope in existence, it has been the target of a large dedicated observing programme, and the two inner planets in the same system have already returned bare-rock-consistent and thick-atmosphere-excluded results.”
III — Biology & Human Enhancement
Biological Immortality
Is the flat mortality hazard measured in hydra a property of the animal or of its laboratory husbandry?
Re-run the hydra demography across a food and temperature gradient, extending the planarian resource-limitation design to the flagship flat-hazard organism.
“Re-run the hydra demography across a food and temperature gradient. This is a direct extension of the planarian resource-limitation design to the flagship organism, it costs a postdoc and glassware, and it would settle whether the strongest flat-hazard result in biology is a property of hydra or a property of husbandry.”
Human Hibernation
Does ultrasonic induction of torpor replicate outside its originating laboratory, and does it work through a human-scale skull?
Independent replication of ultrasonic torpor induction, first in rats at the published parameters and then transcranially in a pig, whose skull acoustics approximate human.
“A negative in the pig would be the most informative single result available today.”
Whole Organ Regeneration
Can any intervention reduce scar formation in a large mammal, and is the closing of the regenerative window a mouse fact?
The scar assay: controlled infarcts in neonatal and juvenile pigs, measuring the percentage of infarct area occupied by collagen at 90 days with and without a candidate intervention, alongside ejection fraction.
“The cheapest decisive experiment is the scar assay, and it is not being run. Take neonatal and juvenile pigs, produce a controlled infarct, and measure the percentage of infarct area occupied by collagen at 90 days with and without a candidate intervention, alongside ejection fraction.”
Limb Regeneration
Can a mammalian induced mass be made to produce the skeletal element correct for its amputation level rather than an extra one?
Score induced skeletal elements by identity rather than presence — morphology plus Shox, Meis and Hoxa13 expression — and report the fraction that are positionally correct rather than ectopic. The brief calls this experiment two the binding one.
“Experiment two is the binding one, and experiments one and five are the cheapest. Nothing beyond link two is worth funding until a mammalian induced mass can be shown to produce the correct element for its amputation level rather than an additional one.”
Artificial Wombs
Does the 336-hour extracorporeal gestation result hold outside the single group that produced it?
An independent laboratory repeating the 336-hour lamb result at the same gestational age, weight band and duration, reporting completion rate, cause of termination and organ weights against in-utero controls.
“The 336-hour result rests on one group and one cohort of six; an independent laboratory running the same gestational age, weight band and duration, reporting completion rate, cause of termination and organ weights against in-utero controls, would either consolidate the best result in the field or reveal it as one lucky animal.”
Synthetic Biology
Is the unpredictability that synthetic biologists report a property of the biology or of the measurement?
A properly powered interlaboratory ring trial on a standard construct — one plasmid, one host, one readout, ten laboratories, pre-registered — run to separate biological unpredictability from measurement variance.
“The cheapest high-value experiment in this subject is an interlaboratory ring trial on a standard construct, and it has never been run.”
Xenobiology
Does genetic information actually cross from XNA into DNA in real microbial communities — is xenobiology's firewall real?
Incubate defined XNA templates with soil, sediment and gut microbial communities and measure, by sequencing, XNA-to-DNA information transfer events per template per unit time against a containment threshold.
“1. Measure the firewall against a real metagenome. Incubate defined XNA templates with soil, sediment and gut microbial communities and measure the rate of XNA-to-DNA information transfer by sequencing.”
Directed Panspermia
Do the elements represented in terrestrial biology correlate with the composition of any stellar class, as directed panspermia's founding paper predicted?
Crick and Orgel's 1973 test, generalised from the molybdenum question: correlate biological element-usage tables against modern stellar abundance catalogues. A null closes the only discriminating observation the founding paper offered; a positive would be the first evidence the hypothesis has ever had.
“The cheapest high-value experiment in this subject was specified in 1973 and has never been run. Crick and Orgel proposed testing whether the elements represented in terrestrial biology correlate with the composition of some stellar class — the molybdenum question, generalised.”
Biological Enhancement
Does polygenic embryo selection deliver the cognitive and educational gain it is sold on?
A prospective cohort of polygenically selected children with the predictor, selection rule and outcome pre-registered, followed to a measured cognitive and educational endpoint against a sibling or non-selected comparison.
“The single most valuable experiment in this subject is a prospective cohort of polygenically selected children, and nobody is running it.”
Genetic Engineering
Can any delivery system edit more than 30% of cells in muscle, central nervous system or haematopoietic stem cells without surgery?
Demonstrate an LNP or capsid achieving above 30% in vivo editing in muscle, central nervous system or haematopoietic stem cells without stereotactic placement or surgical injection, at a dose below the established hepatotoxicity threshold.
“The decisive somatic experiment is a delivery measurement, and it has a number attached. Demonstrate an LNP or capsid achieving above 30% editing in vivo in muscle, central nervous system or haematopoietic stem cells, without stereotactic placement or surgical injection, at a dose below the hepatotoxicity threshold the nex-z programme has now established.”
Designer Organisms
Do gene drives behave in a wild population the way they behave in a cage?
A monitored open field release of a driving construct with pre-registered predictions, supplying the wild genetic diversity, migration, seasonality and population structure that cage work cannot.
“The most valuable experiment in this subject is the one that was about to happen and did not: a monitored field release with pre-registered predictions. Every gene-drive result in the literature is a cage result.”
Brain Preservation
Is a learned behaviour recoverable from a preserved and imaged brain — is the thing preservation preserves the thing that matters?
Train a C. elegans or a Drosophila, preserve it by aldehyde-stabilised cryopreservation, image it, simulate it and test whether the trained state is distinguishable from a naive control. The brief names this link 2 as the binding one for whether preservation is worth doing at all.
“Both connectomes already exist; the experiment costs a fraction of one prize purse, and its endpoints are behavioural rather than interpretive, so it can return a clean negative.”
Cryonics
Does a learned association survive preservation, imaging and simulation — is the structural hypothesis behind cryonics a finding or an assumption?
Train an invertebrate whose complete connectome already exists, preserve it, image it, simulate it and test whether the learned association survives.
“The sixth belongs to the structural claim and is the cheapest decisive experiment in the entire preservation cluster: train an invertebrate whose complete connectome already exists, preserve it, image it, simulate it, and test whether the learned association survives.”
Bioelectric Medicine
Does the permanent double-headed planarian phenotype replicate in a laboratory with no stake in the bioelectric framework?
A pre-registered, blinded, independent replication of the permanent-double-head planarian phenotype under the published protocol, scored as the proportion of animals showing the altered phenotype across successive amputation rounds.
“One: the cheap one. Pre-registered, blinded replication of the permanent-double-head planarian phenotype in an independent laboratory, under the published protocol. Measurement: proportion of animals showing the altered phenotype across successive amputation rounds, with the analysis plan fixed in advance.”
Human Adaptation for Space
Does Dsup protect human cells against the heavy-ion radiation that actually matters in deep space?
Comet and gamma-H2AX assays in Dsup-expressing human cells across iron-56 and carbon tracks at a heavy-ion beamline, against the published X-ray baseline.
“1. Dsup at a heavy-ion beamline. Comet and gamma-H2AX assays in Dsup-expressing human cells across iron-56 and carbon tracks at NSRL, GSI or HIMAC, against the published X-ray baseline.”
Cellular Rejuvenation
Does a change in an epigenetic clock predict a change in function?
Report a functional endpoint — grip strength, wound healing, immune response, cognition, survival — alongside every clock delta, and test across studies whether one predicts the other. The brief names this the binding link, and conceptual rather than technical.
“Which link is actually binding: the first, and it is conceptual rather than technical. Until a clock delta is shown to predict a functional delta, links 3 through 8 are all being scored on an instrument of unknown validity, and a positive result at any of them can be read two ways.”
Longevity Therapies
Do epigenetic-clock deltas predict subsequent morbidity and mortality within individuals rather than across them?
Use existing biobanks with repeat sampling to test prospectively whether within-individual epigenetic-clock deltas predict later morbidity and mortality, validating the endpoint every other longevity trial is scored on.
“The workback plan has nine links and the first one is cheap, decisive and unfunded. (1) Validate the endpoint. Take existing biobanks with repeat sampling and test prospectively whether epigenetic-clock deltas predict subsequent morbidity and mortality within individuals rather than across them.”
Neurogenetics
Can a credible causal psychiatric variant be converted into a measured neuronal mechanism?
Base- or prime-edit a credible causal schizophrenia variant into isogenic human iPSC-derived neurons and measure a synaptic or electrophysiological phenotype.
“Sixteen of the 120 prioritised genes have a credible causal variant identified; this is the experiment that converts one of them into a mechanism, and it is the rate-limiting step for every drug-discovery claim made for psychiatric GWAS.”
Microbiome Engineering
Has microbiome engineering produced one replicated, adequately powered positive trial outside single-pathogen displacement?
One randomised, adequately-powered, replicated positive trial in an indication that is not single-pathogen displacement — the brief's honest scoreboard for the whole field.
“8. Win once outside displacement. One randomised, adequately-powered, replicated positive trial in an indication that is not single-pathogen displacement. None exists. This is the honest scoreboard for the whole field.”
Universal Vaccines
Does escape from stalk-directed immunity cost influenza enough fitness for a universal vaccine to be durable?
A fitness-cliff study: serially passage influenza under stalk-directed monoclonal or polyclonal pressure, isolate escape mutants, and measure their replicative fitness against wild type in primary human airway cultures and in ferrets.
“Serially passage influenza in the presence of stalk-directed monoclonal or polyclonal pressure, isolate escape mutants, and measure their replicative fitness against wild type in primary human airway cultures and in ferrets.”
Biological Computing
Does the reported wetware learning result hold up under independent, blinded, pre-registered replication?
An independent, blinded, pre-registered replication of the original cultured-neuron learning result — the binding link in the brief's wetware chain.
“Two: an independent, blinded, pre-registered replication of the original result.”
Lab-Grown Organs
How far is any engineered organ-scale vascular tree from clinical viability — a factor of two or a factor of a thousand?
A time-to-thrombosis benchmark: perfuse a recellularised or printed organ-scale vascular tree with whole blood at physiological pressure and flow, and publish hours to occlusion against endothelial coverage fraction.
“First, and most important: a time-to-thrombosis benchmark. Perfuse a recellularised or printed organ-scale vascular tree with whole blood at physiological pressure and flow, and publish hours to occlusion against endothelial coverage fraction.”
Human-Machine Symbiosis
Does channel count actually determine brain-computer interface task performance?
Publish the channel-count scaling curve: take one fixed task on one participant with a high-channel implant, subsample channels from the full array down two orders of magnitude, and report task accuracy as a function of channels retained.
“Experiment one, and the cheapest: publish the channel-count scaling curve. Take one fixed task on one participant with a high-channel implant, subsample channels from the full array down two orders of magnitude, and report task accuracy as a function of channels retained. This requires no new surgery, no new hardware and no new consent beyond re-analysis.”
Neuroplasticity Engineering
Does valproate reopen a critical period for adult perceptual learning, and does any gain survive withdrawal?
Replicate valproate properly: parallel-group rather than crossover, a genuine pre-training baseline, a pre-registered primary endpoint on the pitch task, and re-testing at three, six and twelve months after withdrawal.
“One: replicate valproate properly. Parallel-group rather than crossover, a genuine pre-training baseline, a pre-registered primary endpoint on the pitch task, and re-testing at three, six and twelve months after withdrawal. Roughly 120 participants, a generic anticonvulsant and a laptop task. It is the cheapest decisive experiment in this brief and its absence for thirteen years is the field's most conspicuous unforced error.”
Precision Medicine
Can any polygenic predictor plus routine clinical data reach clinical-grade discrimination for a common disease?
An out-of-sample, cross-ancestry validation of a polygenic predictor reaching an area under the curve above 0.85 for a common disease from genotype plus routine clinical variables.
“First: an out-of-sample, cross-ancestry validation of any polygenic predictor reaching an area under the curve above 0.85 for a common disease from genotype plus routine clinical variables. Nothing in this brief's sources approaches it, and the distance from 0.64 to 0.85 is the distance between a research instrument and a clinical one.”
Nanomedicine
Is nanoparticle delivery into tumours governed by endothelial transcytosis or by particle size?
Measure, across a panel of tumour types, whether delivered dose correlates with endothelial transcytosis markers or with nanoparticle hydrodynamic diameter — the two predictions come apart on a single dataset.
“The one experiment that would change everything about the delivery branch is cheaper still. Measure, across a panel of tumour types, whether delivered dose correlates with endothelial transcytosis markers or with nanoparticle hydrodynamic diameter.”
Synthetic Ecosystems
How does the persistence of a closed ecosystem scale with its size, and which functional guilds are lost when?
A microcosm array: 100 or more replicate closed microcosms at logarithmically spaced volumes spanning three or four orders of magnitude, seeded identically and run for years, yielding a size-versus-persistence curve and a per-guild extinction hazard rate.
“The one experiment that would change everything is a microcosm array, and it costs less than a single crewed demonstration month. Build 100 or more replicate closed microcosms at logarithmically spaced volumes spanning three or four orders of magnitude, seed them identically, run them for years, and measure which functional guilds are lost, when, and as a function of what.”
Biological Energy Systems
Do the reported performance gains in microbial electrochemical devices survive independent, protocol-controlled replication?
A round-robin: three independent laboratories, a registered protocol, synthetic wastewater of stated composition, fixed external resistance, thirty degrees Celsius and thirty days, reporting anode-, cathode-, membrane-areal and volumetric power density plus coulombic efficiency for the same device.
“One: the round-robin. Three independent laboratories, a registered protocol, synthetic wastewater of stated composition, a fixed external resistance, thirty degrees Celsius, thirty days, reporting anode-areal, cathode-areal, membrane-areal and volumetric power density plus coulombic efficiency for the same device. Cost is well under a million dollars.”
Future Agriculture
How much nitrogen does the Sierra Mixe maize association actually fix in the field?
An isotopic nitrogen-balance measurement on Sierra Mixe maize across multiple field sites and seasons, reported as percentage of nitrogen derived from atmosphere with error bars.
“First: the isotopic nitrogen-balance measurement on Sierra Mixe maize across multiple field sites and seasons, reported as percentage of nitrogen derived from atmosphere with error bars. It is the ceiling for the associative route and the field has been arguing about a shortcut whose size it has not published widely enough for a reader to find.”
De-Extinction Technologies
Can a mammal be gestated to term entirely outside a uterus, and can a second laboratory repeat it?
Complete full-term ex-utero gestation in a marsupial and have it replicated by a second, unaffiliated laboratory.
“Complete full-term ex-utero gestation in a marsupial and have it replicated by a second, unaffiliated laboratory — the single highest-value experiment in the subject, because marsupials have the shortest gestation and pouch-based development, and because success there converts a scientific unknown into an engineering programme.”
Post-Antibiotic Medicine
Does the complete diagnose-and-narrow pathway, rather than any of its components alone, reduce deaths and antibiotic exposure from resistant infection?
A cluster-randomised pragmatic trial in hospitals of the full bundle -- sub-hour susceptibility testing, protocolised narrowing, enforced stewardship review and narrow-spectrum oral step-down -- against usual care, with 28-day all-cause mortality and days of therapy per thousand patient-days as co-primary endpoints.
“The decisive experiment is a cluster-randomised pragmatic trial of the whole diagnose-and-narrow pathway, with 28-day all-cause mortality and days of therapy per thousand patient-days as co-primary endpoints.”
Immune Engineering
Is drug-free remission after a single CD19 CAR-T infusion a durable cure or a spectacular multi-year remission in B-cell-driven autoimmune disease?
A randomised, controlled trial of CD19 CAR-T in refractory B-cell autoimmunity powered for relapse over three to five years, rather than for response at one year.
“The decisive test is whether drug-free remission after a single CD19 CAR-T infusion holds at three to five years in a controlled cohort large enough to measure a relapse rate.”
RNA Medicines
Can an RNA drug be delivered durably and safely to a therapeutic target outside the liver and central nervous system at a dose that controls the target protein?
A single demonstration of durable, well-tolerated systemic RNA delivery to an extrahepatic, non-CNS tissue such as muscle, controlling the target protein for months from one administration.
“The decisive test is a single demonstration of durable, well-tolerated RNA delivery to a therapeutic target outside the liver and central nervous system — muscle is the nearest — at a dose that restores or suppresses the target protein for months from one administration.”
Microphysiological Systems
Do organ chips predict human toxicity prospectively, and does that predictivity survive transfer to laboratories that did not build them?
A pre-registered, blinded, multi-laboratory trial in which at least three independent labs run the same chip protocol on twenty to forty compounds entering first-in-human studies and deposit public, time-stamped toxicity calls before any clinical data exist.
“Take twenty to forty compounds entering first-in-human studies, distribute them under code to at least three independent laboratories running the same liver-chip or multi-organ protocol, require each laboratory to deposit a public, time-stamped call”
Ex-Vivo Organ Repair
Does perfusion actually repair a donor organ, or does it only identify the organs that were already good enough to transplant?
A paired-kidney trial in which one kidney from each deceased donor goes to perfusion plus a candidate repair intervention and the other to perfusion alone, both transplanted, scored on delayed graft function and one-year function.
“That design removes donor variation, the largest confounder in the field, and it isolates repair from selection”
Pathogen-Agnostic Surveillance
Does agnostic sensing - metagenomic sequencing of wastewater and of undiagnosed severe illness - add true detections over existing targeted surveillance, at what false-alarm rate and at what cost per detection?
A three-year pre-registered parallel run in one defined catchment of a few million people: agnostic metagenomic sequencing of wastewater and of residual severe-illness clinical specimens at declared depth, operated alongside the existing clinical and laboratory surveillance system, with detection criteria, alarm thresholds and adjudication rules fixed in advance, reporting every detection, every false alarm and the cost of each.
“The decisive experiment is a pre-registered parallel run, and nobody has funded it. In one defined catchment of a few million people, operate agnostic metagenomic sequencing of wastewater and of residual severe-illness clinical specimens at a declared depth for three years, alongside the existing clinical and laboratory surveillance system”
Climate-Health Adaptation
Does a defined heat intervention - alerting tied to active outreach to at-risk residents plus a funded cooling option - reduce all-cause mortality, and by how much?
A stepped-wedge or cluster-randomised trial of a defined heat intervention across municipalities, exploiting the staggered administrative rollout of heat plans, with all-cause mortality as the pre-registered endpoint.
“The decisive experiment is a stepped-wedge or cluster-randomised trial of a defined heat intervention, with all-cause mortality as the pre-registered endpoint, and nobody has funded one. Municipalities are natural clusters; heat plans are already rolled out at different times for administrative reasons”
Exposomics
Does untargeted measurement of archived human specimens discover exposure-disease relationships that targeted epidemiology did not already know?
A blinded, pre-registered, multi-cohort replication round for untargeted exposure signals in pre-diagnostic archived specimens, scored against a significance threshold agreed before anyone looks at the data: one discovery cohort, three independent replication cohorts, one blinded laboratory, a pre-specified multiplicity correction, and publication of the null.
“Nobody has funded it. The specimens exist, the instruments exist, and the missing item is an agreement between cohorts to be scored against each other.”
Generative Biomolecular Design
Are generative design methods actually improving at producing molecules that work, and do their in-silico filters have prospective value outside the laboratories that publish them?
A standing, pre-registered, third-party prospective benchmark: a fixed panel of targets spanning easy and hard classes, a fixed number of designs per method, one independent laboratory, one assay protocol, and publication of every result including the nulls.
“Pilot versions already exist as open design competitions run by commercial testing laboratories, which is what makes the standing version a funding decision rather than a research problem.”
Environmental Biotechnology
Is the destruction of perfluoroalkyl acids a biological problem at all, or does it belong permanently to thermal and chemical processes?
A demonstrated biological cleavage of a carbon-fluorine bond in a fully fluorinated perfluoroalkyl acid, with a measured rate and an identified catalyst.
“The decisive result is a demonstrated biological cleavage of a carbon-fluorine bond in a fully fluorinated perfluoroalkyl acid, with a measured rate and an identified catalyst.”
Metabolic Medicines
Does long-term incretin therapy in older adults cost them physical function, falls and fractures, and does concurrent resistance training prevent it?
A randomised trial in adults over 65 with obesity, powered for physical function, falls and fractures rather than weight, comparing incretin therapy alone against incretin therapy plus supervised resistance training and protein targets over at least three years.
“The decisive experiment this brief is waiting on is a randomised trial in adults over 65 with obesity, powered for physical function, falls and fractures rather than weight, comparing incretin therapy alone against incretin therapy plus supervised resistance training and protein targets, over at least three years.”
Mitochondrial Medicine
Does the small carryover of maternal mitochondrial DNA in children born after mitochondrial donation stay below disease threshold, or drift upward over a lifetime as it did in cultured cells?
Decades-long follow-up of the children born after licensed mitochondrial donation, with serial heteroplasmy measurement in accessible tissues alongside growth, metabolic and neurodevelopmental outcomes.
“The decisive result is the long-term follow-up of the children born after mitochondrial donation: serial heteroplasmy measurement in accessible tissues, with growth, metabolic and neurodevelopmental outcomes, over decades.”
Embryo Models
Do integrated stem-cell-based embryo models have developmental potential, or do they stop at a hormonal pregnancy signal?
A properly powered primate transfer experiment: integrated monkey embryo models transferred into enough recipient females, with pre-specified endpoints for hormonal signalling, sac formation and fetal development, and the negatives reported.
“The decisive experiment is the primate transfer test, extended and properly powered: transfer integrated monkey embryo models into a sufficient number of recipient females, with pre-specified endpoints for hormonal signalling, sac formation and fetal development, and report the negatives.”
In-Vitro Gametogenesis
Can a human oocyte be made entirely from pluripotent stem cells in culture, and at what efficiency and epigenetic fidelity?
A human oocyte derived wholly from pluripotent cells in culture, carried to metaphase II and shown to be fertilisable, reported with its efficiency denominator and its methylation state at imprinted loci.
“The decisive experiment is a human oocyte, derived entirely from pluripotent cells in culture, that reaches metaphase II and is shown to be fertilisable, reported with its efficiency denominator and its methylation state at imprinted loci.”
Precision Neuropsychiatry
Can a pre-treatment measurement assign a psychiatric patient to the treatment that will work better for them, when tested with the interaction as the registered primary endpoint?
A prospective randomised trial powered for a biomarker-by-treatment interaction, assigning patients to one of two treatments by a pre-specified marker, with that interaction as the registered primary endpoint.
“The decisive experiment is a prospective, adequately powered randomised trial in which patients are assigned to one of two treatments by a pre-specified biomarker, with the biomarker-by-treatment interaction as the registered primary endpoint.”
Neurodegeneration Reversal
Does clearing amyloid from biomarker-positive people before symptoms appear substantially delay clinical onset, or has the amyloid hypothesis now been tested at the stage most favourable to it?
The secondary-prevention trials: removing amyloid from biomarker-positive people with no symptoms, with clinical onset as the endpoint, in two large industry programmes and one familial-mutation study.
“The decisive experiment is the secondary-prevention trial: removing amyloid from biomarker-positive people who have no symptoms, with clinical onset as the endpoint.”
Biomanufacturing Resilience
Does reserved, ever-warm manufacturing capacity actually deliver released doses on an emergency clock, or is it an untested budget line?
A no-notice activation of reserved capacity in which a regulator names an antigen the contractor has not seen, the reservation is called, and the elapsed time is measured to released, lot-tested doses rather than to an announcement.
“Nobody has scheduled or funded such a drill, which is why the most expensive answer in the field is the least tested.”
Sensory Restoration
Does restored biological hearing from otoferlin gene therapy persist and support spoken-language acquisition better than a cochlear implant does?
Five-year durability and spoken-language outcomes in the first otoferlin gene therapy cohorts, compared against age-matched children who received cochlear implants.
“The decisive result is five-year durability and spoken-language outcome in the first otoferlin cohorts, measured against age-matched implanted children.”
IV — Consciousness & Intelligence
Artificial General Intelligence
Whether machine ability estimates predict performance on instruments that did not exist when the estimates were fitted.
A pre-registered out-of-distribution prediction: fit latent ability for N models on instrument A, publish point predictions for instrument B before B exists, have B built by a disjoint team from a disjoint task generator, and report calibration.
“The decisive experiment is a pre-registered out-of-distribution prediction, and it costs almost nothing except institutional willingness. Fit the latent ability of N models on instrument A; publish point predictions for their scores on instrument B before instrument B exists; have instrument B built by a disjoint team from a disjoint task generator; report calibration.”
Machine Consciousness
Whether consciousness can be assessed in a machine without relying on that machine's own reports.
L1: a report-independent measure of consciousness validated in humans, on which the whole seven-link chain hangs.
“L1 — a report-independent measure validated in humans. The whole chain hangs on it and it is a scientific unknown of the deepest kind; it belongs to Consciousness Research and may never arrive.”
Digital Minds
Whether an AI system's self-reports track its own internal states at all, which is the prior question to any claim about machine experience.
Scale the concept-injection programme — more models, more concepts, pre-registered detection thresholds, and adversarial controls establishing what a system with no introspective access would score — so that the 20% figure becomes a measurement of whether self-report tracks internal state.
“The highest-value experiment is architectural rather than behavioural, and it is already partly specified. Concept injection asks whether a model can detect a manipulation of its own activations.”
Mind Uploading
Whether a preserved brain retains enough to reconstruct an individual's learned behaviour rather than its species' generic behaviour.
Preserve, read out, then behave: preserve a small animal with a known behavioural repertoire by the best available method, reconstruct the connectome and whatever molecular state survives, build the emulation, and test whether it reproduces that individual's learned behaviour.
“The decisive experiment is small, cheap by the standards of this field, and has never been run: preserve, read out, then behave.”
Collective Intelligence
Which aggregation rules tolerate how much correlated error among the individuals being aggregated.
Build the correlated-error curves: simulate the aggregation rules over synthetic populations whose pairwise error correlation is swept from zero to one at fixed individual accuracy, then calibrate on real crowd datasets by measuring the correlation directly.
“Experiment one, and the one that binds: build the correlated-error curves. Simulate majority vote, mean, median, trimmed mean, confidence-weighted mean, surprisingly-popular and market-scoring aggregation over synthetic populations whose pairwise error correlation is swept from zero to one, holding individual accuracy fixed.”
Intelligence Amplification
Whether human-AI complementarity is a stable architecture or a waypoint that closes as automated baselines improve.
Link 5: re-run the routing experiment with the automated baseline advanced one model generation and the routing policy retrained, everything else held, and report whether the complementarity margin holds or shrinks.
“Link 5 binds, and it is the experiment that decides this brief’s central question: re-run Link 4 with the automated baseline advanced one model generation, the routing policy retrained, everything else held. If the complementarity margin holds, augmentation is a stable architecture; if it shrinks, it is a waypoint.”
Memory Engineering
Whether hippocampal stimulation addresses a specific memory or merely modulates a general encoding state.
Derive a stimulation pattern from the model of item A, deliver it while the subject encodes item B, and test both, with a double dissociation — recall of A improves, recall of B does not — as the pass condition.
“The decisive experiment for the prosthesis has never been run and it is not expensive. Derive a stimulation pattern from the model of item A, deliver it while the subject encodes item B, and test both. Pass condition: a double dissociation in which recall of A improves and recall of B does not.”
Brain-Computer Interfaces
Whether an implanted brain-computer interface outperforms existing assistive communication devices on the same task.
Score the best current BCI in its best participant against eye-tracking, a switch scanner and a touchscreen on the same communication task, same scorer, same day, reporting words per minute, error rate, setup time and fatigue.
“One: the head-to-head nobody runs. Take the best current BCI in its best participant and score it against eye-tracking, a switch scanner and a touchscreen on the same communication task, same scorer, same day, reporting words per minute, error rate, setup time and fatigue.”
Human-AI Integration
Whether published human-AI complementarity results survive comparison with a model that abstains instead of always answering.
The abstention control: take any published complementarity result and re-run it against a well-calibrated model that abstains at matched coverage instead of one that always answers.
“Take any published complementarity result and re-run it against a well-calibrated model that abstains at matched coverage instead of one that always answers.”
Consciousness Research
Whether the no-report paradigms a decade of consciousness science rests on measure the same thing as report-based ones.
Run binocular rivalry within-subject with simultaneous report and no-report readouts — optokinetic nystagmus and pupillometry alongside a decoding model — and test whether the decoded percept time-series are statistically identical rather than merely correlated.
“The report-versus-no-report tie-break, which is L1 and is overdue. Run binocular rivalry within-subject with simultaneous report and no-report readouts — optokinetic nystagmus and pupillometry alongside a decoding model — and test whether the decoded percept time-series are statistically identical, not merely correlated.”
Integrated Information Theory
Whether experience tracks causal structure, as integrated information theory requires, or only what is globally broadcast.
The silent-connection experiment: alter a circuit's connectivity without altering ongoing firing — optogenetic silencing of a currently inactive pathway, or pharmacological block of a synapse carrying no traffic — and measure reported experience, where GNWT predicts no change and IIT predicts one.
“The silent-connection experiment is the sharpest available discriminator and has not been run. Take a circuit whose connectivity can be altered without altering ongoing firing — optogenetic silencing of a pathway that is currently inactive, or pharmacological block of a synapse carrying no traffic — and measure reported experience.”
Orch OR
Whether the programme's flagship superradiance result transfers from the bench to the cell.
Ferritin titration against tryptophan superradiance: reproduce the 2024 superradiance measurement, then repeat it across a physiological range of ferritin concentrations.
“(1) Ferritin titration against tryptophan superradiance. Reproduce the 2024 superradiance measurement, then repeat it across a physiological range of ferritin concentrations. Narrow, cheap, and it decides whether the programme's flagship recent result transfers to the cell.”
Distributed Cognition
Whether externally stored information carries the dispositional signatures of internally stored memory.
A longitudinal study of heavy external-memory users — dense digital note-takers and amnesic patients using prosthetic memory systems — probing whether externally stored items show interference, the generation effect, retrieval-induced forgetting and recollection-versus-familiarity recall dynamics.
“The experiment that would make this dispute empirical is available now and appears not to have been run. Rupert’s claim is that notebook-belief and memory-belief differ in dispositional profile. That is testable.”
Cognitive Enhancement
Whether cognitive enhancers raise performance or only raise confidence in it.
The confidence ratio: in any enhancement study, administer the intervention, measure objective performance and self-rated performance on the same task, and publish the ratio.
“One: the confidence ratio. In any enhancement study, administer the intervention, measure objective performance and self-rated performance on the same task, and publish the ratio.”
Synthetic Consciousness
Whether integrated information can be measured practically with proven bounds, which would supply the field's only design objective or kill it.
A practical measure of integrated information, rigorously tested where the ground truth can be established, with proven bounds rather than correlations.
“The highest-value experiment is the one the approximation literature explicitly asked for and nobody has run. A practical measure of integrated information, rigorously tested in an environment where the ground truth can be established, with proven bounds rather than correlations — the difference between an approximation and a proxy.”
Cognitive Architectures
Whether the match-cost constraint that shaped forty years of symbolic cognitive architecture survives modern indexing.
Measure Soar 9.6.5's match cost per decision cycle as a function of rule count, then reimplement the match over a modern index — the structures a database query planner or vector store would use — and publish the curve again.
“X1 — measure the match-cost curve, then swap the index. Take Soar 9.6.5, generate rule sets spanning several orders of magnitude in size, and publish match cost per decision cycle as a function of rule count, with working-memory size held fixed and then varied.”
Neural Interfaces
Whether electrode encapsulation scales sublinearly with implanted surface area, which sets the ceiling on channel density.
The density study: implant the same electrode material at four densities in one species and measure yield at two years as a function of implanted surface area per cubic millimetre, with sublinear scaling of encapsulation as the pass condition.
“The density study, which is the scientific binding link. Implant the same electrode material at four densities in one species and measure yield at two years as a function of implanted surface area per cubic millimetre. Pass condition: sublinear scaling of encapsulation with surface area.”
Artificial Creativity
Whether transformational creativity can be measured at all, before any such measure is applied to machine output.
L3, the binding link: build a gold-standard set of historical H-creative artifacts labelled by domain experts as exploratory or transformational, show a candidate metric separates them, and only then apply it to machines.
“L3 — the binding link: a validated instrument for transformational creativity. Construct a gold-standard set of historical H-creative artifacts labelled by domain experts as exploratory or transformational; demonstrate that a candidate metric separates them; only then apply it to machine output.”
Intelligence Measurement
Whether benchmark scores are measurements of a latent ability or descriptions of one particular test.
M4, the binding link: fit latent ability for N models on instrument A, pre-register point predictions of their scores on an instrument B built by a disjoint team from a disjoint generator, publish before B exists, and report calibration.
“M4 — the pre-registered out-of-distribution prediction. Fit latent ability for N models on instrument A. Pre-register point predictions of their scores on instrument B, built by a disjoint team from a disjoint generator.”
Future Education Systems
Whether reported education effect sizes survive measurement on an independently constructed standardised instrument.
Dual-instrument reporting: every study reports its effect size on both a local measure and an independently constructed standardised measure.
“Experiment 4 — dual-instrument reporting. Binding, and the highest-value action available. Every study reports effect size on both a local and an independently constructed standardised measure. This is a reporting norm rather than an experiment; it costs one extra assessment; and it would have prevented the two-sigma myth, most of the intelligent-tutoring literature’s overclaiming, and the retracted meta-analysis.”
AI Governance
Whether a measured dangerous-capability score is a ceiling on the model or an artefact of how hard anyone tried to elicit it.
The elicitation-gap curve (L3): on fixed weights and a fixed dangerous-capability task, measure best performance under naive prompting and then under 1, 4, 16 and 64 expert-hours of scaffolding and tool provision by a blind red team, publishing performance against expert-hours with a fitted plateau.
“Experiment one: the elicitation-gap curve (L3). Take a fixed set of weights and a fixed dangerous-capability task. Measure best performance under naive prompting; then under 1, 4, 16 and 64 expert-hours of scaffolding, tool provision and prompt engineering by a red team blind to the earlier results.”
Artificial Scientists
Whether an automated verifier can separate true claims from plausible false ones at a characterised error rate.
An ROC curve for an automated verifier: run it against roughly 200 published claims of known replication status and report sensitivity and specificity as a full curve rather than a single accuracy point.
“Experiment one, and the one that unblocks the field: an ROC curve for an automated verifier. Take a corpus of roughly 200 published claims with known replication status — the large reproducibility-project corpora, or the cancer-biology set Eve was run against — and measure an automated verifier’s sensitivity and specificity against ground truth.”
Cognitive Liberty
Whether neural decoding generalises to a new subject without subject-specific training data, which governs the whole threat model.
A zero-shot cross-subject decoding benchmark: decoding accuracy on a held-out subject with zero subject-specific training data, reported against the within-subject ceiling, on a public evaluation set fixed before the models are built.
“A zero-shot cross-subject decoding benchmark. Decoding accuracy on a held-out subject with zero subject-specific training data, reported against the within-subject ceiling, on a public evaluation set fixed before the models are built.”
Human Cognitive Augmentation
Which cognitive-augmentation effects survive adjustment for publication bias.
Apply publication-bias adjustment — selection models, PET-PEESE, robust Bayesian multilevel meta-analysis — to the existing brain-training, working-memory-training and stimulation corpora, and report the adjusted pooled effect and its credible interval per corpus.
“Experiment one, and it needs no new data. Apply publication-bias adjustment — selection models, PET-PEESE, robust Bayesian multilevel meta-analysis — to the existing brain-training, working-memory-training and stimulation corpora. Measurement: the adjusted pooled effect and its credible interval per corpus.”
Multi-Agent Intelligence Systems
Whether published multi-agent gains survive comparison with single-agent methods at equal compute.
The compute-matched control: for each method in MASLab hold total tokens, wall clock and dollars constant and compare against greedy decoding, self-consistency at matched sample count, and best-of-N with a reward model at matched N, reporting paired differences with confidence intervals.
“Experiment one, and the binding link: the compute-matched control. For each method in MASLab, hold total tokens, wall clock and dollars constant and compare against three single-agent controls — greedy decoding, self-consistency at matched sample count, and best-of-N with a reward model at matched N.”
Agentic Autonomy and Control
Whether authorization architecture alone can hold the conversion rate from prompt injection to unauthorized agent action near zero at a task-completion cost enterprises will pay, given that no model reliably separates instructions from data.
An adversarial architecture trial: the same agent, tasks and standing red team run once with coarse long-lived credentials as standard practice and once with every credential short-lived, audience-bound, attenuated to the task and revocable, measuring injection-to-unauthorized-action conversion as the primary endpoint and task completion as the cost axis.
“The decisive experiment is an adversarial architecture trial: the same agent, the same tasks, the same standing red team, run once with the coarse long-lived credentials that are standard practice and once under an architecture in which every credential is short-lived, audience-bound, attenuated to the task and revocable, with the conversion rate from injection to unauthorized action as the primary endpoint and task completion as the cost axis.”
AI-Cyber Convergence
Whether AI-driven automation favours cyber attackers or defenders, decided by whether autonomous patch cadence across a real fleet can beat autonomous exploitation cadence on the same disclosed vulnerabilities.
A scored, adversarial head-to-head that pits an autonomous defender against an autonomous attacker on the same live target, measuring whether mean-time-to-patch across a representative fleet falls below mean-time-to-exploit on the same disclosed flaws.
“The decisive experiment this brief calls for is a scored, adversarial contest that pits an autonomous defender against an autonomous attacker on the same live target”
Synthetic Data and Training Provenance
Does training a frontier-scale model on a current, synthetically contaminated web crawl measurably degrade capability, calibration and tail coverage relative to a pre-contamination crawl at matched compute and curation?
The crawl-vintage comparison: train two otherwise identical frontier-scale models, one on a web snapshot predating late 2022 and one on a current snapshot, under the same curation pipeline and compute budget, and measure capability, calibration and tail coverage on contamination-proof evaluations.
“The decisive experiment is the crawl-vintage comparison: train two otherwise identical frontier-scale models, one on a web snapshot predating late 2022 and one on a current snapshot, under the same curation pipeline and compute budget, and measure capability, calibration and tail coverage on contamination-proof evaluations. Any frontier laboratory could run it today; none has published one.”
Embodied AI and Robot Foundation Models
Whether developer-reported success rates of robot foundation models survive independent, statistically powered evaluation in facilities and conditions the developer did not curate.
A frontier vision-language-action model evaluated by independent labs in unseen facilities, with pre-registered success criteria and hundreds of trials per condition, holding its developer-reported success rates within stated confidence intervals.
“The decisive test is independent, statistically powered replication: a frontier vision-language-action model, run by evaluators its developer does not employ, in facilities it has never seen, holding its reported success rates within stated confidence intervals.”
Fault-Tolerant Quantum Computing
Does exponential logical-error suppression survive the code distances, run times and correlated-error environment that useful algorithms require?
Run a single logical qubit at code distance 15 or more with real-time decoding, sustaining a logical error rate near one in a million cycles for hours despite radiation-induced burst errors.
“The decisive demonstration is endurance at scale: a single logical qubit at distance 15 or more, decoded in real time, holding a logical error rate near one in a million cycles for hours, through the radiation bursts that currently floor repetition-code performance at one error in ten billion cycles.”
Privacy-Enhancing Computation
Does the privacy parameter used by the largest differentially private release in the world actually stop a funded attacker from re-identifying the people in it?
An independent re-identification study run against the 2020 US redistricting file at its production privacy-loss budget of 19.61, by a team with no stake in the answer, using the commercial identity data an attacker would buy.
“The decisive experiment is a published re-identification study run against a production release at its production privacy parameter, by a team with no stake in the answer, buying the same commercial data an attacker would buy.”
AI Companions
Whether companion use substitutes for human contact or scaffolds it, and whether the observed relationship between heavier use and worse loneliness is causal.
A preregistered randomised trial of at least six months with companion use as the randomised factor, an active human-contact comparison rather than a waitlist, and objectively measured social contact as the primary outcome.
“The decisive experiment is a preregistered randomised trial of at least six months in which companion use is the randomised factor, the comparison is an active human-contact condition rather than a waitlist, and the primary outcome is objectively measured social contact rather than a self-reported loneliness score.”
Autonomous Materials Discovery
Does an autonomous materials pipeline convert candidates into confirmed new materials at a better rate, and a lower cost per confirmed material, than a matched human group?
A blinded, matched-cost, end-to-end yield trial: one target list of predicted-stable compositions with no prior literature report, split between an autonomous laboratory and a matched human group on equal budget and time, with every product from both arms identified by an independent crystallography group that does not know which arm produced it. The measured quantity is the full funnel, ending in cost per confirmed novel phase.
“The decisive experiment is a blinded, matched-cost, end-to-end yield trial, and nobody has run it.”
V — Energy Systems
Commercial Fusion
Whether a fusion plant can breed more tritium than it burns, the assumption under every commercial roadmap.
Close a tritium loop: breed tritium in a blanket, extract it, purify it, inject it, burn it, and account for the inventory, at any scale and any breeding ratio, published.
“The highest-value experiment in the subject is the one nobody has run: close a tritium loop. Breed tritium in a blanket, extract it, purify it, inject it, burn it, and account for the inventory — at any scale, with any breeding ratio, published.”
Advanced Fission
Whether thorium conversion in a molten salt reactor works at a ratio anyone outside the operating institute can verify.
A published conversion ratio from TMSR-LF1 at Wuwei, in a peer-reviewed source, above the 0.1 the institute has reported.
“The natural experiment worth watching in salt is at Wuwei. TMSR-LF1 is the only operating molten salt reactor in the world and the only vehicle testing thorium conversion in a salt. The decisive readout is not another announcement but a published conversion ratio, from a peer-reviewed source, above the 0.1 the institute has reported”
Space-Based Solar Power
Whether power beamed from orbit can be recovered on the ground as usable electricity rather than merely detected.
The 2023 Caltech orbital mission, the only space-to-ground beaming attempt on record, which returned a detection rather than a transfer, lost 14.7% of transmitted power in eight months and jammed its deployable structure twice.
“The decisive experiment has already been run once and it produced a negative result that the field has not fully absorbed. The 2023 Caltech mission is the only space-to-ground beaming attempt on record.”
Superconducting Infrastructure
Whether superconducting cable works at transmission voltage and length, or stays confined to short distribution-level demonstrators.
Building SuperLink, the 110 kV, 500 MW, roughly 12 km Munich link, rather than the 120 to 150 metre substation demonstrator that is all that exists.
“SuperLink is the experiment that would settle the transmission question, and it has not been built. The full project is 110 kV and 500 MW”
Wireless Energy Transmission
Whether dynamic wireless road charging survives commercially once built, or is removed after the demonstration ends.
The deployment record read as results: Gotland's road dismantled at the end of its project, and all four original Korean OLEV bus lines shut down with the 2014 Gumi commercial route abandoned.
“The decisive experiments in this subject have largely been run, and several of them returned answers by being taken apart. The deployment record below is the natural experiment, and it should be read as a set of results rather than as a list of projects.”
High Temperature Superconductors
Whether REBCO tape price falls as manufacturing volume rises, or is set by deposition physics rather than by scale.
Whether the 77 K self-field price quoted in the peer-reviewed literature falls below roughly 50 euro per kA-m in general availability, rather than in a single contracted forward delivery, before 2030.
“The clearest single test of the position this brief takes is a price series: whether the 77 K self-field figure quoted in the peer-reviewed literature falls below roughly 50 €/kA·m in general availability, rather than in a single contracted forward delivery, before 2030.”
Energy Storage Revolutions
Whether long-duration storage plants deliver the energy, cycles, availability and round-trip efficiency their nameplates claim.
Publishing a single year of measured delivered energy, achieved cycles, availability and round-trip efficiency from Rudong, Goderich, Cambridge or Big Stone.
“And the decisive missing experiment is the one nobody is running: publishing operating data. A single year of measured delivered energy, achieved cycles, availability and round-trip efficiency from Rudong, Goderich, Cambridge or Big Stone would settle more of this subject than any new demonstrator.”
Geothermal Megaprojects
Which half of the shale completion toolkit transfers to hot crystalline rock, and therefore whether enhanced geothermal can produce commercial flow rates.
Utah FORGE's controlled comparison of unpropped against propped stimulation in the same formation, which returned 0.7 kg/s against 26 kg/s.
“The decisive experiment already ran and its result is the thirty-seven-fold number. Utah FORGE compared unpropped against propped stimulation in the same formation and got 0.7 kg/s against 26 kg/s”
Ocean Thermal Energy Conversion
Whether an operating OTEC plant delivers positive net power once its parasitic loads are metered, rather than only gross.
Instrument an existing plant at Kumejima or Kailua-Kona and publish the net output and the parasitic breakdown at a stated temperature difference over a stated period.
“The experiment that would actually settle the central question has a clean specification and nobody is running it. Instrument an existing plant — Kumejima or Kailua-Kona, both of which have run for over a decade — and publish the net output and the parasitic breakdown at a stated temperature difference over a stated period.”
Hydrogen Economies
Whether the 2024-26 collapse in announced hydrogen projects was a one-off shakeout or a durable trend.
More than 100 GW of announced electrolysis must take a final commitment before the end of 2027 or lose any chance of operating by 2030, with 22 Mt of announced production at risk without a decision by early 2027.
“Two decisive experiments are already scheduled and need no new apparatus. More than 100 GW of announced electrolysis either commits before the end of 2027 or loses any chance of operating by 2030”
Advanced Battery Technologies
Whether solid-state batteries reach production vehicles on the announced schedule, or the deadline is met only by shrinking the scope to limited batches.
Whether Toyota ships a solid-state vehicle in 2027 or 2028, read not off the car but off Idemitsu Kosan's solid-electrolyte plant, due end of 2027 at several hundred tonnes a year.
“The decisive live experiment for solid-state is dated and public: whether Toyota ships in 2027 or 2028. The enabling test is not the vehicle but the electrolyte plant.”
Small Modular Reactors
Whether modular construction actually delivers the schedule and cost learning its economic case assumes, or repeats the negative-learning record of large nuclear.
The Changjiang side-by-side: a 125 MWe ACP100 built against a 1,100 MWe unit on one site, one owner, one regulator, with the follow-on test being whether a second ACP100 unit is materially faster than the first.
“Changjiang is a controlled comparison no study could have designed. Two reactors, one site, one owner, one regulator, one labour market, four months apart, differing by a factor of nine in size and by the modular claim itself.”
Nuclear Waste Solutions
Whether the American dry-storage canister population is developing chloride-induced stress corrosion cracking.
A programme of dry-storage canister inspections at scale, with rigorous criteria and mature techniques, against a field record that is close to two horizontally stored canisters inspected at Calvert Cliffs in 2012.
“The independent review board's phrasing — very few inspections, none with rigorous criteria and mature techniques — describes an experiment that is available, affordable and not being done at scale. It is the cheapest decisive experiment in this brief.”
Lunar Energy Infrastructure
Whether a lunar outpost should be powered by vertical solar plus storage or by a 40-to-100 kWe reactor at a named site.
A published site-specific trade study: worst-case darkness duration for a named candidate ridge, a stated survival-power duty cycle, and a landed-mass comparison between vertical solar plus storage and a 40-to-100 kWe reactor over ten years.
“The experiment that would settle the central question is a published trade study, and it is cheap. Nothing needs to be launched.”
Zero Carbon Industrial Systems
Whether limestone calcined clay cement can be admitted to the European concrete standard, which is what gates the sector's largest near-term emissions cut.
A standards trial generating durability and strength evidence on LC3-50 concrete over the periods a code committee accepts, so that EN 206 can admit CEM II/C.
“The cheapest high-value experiment is a standards trial, not a laboratory one. EN 206 admitting CEM II/C requires durability and strength evidence on LC3-50 concrete over the periods a code committee accepts.”
Artificial Photosynthesis
Whether direct artificial photosynthesis beats photovoltaics plus electrolysis on the same land area.
A year-long side-by-side trial on equal land under identical irradiance, best photocatalytic or photoelectrochemical array against a commercial photovoltaic field feeding a commercial electrolyser, publishing hydrogen delivered, capital cost, water consumed, degradation and downtime.
“The highest-value experiment is the comparison the literature avoids, run properly. Take a fixed land area under identical irradiance; on one half deploy the best available photocatalytic or photoelectrochemical array, on the other a commercial photovoltaic field feeding a commercial electrolyser”
Ultra Efficient Computing Energy Systems
Whether efficiency gains per computation reduce total computing energy, or are consumed by rebound.
A hyperscaler's own annual accounts, reporting a 33-fold reduction in energy per median text prompt alongside a 37% increase in total electricity consumed in the same year.
“The most useful natural experiment is one company's annual accounts. A hyperscaler's own disclosures for the same year report a 33-fold reduction in energy per median text prompt and a 37% increase in total electricity consumed, the largest load growth in its history.”
Industrial Heat Electrification
Whether electrothermal storage delivers industrial steam at the cost, efficiency and availability its vendors claim, once measured by someone other than the vendor.
A year of independently metered operating data from one of the 100 MWh-class thermal batteries running under commercial duty: steam delivered, realised charging price, round-trip efficiency as operated, and availability against the host plant's schedule.
“The decisive result is a year of independently metered operating data from one of the 100 MWh-class thermal batteries now running under commercial duty: tonnes of steam delivered, the realised electricity price paid to charge, round-trip efficiency as operated, and availability against the host plant's schedule.”
Perovskite Tandem Photovoltaics
Whether production perovskite-silicon tandem modules degrade slowly enough in real fields (near 1% per year) to support silicon-grade warranties and bankability.
A published, third-party, multi-year degradation dataset on production tandem modules operating in customer fields alongside silicon controls - fielded product, independent instruments, several years, published rates.
“The single result most able to change this assessment is a published, third-party, multi-year degradation dataset on production tandem modules operating in customer fields alongside silicon controls.”
Floating Offshore Wind
Whether the moorings, connectors and dynamic cables of floating wind achieve a failure rate per component-year that lenders and insurers can price.
Pool and publish the component-level service record of the entire global floating fleet, with installation date, inspection history, failure mode, repair duration and lost production for every mooring line, anchor, dynamic cable, bend stiffener and buoyancy module.
“It requires no new hardware, no new science and no new vessel. It requires operators to agree to disclose, which is why it does not exist.”
Wave and Tidal Energy
Whether tidal stream is actually on a cost-reduction path, or whether its contracted strike prices reflect a cost that does not fall.
The contracted UK tidal-stream fleet publishing, on a common basis, its achieved capacity factor, availability, operating cost per megawatt-hour, and every component recovery and replacement event with duration and cause.
“The United Kingdom’s ring-fenced contracts commit a small fleet of tidal-stream capacity to deliver across the second half of this decade at a known strike price.”
VI — Climate & Planetary Engineering
Climate Engineering
Whether the projected Sahel precipitation loss under stratospheric aerosol injection is a robust physical result or an artefact of model structure.
A GeoMIP-successor model intercomparison with common injection strategies at higher resolution, evaluated on monsoon dynamics and reported per model rather than as an ensemble mean.
“First, a GeoMIP-successor model intercomparison with common injection strategies at higher resolution, evaluated on monsoon dynamics and reported per model rather than as an ensemble mean. It requires no release, no permit and no consent, and would settle whether the Sahel result is robust or model-structural.”
Carbon Capture at Scale
Whether geological storage can hold its permitted injection rate for years at a time, which sets the confidence interval on every carbon removal pathway.
A geological storage complex operating at its permitted injection rate for five consecutive years, with Northern Lights as the test and Gorgon as the counter-case.
“Fourth, a geological storage complex operating at its permitted injection rate for five consecutive years. Gorgon is the counter-case and Northern Lights is the test. This is the single experiment that would most change the confidence interval on every removal pathway, because every one of them ends in a well.”
Desert Greening
Whether the measured greening of the Sahel coincides with a gain or a loss in biodiversity on the same ground.
Co-measuring greenness and counterfactual biodiversity on the same plots in Nigeria and Senegal, with matched controls and NDVI or biomass and species richness taken from the same footprints in the same years.
“One: co-measure greenness and counterfactual biodiversity on the same plots. Nigeria and Senegal, matched controls, NDVI or biomass and species richness from the same footprints in the same years. This resolves the sharpest contradiction in the subject and requires a field season, not a programme.”
Ocean Engineering
Whether the 1 mm seabed blanketing limit written into draft deep-sea mining regulation is an achievable specification or an aspiration.
A full-scale collector trial at commercial throughput with plume instrumentation at 5, 50 and 500 kilometres, measuring sediment accumulation depth and suspended load against range and time over at least one seasonal cycle.
“The measurement: sediment accumulation depth and suspended load as functions of range and time, over at least one seasonal cycle, with the collector operating at commercial throughput rather than at pre-prototype scale. The decision it closes: whether a 1 mm blanketing limit is a specification or an aspiration.”
Arctic Engineering
Which mechanism actually drives the observed displacement of embankments built on ice-rich permafrost.
Running settlement rods, lateral displacement and inclinometer arrays and distributed temperature-sensing strings together for at least five years on an already-instrumented ice-rich permafrost embankment, and publishing an attribution of 80% or more of displacement to one mechanism.
“Experiment one: settle the failure mode on ground that is already instrumented. Take an embankment on ice-rich permafrost — the Inuvik–Tuktoyaktuk corridor is the obvious candidate — and run vertical settlement rods, lateral displacement and inclinometer arrays, and distributed temperature-sensing strings together for at least five years.”
Weather Modification
Whether seeded ice-nucleating particles persist downwind long enough for the design assumption under every cloud-seeding trial to hold.
A one-season persistence test instrumenting ice-nucleating-particle concentrations downwind of seeded storms at one day, one week and one month against matched unseeded control sites.
“The cheapest high-value experiment available today is the persistence test, and nobody is running it. Instrument ice-nucleating-particle concentrations downwind of seeded storms at intervals of one day, one week and one month, against matched unseeded control sites, for one season.”
Atmospheric Management
Whether iron-salt chlorine chemistry has a stable sign for net methane removal across the ambient atmospheric envelope.
Chamber measurement of chlorine yields extending the Fe(III)/sea-salt work across the NOx, humidity, ozone and SO2 parameter space, establishing where net methane removal changes sign.
“Experiment one, and the cheapest high-value move available today: chamber chlorine yields across the full ambient envelope. Extend the Fe(III)/sea-salt chamber work to the NO x , humidity, ozone and SO 2 parameter space implied by the sign-dependence preprint, and establish where net methane removal changes sign.”
Water Infrastructure Megaprojects
Whether a widely cited systematic review mislabelled desalinated-water and brine quantities, leaving global brine understated by about 1.5 times.
Obtaining Jones et al. (2019) and reading its two headline quantities against the brine-to-product ratio implied by the 2026 review's own number.
“The first required experiment is not an experiment. It is a library retrieval, and it should be done before anything else in this brief is treated as settled. Obtain Jones et al. (2019) and read its two headline quantities.”
Planetary Cooling Concepts
Whether the cloud-cover-dominance result behind marine cloud brightening, and the climate sensitivity revision implied by it, survives independent replication.
Replicating the cloud-cover-dominance result on existing satellite archives with a different volcano or ship-track dataset and an attribution method other than the original machine-learning implementation.
“The cheapest high-value experiment in this brief requires no release, no permit and no new instrument. Replicate the cloud-cover-dominance result using existing satellite archives, a different volcano or ship-track dataset, and an attribution method that is not the original machine-learning implementation.”
Continental Irrigation Systems
Whether a continental-scale water transfer perturbs precipitation and surface temperature enough to make the scheme climate engineering.
Running a NAWAPA-scale transfer in a modern earth-system model and reporting precipitation and temperature change against transferred volume for donor, recipient and teleconnected domains.
“One. Run a NAWAPA-scale transfer in a modern earth-system model. Report change in precipitation and surface temperature as a function of transferred volume, for the donor and recipient domains and for the teleconnected regions. The established scaling result (Chen & Xie 2010) makes this a well-posed experiment; the continental case has never been run.”
Floating Cities
Whether any jurisdiction allows a floating dwelling and its berth to be registered as a single financeable unit, and which statutes block it.
Attempting, in a cooperative jurisdiction, to register a security interest over an existing floating dwelling and its berth as one unit, and recording which office refuses and on what statutory ground.
“Experiment 1: the registry filing test. Take an existing Dutch or Scandinavian floating dwelling and attempt, in a cooperative jurisdiction, to register a security interest over both the module and its berth as a single unit. Record precisely which office refuses, and on what statutory ground.”
Polar Development
Whether pumping seawater onto Arctic ice can work as a global climate intervention rather than only as a regional ice-preservation measure.
Zampieri and Goessling's model of sea-ice-targeted geoengineering by seawater pumps, which returned a global annual-mean near-surface air temperature reduction of 0.02 K.
“Link four has been tested directly and it failed. Zampieri and Goessling modelled sea-ice-targeted geoengineering by seawater pumps and found global annual-mean near-surface air temperature reduced by 0.02 K, against real regional Arctic cooling.”
Sustainable Megacities
Whether tenure security is separable from and prior to physical upgrading in informal settlements.
A settlement-level four-arm trial - tenure security alone, physical upgrading alone, both, neither - measured on health, school retention, household investment and income over at least five years.
“3. The upgrading trial that separates tenure from bricks. A settlement-level design with four arms — tenure security alone, physical upgrading alone, both, neither — measured on health, school retention, household investment and income over at least five years.”
Geoengineering Governance
Whether a functioning treaty body can extend an existing assessment framework to cover a solar geoengineering technique.
The London Protocol intersessional correspondence group's October 2026 report on how existing instruments apply, with marine cloud brightening among the techniques considered for Annex 4 listing.
“That is a live test of whether a functioning treaty body can extend a working framework to a solar geoengineering technique , and its outcome is the single most informative datum this subject will produce this decade.”
Biodiversity Restoration
Whether the finding that only a small minority of taxa benefit from protected areas is a property of protected areas or a property of Finland.
Running the Santangeli design continentally: a counterfactual, multi-taxon, matched-site occupancy comparison over four decades, replicated across several biogeographic regions.
“One: run the Santangeli design continentally. A counterfactual, multi-taxon, matched-site occupancy comparison over four decades, replicated across several biogeographic regions rather than one country, would establish whether the Finnish result — a small minority of taxa benefiting, mostly through slower decline — is a property of protected areas or a property of Finland.”
Coastal Defense Systems
Whether a coastal authority anywhere has lawful power to withdraw protection, what standing a landowner has to stop it, and what compensation follows.
An audit of the statute book across twenty coastal jurisdictions establishing the power to withdraw protection, landowner standing to prevent it, and the compensation consequences.
“Six. Audit the statute book. For twenty coastal jurisdictions, establish whether an authority has a lawful power to withdraw protection, what standing a landowner has to prevent it, and what compensation follows.”
Climate Migration Planning
Whether the most-cited claims in climate-migration policy documents are supported by the sources cited for them.
A systematic citation audit of the fifty most-cited claims in climate-migration policy documents.
“A systematic citation audit of the fifty most-cited claims in climate-migration policy documents is a weekend of work and would probably be the highest-value output in the subject.”
Planetary Stewardship
Whether disagreement over national planetary-boundary allocations is normative rather than empirical.
Recomputing one national allocation of the published nitrogen ceiling under equal-per-capita, historical-responsibility, capability and grandfathering rules, propagating the ceiling's interval through each, and publishing the spread.
“The predicted result is that the between-rule spread exceeds the within-rule uncertainty for most states, and that for some the two are comparable — which would establish, quantitatively, that the argument about allocation is normative rather than empirical.”
Carbon Removal Verification
Do the market's registries, applied to the same physical removal deployment, issue the same number of tonnes?
A registry round-robin: one instrumented removal deployment credited in parallel under Isometric, Puro.earth and Verra rules, with all measurements public and the spread in issued tonnes published as the result.
“The decisive test is a registry round-robin: one instrumented removal deployment credited in parallel under the rules of Isometric, Puro.earth and Verra, with every measurement public and the spread in issued tonnes published as the result.”
Compound Climate Hazards
Whether dependence models fitted to the observed climate record predict joint hazard exceedances better than the independence assumption built into design standards and pricing.
A coordinated out-of-sample verification: fit the field's dependence models to the instrumental record up to a cutoff, freeze them, and score their predicted joint exceedances against the decades observed since, region by region and hazard pair by hazard pair, with independence as the null model.
“Fit the field’s dependence models — the copula families, the conditional models, the model-ensemble dependence structures — to the instrumental record up to a cutoff, freeze them, and score their predicted joint exceedances against the decades already observed since, region by region and hazard pair by hazard pair, with independence as the null model.”
Wildfire Systems
Do fuel-treatment severity reductions measured mostly under moderate fire weather hold under the extreme fire weather that now produces most burned area?
A registered, prospective measurement protocol across wildfire encounters with recently treated forest in extreme seasons - pre-registered treatment polygons, standardized severity metrics, weather percentile at encounter, results published whether or not the treatment held.
“The decisive test is already running: each extreme season drives wildfire into thousands of hectares of recently treated forest, and a registered, prospective measurement protocol across those encounters would settle whether severity reductions measured mostly under moderate weather survive the conditions that now do most of the burning.”
Earth-System Digital Twins
Does an operational Earth-system digital twin improve real local decisions, scored on outcomes, over the incumbent forecast chain?
A pre-registered paired trial in which matched decision units make real operational choices, one arm served by the twin chain and one by the incumbent products, with the arms scored on decision outcomes rather than forecast skill.
“The decisive test is a pre-registered paired trial in which matched decision units make real operational choices, one arm served by the twin chain and one by the incumbent products, and the arms are scored on decision outcomes rather than on forecast skill.”
Climate Overshoot and Lock-In
Is the reversibility assumption underneath overshoot planning sound for the largest tipping element in the climate system?
Sustained measurement of the Atlantic meridional overturning circulation by the existing observing arrays, read together with the South Atlantic freshwater-transport fingerprint, to see whether a decline coherent across latitudes coincides with movement toward the model-identified tipping regime.
“If the overturning arrays show a decline that is coherent across latitudes and the fingerprint moves toward the model-identified tipping regime, the reversibility assumption underneath overshoot planning fails for the largest single element in the system, and it fails while the temperature is still rising.”
AI Weather Prediction
Has machine learning replaced numerical weather prediction, or only its cheapest stage?
An end-to-end machine-learning forecast system initialised from raw observations, with no physics-based analysis anywhere in the chain, evaluated against independent observations on extreme-value metrics rather than against reanalysis.
“The decisive demonstration is an end-to-end machine-learning forecast system, initialised from raw observations with no physics-based analysis anywhere in the chain, evaluated against independent observations on extreme-value metrics.”
Extreme Heat Survivability
Does the housing stock of a heat-exposed city keep people alive through a multi-day heat event once the cooling stops, and by how wide a margin?
A stratified measurement of indoor temperature and humidity, with power-state metadata, across ordinary dwellings through a real multi-day heat event that includes a real power interruption.
“The decisive result is a measured indoor-conditions dataset spanning a real multi-day heat event that includes a real power interruption, in a stratified sample of ordinary dwellings.”
Blue Food Systems
Can land-based grow-out of a high-value farmed fish be produced at a cost that competes with an open net pen, across complete production cycles rather than in a projection?
An audited, multi-cohort, full-cycle cost of production published from a commercial-scale land-based grow-out facility.
“The decisive result is an audited, multi-cohort, full-cycle cost of production from a commercial-scale land-based grow-out facility.”
VII — Civilization-Scale Infrastructure
Continental Transportation Systems
Is converting a corridor's track gauge cheaper than living with the break of gauge at the volumes that corridor actually carries?
Measure the per-container cost of a gauge break on a defined China-EU or Ukraine-EU corridor as dwell time, transhipment handling cost and reliability penalty, reported as a distribution, and compare it against the amortised cost of conversion at the corridor's actual volume.
“That comparison is computable with data that already exists, it decides whether the Ukrainian conversion programme is rational, and nobody in the Anglophone literature appears to have computed it. It is the most tractable open question in this brief.”
Northern Development Corridors
Does corridor capital deliver more cost-of-living benefit per northern household than the same money spent on air service, local energy or local food?
Compare corridor capital per household served, amortised over asset life with the maintenance path included, against the cost-of-living reduction the same capital buys through runway extension and air-service reliability, local energy, or local food production.
“The decisive measurement is a comparison nobody runs. Take corridor capital per household served, amortised over the asset life with the maintenance path included, and compare the cost-of-living reduction it delivers against the same capital spent on runway extension and air-service reliability, on local energy, and on local food production.”
Smart Cities
Did smart-city programmes improve the outcomes they were sold on, measured against comparable cities that were shortlisted but not selected?
Run the difference-in-differences already sitting in India's Smart Cities Mission: 100 cities with staggered implementation against shortlisted-but-not-selected controls, on outcome metrics the cities were reporting before the programme began.
“The cheapest decisive study in the subject already has its comparison group and has not been run. India's Mission gave 100 cities staggered implementation across ten years, and it shortlisted cities that were not ultimately selected.”
High Speed Transit Networks
How much of the spread in high-speed rail cost per kilometre is terrain, standards, mitigation and procurement rather than national competence?
Decompose per-kilometre outturn cost for at least five national high-speed programmes into terrain-forced, standard-forced, mitigation-forced and procurement-forced components, on one explicitly stated scope boundary and a single price year.
“One: the four-way cost decomposition. Take at least five national high-speed programmes and separate per-kilometre outturn into terrain-forced, standard-forced, mitigation-forced and procurement-forced components, on a single scope boundary that states explicitly whether it includes rolling stock, stations, depots, land, electrification and interest during construction, at a single price year.”
Underground Cities
What a subsurface cadastre actually costs to build, a price every downstream underground planning instrument has assumed without testing.
Building one subsurface cadastre for a mid-sized city from geophysical proxy indicators, cheap monocular or photogrammetric capture of accessible voids and existing utility records, and publishing coverage, accuracy and total cost per square kilometre.
“The deliverable is not the map; it is the price of the map. Every downstream instrument has been blocked for thirty-five years on the presumption that this is expensive, and the presumption is untested.”
Arcologies
At what population, vertical extent and compartment count performance-based fire engineering stops being able to demonstrate a margin against tenability limits.
Running the current performance-based design toolchain against a family of hypothetical occupancies scaling population, vertical extent and compartment count, recording where a demonstrable margin with defensible input uncertainty is lost.
“1. Find the population threshold at which performance-based fire engineering stops being demonstrable. This is first because it bounds everything else and because nobody has published it.”
Automated Construction Systems
Whether construction's measured productivity stagnation is substantially an artefact of deflator-based output measurement.
Building a hedonic construction output index adjusted for regulated performance attributes, back-cast thirty years across two national statistical systems, and comparing its growth rate against the current deflator-based series.
“The decisive comparison is the index's growth rate against the current deflator-based series.”
Spaceports
Is launch cadence limited by ground range and airspace institutions rather than by vehicles, and what is a range window actually worth?
Publish a crossover of launch cadence against aviation delay cost on a dense route network, the measurement that pricing or allocating range windows and airspace closures requires.
“This is the binding link and it is institutional. The measurement is a published crossover of launch cadence against aviation delay cost on a dense route network. No such mechanism exists in any jurisdiction and no such study has been found.”
Autonomous Supply Chains
Whether warehouse robotics actually raises picks per labour hour at constant SKU mix and order profile.
A before-and-after measurement of picks per labour hour on the same SKU mix and order profile, with a control facility.
“One: picks per labour hour, before and after, same SKU mix and order profile, with a control facility. This is the measurement the entire warehouse-robotics case rests on, it is trivially available to any operator, and it is essentially unpublished.”
Future Ports and Shipping
Whether terminal automation's productivity gains come from the technology itself or from implementation, integration and training.
Computing between-terminal productivity variance among automated terminals against manned ones; higher variance among the automated would show implementation rather than technology is the causal factor.
“The highest-value unrun experiment in this subject is cheap, and it is a variance test. If terminal automation's gains are real but conditional on integration and training — which is what the Mediterranean study's own hedge says — then automated terminals should show higher between-terminal variance in productivity than manned ones, not a uniform advantage.”
Energy Corridors
Whether a cross-border capacity-allocation and cost-recovery regime holds when the exporting system is itself short of power.
A rule set compensating jurisdictions crossed but not served by a corridor and preserving capacity rights against a national curtailment override, tested on revealed behaviour in a correlated stress event.
“A rule set under which a jurisdiction crossed by, but not terminating, a corridor is compensated, and under which capacity rights survive a national curtailment override. The measurement is revealed behaviour in a correlated stress event: does the corridor deliver across a border when the exporting system is itself short?”
Intercontinental Rail Systems
Does existing traffic across each candidate strait come anywhere near the throughput a fixed intercontinental rail link would have to carry?
Assemble, for each of the three crossings, current annual tonnage and passenger movements by ferry, short-sea shipping, air and land detour, split by direction and commodity value class, and express existing flow as a ratio to the throughput a link would need.
“The output is a single ratio per crossing: existing flow against the throughput a link would need. Nobody has published it, and it decides everything downstream.”
Future Housing Systems
What fraction of an industrialised building cost reduction reaches the sale price rather than the land price, in constrained against elastic markets.
Delivering an identical industrialised product into a high-elasticity and a low-elasticity market, matched on specification and boundary, and measuring what share of the hard-cost reduction appears in sale price versus land price.
“The second experiment is the pass-through test, and it is the one that decides the programme. Deliver an identical industrialised product into a high-elasticity and a low-elasticity market, matched on specification and boundary, and measure what fraction of the hard-cost reduction appears in the sale price versus the land price.”
Megaproject Governance
Whether megaproject cost overruns are forecasting error or strategic misrepresentation.
A two-by-two of competitive against non-competitive approval crossed with externally set against promoter-set uplift, with outturn as the dependent variable and difference-in-differences as the estimator.
“Three, and the highest-leverage unrun study in the field: the discriminating test between error and misrepresentation. A two-by-two: competitive against non-competitive approval, crossed with externally set against promoter-set uplift, with outturn as the dependent variable and difference-in-differences as the estimator.”
Infrastructure Resilience
Whether restoration capacity rather than failure probability is the binding resilience variable, which would mean much hardening expenditure targets the wrong term.
Publishing the time-to-restore distribution for one regulator's jurisdiction over one decade, with damage, event class and resources brought to bear recorded, then running a variance decomposition against physical damage.
“The decisive analysis is a variance decomposition: does time-to-restore vary more across events with similar physical damage than physical damage itself varies?”
Robotics in Infrastructure
Is construction's flat measured productivity a real stagnation, or an artefact of price and quality measurement that hides genuine robotic gains?
Recompute the standardised-product output-per-worker-hour series with an explicit hedonic adjustment for code-driven quality content, testing directly whether the physical productivity measures are themselves contaminated.
“Two: the same series with an explicit hedonic adjustment for code-driven quality content, which is the direct test of whether physical measures are themselves contaminated and therefore the sharpest available challenge to Goolsbee and Syverson's strongest evidence.”
Industrial Ecology
Is waste law the binding constraint on industrial symbiosis, so that a reclassification pipeline rather than a park-building programme is the instrument that works?
Run a difference-in-differences on end-of-waste reclassification: a by-product stream reclassified out of waste law on a known date against a comparable non-reclassified control, measuring exchange formation counts before and after with material price controlled.
“The decisive experiment is a difference-in-differences on end-of-waste reclassification. Take a specific by-product stream reclassified out of waste law on a known date, take a comparable non-reclassified stream as control, and measure exchange formation counts before and after, controlling for the material's price.”
Circular Infrastructure Systems
What does it cost to certify reclaimed structural steel, and what fraction of a demolition batch actually passes?
Put a defined batch of demolition steel sections through a published test protocol with a named certifying body issuing or refusing a document, then publish the cost per tonne of establishing conformity and the fraction of the batch that passes.
“The first experiment is a certification pilot and it is embarrassingly cheap. Take a defined batch of structural steel sections from one demolition.”
Civilization Resilience Planning
Whether a collapsed industrial society could bootstrap back with the energy return, metallurgical sequence and population it would actually have.
A rigorous energy-and-materials analysis of industrial bootstrap pathways under a salvage endowment, closing the energy return on accessible surface coal and shallow oil under pre-industrial extraction, the metallurgy sequence, and the minimum population and knowledge base.
“The first required experiment is a paper, and this brief can specify its contents precisely enough that somebody could go and write it. A rigorous energy-and-materials analysis of industrial bootstrap pathways under a salvage endowment would have to contain three things.”
The Physical Stack of AI
Which layer of the physical stack — packaging, machinery, interconnection, or the permit — actually binds the AI buildout, and for how long.
Whether the announced 2028 cohort of gigawatt campuses — five Stargate-class sites totalling roughly 8 GW with fourth-quarter-2028 targets, plus the Crane nuclear restart — reaches powered operation on schedule, with the pattern of slips localising the binding layer.
“The decisive test is already running. The announced 2028 cohort — five Stargate-class sites carrying roughly 8 GW of stated capacity against fourth-quarter-2028 targets, the Crane nuclear restart, and the giant load queues behind them — will either reach powered operation on schedule or slip, and the pattern of the slips will localise the binding layer of the stack in a way no forecast can.”
Memory-Safe Computing Transition
Can automated translation convert large legacy C codebases into maintainable, semantics-preserving Rust at a cost that makes the memory-unsafe legacy tail tractable?
The independent TRACTOR evaluation: MIT Lincoln Laboratory scores automated C-to-Rust translation against escalating benchmark batteries every six months; success on realistic million-line codebases would turn billions of unsafe lines into a finite engineering bill, while a plateau at transliterated output on simplified C would leave the transition a new-code-only story.
“The decisive test is already running: the TRACTOR evaluation of whether automated translation can turn real C codebases into Rust that maintainers would accept without a full manual rewrite. MIT Lincoln Laboratory scores performer output against benchmark batteries released every six months, and a Round 1 evaluation report has been published.”
Post-Quantum Migration
Do the structured-lattice assumptions beneath ML-KEM and ML-DSA withstand sustained cryptanalysis while the migration that bets almost everything on them completes?
A practical attack on the module-lattice problems underlying ML-KEM and ML-DSA, of the kind that broke SIKE in 2022; the brief treats the worldwide cryptanalytic effort against these assumptions as its decisive experiment, already running, with even sustained security-estimate erosion short of a break forcing parameter escalation.
“The single result most able to change this assessment is a practical attack on the module-lattice problems beneath ML-KEM and ML-DSA. The 2022 break of SIKE showed what such an event looks like: a scheme that had survived years of public review fell to a classical attack running in about an hour on a single core.”
Quantum Networks
Whether a quantum repeater can outperform direct photon transmission over deployed fibre, turning the quantum internet from architecture into demonstrated infrastructure.
A memory-enhanced repeater link over deployed fibre, between independently operated nodes in different buildings, delivering entangled pairs at a higher rate than direct transmission through the same fibre allows.
“The decisive demonstration is a memory-enhanced repeater link over deployed fibre, between independently operated nodes in different buildings, that delivers entangled pairs at a higher rate than direct transmission through the same fibre would allow. Nobody has done it; every credible roadmap treats it as the gate through which the field must pass.”
Extreme-Environment Test Infrastructure
Does IFMIF-DONES come online and produce the first fusion-spectrum irradiation data, converting fusion's materials case from extrapolation into measurement?
First beam on target at IFMIF-DONES followed by the first published fusion-spectrum irradiation data for a reduced-activation steel, which the brief treats as the single result most able to change its assessment.
“The decisive result this brief tracks is first beam on target at IFMIF-DONES followed by the first published fusion-spectrum irradiation data for a reduced-activation steel. Every fusion first-wall schedule on the map extrapolates across a spectral gap that only this facility is funded to close.”
Additive Manufacturing Qualification
Whether in-situ monitoring can demonstrate regulator-grade probability of detection for critical defects across machines, making certification of fracture-critical additive parts repeatable rather than bespoke.
A blind, multi-site round robin printing nominally identical fracture-critical builds on at least five machines, with each site's in-situ monitoring verdicts locked in escrow before inspection, then computed tomography and fatigue testing to failure on every part, published as probability-of-detection curves against defect size.
“The decisive test is a blind, multi-site round robin: nominally identical fracture-critical builds on at least five machines, with in-situ monitoring verdicts locked before inspection, followed by computed tomography and fatigue testing to failure on every part, published as probability-of-detection curves against defect size.”
Critical Dependency Atlas
Does an audited cross-sector dependency inventory predict the failure set a real multi-sector event actually realises?
Build an audited power-water-gas-communications dependency inventory for one mid-sized region, register in advance the failure set it predicts for a defined hazard, and let the region's next severe event adjudicate the map.
“The decisive test is an audited cross-sector dependency inventory for one mid-sized region, scored against the next real event.”
Digital Chokepoints
Whether the redundancy built into the world's submarine cable network protects anything, which turns entirely on how long faults actually take to repair.
Publication by the maintenance-zone consortia of the time-to-restore distribution for submarine cable faults, broken down by cause, sea area and permitting regime.
“Every claim made about cable resilience is a claim about that distribution: whether a second cable helps depends on how long the first one stays down, and whether a repair fleet is adequate depends on the tail rather than the mean.”
Next-Generation Networks
Can the candidate 6G upper mid-band be covered from the tower grid that already exists, or does it require a second national build-out?
An independent, co-sited field comparison of uplink coverage at 7 GHz against 3.5 GHz on the same towers, using commercial handset transmit power and a published, repeatable methodology.
“The decisive measurement is an independent, co-sited comparison of uplink coverage at 7 GHz against 3.5 GHz on the same towers, with commercial handset transmit power and a published methodology.”
Manufacturing Digital Twins
Does installing a coupled model of a production system change operational outcomes, once the twin is separated from the sensors, processes and management attention that arrive with it?
A stepped-wedge deployment across several comparable lines or sites: randomise the order in which the twin is switched on, register the primary outcome before the first switch -- unplanned downtime hours, first-pass yield, energy per unit -- and publish the result whichever way it falls.
“The decisive experiment is a stepped-wedge deployment with pre-registered outcomes, and no operator has published one.”
Zero-Carbon Aviation
Can the largest single term in aviation's radiative forcing, contrail cirrus, be reduced at network scale by routing today's aircraft differently, and is the extra fuel burnt worth the forcing avoided?
A full-year trial across one airline's entire operation in which eligible flights are randomised into contrail-avoidance routing, contrail formation is verified from satellite rather than from the forecast that generated the decision, and the extra fuel burnt is accounted against the forcing avoided.
“The decisive experiment is a network-scale, independently verified contrail avoidance trial, and nobody has funded one.”
Advanced Air Mobility
Whether a certificated powered-lift fleet in scheduled service achieves the dispatch reliability, operating cost, vertiport throughput and community noise outcome the field has projected.
One full year of scheduled revenue operations by a type-certificated powered-lift fleet, publishing dispatch reliability against schedule, direct operating cost per block hour with pack amortisation separated, achieved movements per hour at the busiest vertiport, and noise complaints per hundred movements.
“Every contested claim in the field resolves against that dataset and none of them resolves without it.”
VIII — Governance & Institutions
Future Democracies
Whether the recommendations of citizens' assemblies survive into legislation, and at what rate.
Independently track every recommendation a deliberative body makes to its legislative fate, done for ten assemblies across five countries under a published and contested coding scheme.
“The decisive experiment is cheap, obvious and almost never run: independently track every recommendation from a body to its legislative fate.”
AI-Assisted Governance
Whether AI assistance raises administrative output, rather than the self-reported minutes the field currently measures.
Instrument a caseload rather than a diary: cases cleared per week, decisions reversed on appeal, backlog age and error rate at quality-control sample.
“Measure output, not minutes. Instrument a caseload rather than a diary: cases cleared per week, decisions reversed on appeal, backlog age, error rate at quality-control sample. This is the study that does not exist, and constructing it requires no new method — every one of those numbers is already produced by the administrations in question for other purposes.”
Digital Constitutional Systems
Whether a statute encoded as machine-readable rules is reproducible across independent competent coding teams.
An inter-coder replication: run independent teams across statutes of different drafting styles and measure the residual divergence in the encodings they produce.
“The cheapest decisive experiment is the inter-coder replication, and it would settle a constitutional argument for the price of a research assistant. Witt and colleagues showed divergent interpretive choices surviving an agreed vocabulary over two weeks on one statute. Run the same design with independent teams across statutes of different drafting styles and measure the residual divergence.”
Institutional Design
How many of the eleven commons design principles survive once measured coder disagreement is taken out of the headline result.
Re-estimate the commons corpus test under the measured coder-disagreement distribution and report how many of the eleven principles survive.
“The cheapest decisive experiment in the subject has been specified by the people best placed to run it and has not been run. Re-estimate the commons corpus test under the measured coder-disagreement distribution and report how many of the eleven principles survive.”
Scientific Governance Models
Whether partial lottery allocation of research funding produces different outcomes from panel ranking.
Analyse the partial-randomisation arms already executing at funders since 2013, where assignment above the quality threshold is random by construction, against the outcome data accumulating in the funders' own reporting systems.
“The highest-value experiment in this subject has already been run and simply not analysed. At every funder using partial randomisation, assignment above the quality threshold is random by construction.”
Future Legal Systems
Whether algorithmic risk assessment in courts operates as decision support or as the decision.
Publish judicial override rates and outcomes by defendant group for deployed risk-assessment tools.
“The cheapest decisive experiment is to publish override rates and outcomes by defendant group. Two of the field studies underpinning this brief reconstructed judicial compliance indirectly from administrative data — that is how the 90%-modelled against 29%-actual figure for immediate non-financial release was obtained. No jurisdiction publishes it directly.”
Public Policy Foresight
Whether published national risk assessments were right, which nobody has ever scored.
Retrospective scoring of the eighteen years of ranked, banded, dated assessments already published in the UK national risk register.
“The cheapest and most valuable experiment in this entire subject is retrospective scoring of documents that already exist. The UK national risk register has been published since 2008 with ranked, banded, dated assessments.”
Long-Term Institutions
Whether future-generations bodies have any measurable effect on legislation.
Publish the counts: how many legislative opinions a future-generations body issued, how many constitutional referrals it proposed, how many were made and how many succeeded.
“The most valuable single measurement is also the cheapest: publish the counts. Hungary's ombudsman office could state, in a paragraph, how many legislative opinions it issued, how many Constitutional Court referrals it proposed, how many were made, and how many succeeded. Every future-generations body in the world could do the same.”
Future Federalism
Whether reforming a fiscal equalisation formula changes subnational behaviour and outcomes.
Pre-register the evaluation of an equalisation reform before the formula changes, using the announced Canadian, Australian and German reform dates.
“The cheapest high-value experiment is to pre-register the evaluation of an equalisation reform before the formula changes. The dates are announced years ahead and the outcome data exist: Canada's programme renews 31 March 2029 , Australia's Productivity Commission reports finally on 31 December 2026 , Germany reformed in 2020 with parameters published in advance.”
Global Cooperation Models
Whether the Montreal Protocol caused the ozone outcome, or coincided with substitution that commercial incentives would have driven anyway.
A selection-corrected re-analysis of the Montreal Protocol itself, using staggered ratification dates and an outcome series measured by agencies outside the regime.
“The experiment this field most needs is a selection-corrected re-analysis of the Montreal Protocol itself — the one case with clean independent outcome data and near-universal membership. Nobody has separated the treaty's effect from the commercial incentives that would have driven substitution anyway, and until somebody does, the field's best evidence is also its least examined.”
Existential Risk Governance
Whether existential-risk regimes prevent the outcomes they target or merely select the states that were never going to produce them.
A prevention institution with a control group: exploit the staggered adoption of the Additional Protocol across states, or the staggered establishment of national AI evaluation bodies, as natural experiments with real variation in timing.
“Sixth, and the only design that would settle anything: a prevention institution with a control group. Nothing in this brief has one. The nearest available candidates are staggered adoption of the Additional Protocol across states, and the staggered establishment of national AI evaluation bodies across countries with otherwise comparable technology sectors.”
Technocracy and Democracy
Which of five mutually inconsistent results on central bank independence and inflation is an artefact of specification.
A pre-registered re-analysis of the five conflicting independence-inflation results on a common sample, varying the index choice, the sample split, the estimator and the treatment of initial inflation systematically.
“A pre-registered re-analysis on a common sample, with the index choice, the sample split, the estimator and the treatment of initial inflation varied systematically, would settle which of the five is an artefact of specification. The data are public. Nobody has done it.”
Future Civil Services
Whether AI tools raise civil-service throughput, measured as cases cleared rather than minutes reported.
Add one more arm to the existing departmental evaluations: random assignment of licences within a single processing function, with cases cleared per week as the dependent variable.
“The cheapest decisive experiment has already been designed twice and simply needs an output measure attached. Three United Kingdom departments evaluated the same product in the same quarter with three designs.”
Digital Citizenship
Whether the measured leakage reduction belongs to biometric authentication or to the removal of the payment intermediary.
A four-arm trial on one welfare programme in one state, randomising the channel change and the authentication layer separately.
“The highest-value experiment is the decomposition the two Indian trials came within one design choice of running. Randomise the channel change and the authentication layer separately — four arms: unchanged channel with and without biometric authentication, disintermediated channel with and without — on one welfare programme in one state.”
Civic Technology
What signature threshold a civic participation platform should set, and what that threshold does to proposals enacted and to later participation.
Randomise the signature threshold across comparable municipalities and measure proposals cleared, proposals enacted and subsequent participation.
“The obvious experiment is to randomise the threshold. These platforms are software and the signature requirement is a configuration value. Varying it across comparable municipalities and measuring proposals cleared, proposals enacted and subsequent participation would answer the field's central design question within a single budget cycle. It has never been done, and there is no technical or ethical barrier to doing it.”
Scientific Advisory Institutions
Whether formal scientific assessments change the decisions they are commissioned to inform.
Vary the advice and hold the decision-maker constant: a matched comparison of technical decisions taken with and without a formal assessment, controlling for salience and contestedness.
“The experiment nobody has run is still the obvious one: vary the advice, hold the decision-maker constant. Legislatures and executives make hundreds of technical decisions a year and commission formal assessments for a small, non-random subset.”
Distributed Governance
Whether token-governed organisations satisfy the commons design principles, and whether satisfying them predicts survival.
Apply the eleven commons design principles as a scored checklist to token-governed organisations and test whether configural satisfaction predicts survival, as it does for irrigation systems and fisheries.
“An experiment that would settle the central claim and has not been run: apply the eleven commons design principles as a scored checklist to token-governed organisations and test whether configural satisfaction predicts survival, as it does for irrigation systems and fisheries. The coding instrument exists and has been validated on 69 cases; nobody has pointed it at this domain.”
Future Public Administration
Whether making evaluation a condition of funding produces the outcome evidence public administration currently lacks.
Make an adequate evaluation plan a gate on project funding rather than a statistic government reports about itself.
“The cheapest high-value intervention is to make evaluation a condition of funding rather than a nice-to-have. The measurement already exists: 34% of the portfolio has an adequate plan, 66% does not, and 55% of the shortfall group produced no evidence of a plan at all.”
Civilizational Planning
Whether a statutory future-generations duty changes measured outcomes against a comparable jurisdiction without one.
A difference-in-differences comparison of Wales, which legislated a statutory future-generations duty in 2015, against the rest of the United Kingdom, which did not.
“The single highest-value experiment in this subject is already set up and nobody has run it. Wales legislated a statutory future-generations duty in 2015 and the rest of the United Kingdom did not. That is a natural experiment with a control group and roughly eleven years of indicator data collected on a comparable statistical basis. At the institutional level it remains unrun.”
Content Authenticity Infrastructure
Whether legal compulsion can raise the share of AI-generated media that reaches users carrying an intact machine-readable mark above the pre-mandate 35 to 45 percent platform baseline.
The three-jurisdiction natural experiment formed by China's labeling measures (September 2025), EU AI Act Article 50 (August 2026) and California SB 942 (August 2026): platform transparency reports by the end of 2027 show either the labelled share of AI-generated media rising decisively above the 35 to 45 percent baseline, or staying flat while stripping and open-weight generation absorb the mandate.
“The decisive test is already running: three mandatory-marking regimes came into force within twelve months of each other, and by the end of 2027 platform transparency reports will show whether the labelled share of AI-generated media rises decisively above the 35 to 45 percent baseline or stays flat while stripping and open-weight generation absorb the mandate.”
AI-Biology Governance
Whether the synthesis screening layer actually stops live orders, as opposed to scoring well on a curated benchmark.
A standing, blinded order-placement audit of the whole synthesis provider population: benign test orders submitted under varied customer identities at unannounced intervals, scored on whether each order was stopped, queried or shipped, with per-provider results published and an adjudication step separating the explanations the June 2025 twelve-order exercise could not distinguish.
“The decisive experiment is a standing, blinded order-placement audit of the whole provider population, with per-provider results published.”
Military AI and Strategic Stability
Whether the Seventh CCW Review Conference in November 2026 adopts a negotiating mandate for a binding autonomous-weapons instrument, settling if consensus arms control can still bind a militarily significant technology before mass adoption.
The November 2026 CCW Review Conference decision: with a negotiating majority above 70 states and a 156-vote General Assembly resolution behind it, either it adopts a mandate to negotiate a legally binding instrument on autonomous weapons, or it lets a decade of preparatory work lapse into another mandate cycle; the brief treats either outcome as settling the live institutional question.
“The decisive near-term result is institutional rather than technical: whether the Seventh CCW Review Conference in November 2026 adopts a negotiating mandate for a legally binding instrument on autonomous weapons, or lets a decade of preparatory work lapse into another mandate cycle.”
Information Integrity
Whether any single information-integrity intervention produces a durable downstream civic effect rather than an immediate change in sharing behaviour.
One randomised treatment assigned at the account level, with exposure, sharing, belief and a pre-registered downstream civic behaviour measured on the same subjects and reported at one week, one month and six months.
“The decisive experiment is a platform-scale field trial that carries one intervention all the way through the chain and reports durability at six months.”
Standards as Governance
Can delegated private standardisation supply the operative content of frontier-technology regulation on a legislated schedule?
Whether European harmonised standards for high-risk artificial intelligence are delivered and cited in the Official Journal in time for the December 2027 obligations, or the deadline is moved a second time.
“No new institution is required to run this experiment and no one can stop it.”
Research Security
Do national research-security regimes reduce meaningful risk, and at what cost in collaboration, delay and exclusion?
A difference-in-differences analysis of the staggered introduction of field-specific research-security lists, comparing in-scope and out-of-scope research areas before and after each date of effect on grant application volumes, partner composition, international co-authorship and time-to-award, with cross-country variation separating policy effect from global trend.
“The data already exist in funder administrative records and in bibliometric databases, no new instrument is required, and the cost is an analyst-year.”
Science Diplomacy
Does scientific cooperation survive strategic rivalry because of shared scientific values, or because separation is physically and financially expensive?
Assemble and analyse the three outcome series produced by the single 2022 shock across three flagship arrangements - component delivery schedules, co-authorship by affiliation, user counts and beam-time allocations, crew rotation and reboost manoeuvres, before and after - to test whether continuity tracks the cost of separation rather than the strength of scientific ties.
“The prediction from this brief is that continuity tracks the cost of separation and not the strength of scientific ties”
Civilizational Archives
Can a sealed long-term deposit actually be recovered and understood by someone outside the institution that made it?
A blind read-back trial: give a sealed long-term deposit made years ago - an Arctic film reel, a micro-etched disc, an offline tape set - to an independent team with no access to the depositing institution and no documentation beyond what physically accompanies the artefact, and measure bit recovery, meaning recovery and elapsed time.
“Everything this field believes about bootstrap usefulness is an untested assumption until somebody runs that trial and publishes the recovery rate.”
Crisis Science Institutions
Has any post-2020 reform actually shortened the time from identifying a novel pathogen to randomising the first patient under a pre-authorised protocol?
Publish the interval from pathogen identification to first patient randomised under a pre-authorised protocol, for every country claiming the capability, and score every preparedness exercise on it between now and the next outbreak.
“The decisive test is activation time, and it is measurable to the day.”
IX — Economics & Society
Post Scarcity Economics
Are falling delivered prices driven by learning curves that can keep falling, or by non-curve costs that are not falling at all?
Decompose delivered prices into experience-curve and non-experience-curve components across several jurisdictions, decades and goods, using data that already exists.
“The highest-value experiment is also the cheapest: decompose delivered prices into experience-curve and non-experience-curve components across jurisdictions, using data that already exists. The Berkeley and Brattle work does this for US electricity for one period. Doing it for several countries, several decades and several goods would convert the central dispute from an argument about anecdotes into a measured ratio.”
Universal Basic Abundance
Does a universal transfer move rents, wages and prices once it is paid to an entire labour and housing market rather than to scattered households?
Run a saturation design: randomise the transfer at the level of a labour and housing market rather than the household, and observe rents, wages, prices, firm entry and labour demand, which is the direct test of the inflation objection.
“The single most valuable experiment is a saturation design: randomise at the level of a labour and housing market, not at the level of a household.”
Future Labour Markets
Do firm, industry and regional estimates of automation's employment effect agree once a single design produces all three?
Estimate firm, industry and regional effects from one design on linked employer-employee registers, instead of the two-level-at-a-time national estimates that currently disagree.
“The highest-value experiment is also the cheapest: report firm, industry and regional effects from one design. France and the Netherlands each estimate two of the three levels and find them inconsistent; nobody estimates all three together.”
Human Flourishing
Do cross-country wellbeing comparisons measure different lives, or different ways of mapping the same life onto a number?
Measure the reporting function directly: present the same described life to respondents across countries, languages and income levels, anchor with vignettes at scale, and estimate each person's mapping from described state to reported integer.
“The highest-value experiment is also the cheapest, and it is not a survey. Measure the reporting function directly: present the same described life to respondents across countries, languages and income levels, anchor with vignettes at scale, and estimate the mapping from described state to reported integer per person.”
Wealth Distribution Systems
Are predistributional instruments the high-leverage lever on wealth inequality, given that most of the gap is pre-tax and almost none of it has been measured?
Build an evaluation programme for predistributional instruments, minimum wage schedules, sectoral bargaining coverage, licensing, zoning, education financing and healthcare financing, giving them the elasticity literature that wealth taxation already has.
“Fourth, and the one that would matter most: build an evaluation programme for predistributional instruments. If two-thirds to 90% of the gap is pre-tax, then minimum wage schedules, sectoral bargaining coverage, licensing, zoning, education financing and healthcare financing are the high-leverage instruments, and none has anything resembling the elasticity literature that wealth taxation has.”
Future Capital Markets
Does private capital actually beat public markets once every fund reports a public market equivalent on a common fee convention?
Require every fund marketing to institutional or retail investors to report a public market equivalent against a named benchmark on a standardised fee treatment, alongside the IRR.
“The highest-value experiment in this subject is a disclosure rule, and it is cheap. Require every fund marketing to institutional or retail investors to report a public market equivalent against a named benchmark on a standardised fee treatment, alongside the IRR. The data already exist inside the funds.”
Innovation Ecosystems
Do cluster interventions cause the regional outcomes they claim, or do they select places that were going to do well anyway?
Randomise or quasi-randomise a cluster intervention by running a lottery among the eligible, which every oversubscribed cluster programme is administratively able to do, and pre-register it.
“The highest-value experiment is the one the field's own surveys have been asking for and nobody has run: randomise or quasi-randomise a cluster intervention. Eligible regions exceed available funding in every cluster programme in existence, which means a lottery among the eligible is administratively available and ethically defensible.”
Digital Economies
Does the gap between individual and collective valuation of network goods recur beyond one platform, and must market-definition doctrine therefore carry a correction factor?
Replicate the collective-versus-individual valuation design on other network goods, messaging, marketplaces and operating systems, and see whether the substitution gap recurs.
“The highest-value experiment is the one already invented and run once: replicate the collective-versus-individual valuation design on other network goods. If the gap between individual and collective substitution recurs for messaging, marketplaces and operating systems, the standard market-definition instrument has a measured bias and competition authorities need a correction factor.”
AI Driven Productivity
Do measured AI gains on intermediate outputs survive through to shipped, sold or delivered units outside software?
Replicate the intermediate-to-final-output decomposition outside software: take a setting with a measured AI gain on an intermediate output and follow it through to a shipped, sold or delivered unit.
“The highest-value experiment is the cheapest and has been specified for two years: replicate the intermediate-to-final-output decomposition outside software. Take any setting with a measured AI gain on an intermediate output and follow it to a shipped, sold or delivered unit.”
Resource Economies
Do resource funds receive what their own inflow rules require, and is the Nigerian shortfall the tail of the distribution or its median?
Audit fund inflow rules against realised inflows across every country with a resource fund, using the published rules and the published deposits.
“The cheapest high-value study in this subject is an audit nobody has run: fund inflow rules against realised inflows, across every country with a resource fund. The rules are published, the deposits are published, and the difference is arithmetic.”
Future Trade Systems
Has trade genuinely diversified away from China since 2018, or has a single alternative supplier taken most of the share China lost?
Recompute the supplier-concentration result on published trade data for 2023 to 2026, testing whether a single supplier still takes at least three quarters of the lost market share for most affected products.
“The cheapest high-value study is a repeat of an existing calculation on newer data: recompute the supplier-concentration result for 2023–26. The original finding — that for over 70% of affected products a single supplier took at least three quarters of China's lost market share — is the sharpest available refutation of the de-risking story, and it is a straightforward calculation on published trade data.”
Circular Economies
Does eco-modulated producer responsibility change design decisions, or does the instrument fail at any fee level rather than only at the level it has been set at?
Apply the estimator that produced the null on flat fees to the dated introduction of modulated fees across jurisdictions, using the same panel of countries and materials.
“The highest-value study available is an evaluation of an eco-modulated scheme using the design that produced the null on flat fees. The original panel used temporal variation in the fees actually charged across 25 countries and four materials; several jurisdictions have since introduced modulated fees on dated schedules.”
Future Taxation Models
Why did the global minimum tax raise roughly a third of its forecast: substance-based exclusions, transition rules, or adoption gaps?
Publish five years of global minimum tax outturn by jurisdiction against jurisdiction-level projection, which separates the substance-based exclusion, the transition rules and adoption gaps as explanations.
“The highest-value experiment is publication rather than policy: five years of global minimum tax outturn by jurisdiction. The first year came in at roughly a third of forecast and the only public figures reach this brief through an interested secondary source.”
Scientific Funding Models
Do lottery allocation, milestone termination and people-not-projects funding outperform conventional peer review, on comparisons funders already hold?
Report the quasi-random allocation comparisons already sitting in funders' records, wherever part of a portfolio was assigned by a rule that allocates quasi-randomly above a quality threshold.
“The cheapest and highest-value experiment in this subject has already been run and merely needs reporting. Wherever a funder has allocated part of a portfolio by a rule that assigns quasi-randomly above a quality threshold, the comparison exists in its records.”
Civilization Scale Investment
Will investors buy an ultra-long, consumption-linked sovereign instrument at size, or is the long-tenor market a theory with no bid behind it?
Issue an ultra-long sovereign instrument whose coupon is linked to consumption or output rather than to a nominal rate, and publish the order book: the tenor hypothesis predicts placement failure at size among solvency-constrained institutions, the design literature predicts the opposite.
“The decisive experiment is an issuance, and it is available to any large sovereign that wants it. Issue an ultra-long instrument whose coupon is linked to consumption or output rather than to a nominal rate, and publish the order book.”
Economic Resilience
What does a strategic reserve release actually do to prices, measured by a method fixed before the release rather than chosen afterwards?
Pre-register the evaluation of the next strategic reserve release before it happens, agreeing comparator markets, price windows and model specification in advance.
“The cheapest high-value experiment in this brief costs nothing and has never been run: pre-register the evaluation of the next strategic reserve release before it happens. Two published estimates of the 2022 release differ by a factor of three and the wider one comes from the party that authorised it.”
Reputation Economies
Does reputation survive being transferable, if every transfer is labelled and the chain of custody is visible?
Build a market in provenance-labelled reputation: permit transfer, record every transfer publicly, display the chain of custody, and measure whether a labelled name still commands a premium, how fast that premium decays across transfers, and whether pooling appears anyway.
“The experiment that would move this subject most is the one nobody has run: build a market in provenance-labelled reputation. Permit transfer, record every transfer publicly, display the chain of custody, and measure whether a labelled name still commands a premium, how fast the premium decays across transfers, and whether pooling appears anyway.”
Human Development Metrics
Do development rankings survive a change in the aggregation rule, or do they carry an unstated dependence on choices their users never see?
Publish a parallel HDI under alternative aggregation, non-compensatory, with different goalposts and no income cap, and report how far the rankings move.
“The cheapest high-value study is a recomputation, not a survey. Publish a parallel HDI under alternative aggregation — non-compensatory, different goalposts, no income cap — and report how far rankings move.”
Future Philanthropy
Do dormant donor-advised fund accounts eventually grant, and after how long?
Build an account-level panel of donor-advised fund behaviour with dormancy tracked to resolution, from records sponsors already hold.
“The most valuable study available is the one the sector's own data would support tomorrow: an account-level panel of donor-advised fund behaviour with dormancy tracked to resolution. The existing study covers 2014–2022 and shows 22% of accounts granting nothing in a three-year window.”
Critical Mineral Midstream
Whether policy-built refining and separation capacity outside China durably reduces measured midstream concentration, or reverts once Chinese prices collapse.
The running natural experiment in rare earths: whether the top-supplier refining share, 90% in 2023 and 85% in 2025, continues toward the roughly 70% the IEA projects for 2035 under the US price floor, Lynas heavy separation and USD 65 billion of public finance, or reverts in the first Chinese price collapse.
“The decisive test is already running: whether the rare-earth top-supplier refining share, 90% in 2023 and 85% in 2025, continues down to the roughly 70% the IEA projects for 2035, or reverts the first time Chinese prices collapse and the subsidised plants must sell into them.”
Compute Concentration
Does concentrated access to frontier compute convert into concentrated capability, or does capped access keep converging on the frontier?
The export-control natural experiment already under way: comparable, independently run evaluations of models trained under the compute ceiling against uncapped models released in the same window, with training-compute estimates attached, through 2026 and 2027.
“The decisive result is the one already running: whether models trained under the export-control compute ceiling stay within about a year of the frontier through 2026 and 2027.”
Technological Sovereignty
Whether a published, dated technological-sovereignty target functions as a binding instrument or as decoration.
The 2030 measurement of the European Chips Act's 20%-of-global-production-value target, and whether missing it forces a public re-scoping of the programme or is absorbed without consequence.
“The decisive test in this subject is already scheduled and nobody scheduled it as a test: the European Chips Act’s 20% of global production value by 2030. It is the only technological-sovereignty target anywhere that is numeric, dated, externally measurable and already scored by an independent audit body that has projected a miss.”
Material Passports
Does mandating a machine-readable material record change what actually happens to the material?
A pre-registered before-and-after study of the EU battery passport obligation taking effect on 18 February 2027, fixing recovery baselines per Member State and per chemistry beforehand, with recovered mass and recovered value per tonne as endpoints and non-EU recyclers of comparable chemistries as the comparison arm.
“The decisive test is already scheduled and nobody has designed it: the battery passport becomes obligatory on 18 February 2027, which creates a dated before-and-after on a defined product class whose recovery statistics are already collected.”
Climate Adaptation Finance
Does money spent on adaptation buy a verified physical change, and does that change reduce measured loss?
Publish one audited, matched-sample series of claim frequency, claim severity and premium for wind-retrofit-designated homes against comparable undesignated homes across several Gulf storm seasons, alongside the grant cost, in a form an outside analyst could re-run.
“The decisive measurement is a matched publication of claim frequency, claim severity and premium for designated versus comparable undesignated homes across several storm seasons.”
Aging Societies
What does a large, rapid immigration expansion and its equally rapid reversal actually do to wages, rents, output per head, the age structure and the sectors that staffed themselves from the inflow?
Evaluate Canada's 2021-2024 immigration expansion and its 2024-2027 contraction as a two-directional natural experiment, using the quarterly demographic, wage, housing and fiscal data and the administrative records that already exist.
“The decisive test is a natural experiment already running in Canada, and it runs in both directions.”
X — Historical & Frontier Science Studies
History of Electrogravitics
Whether the founding citation of the electrogravitics suppression literature refers to an article that exists.
Requesting Interavia vol. 11 no. 5 (1956), pp. 373-374 through interlibrary loan: either the seventy-year-old citation is verified, or the founding citation of the suppression literature is formally recorded as untraceable.
“Two: request Interavia vol. 11 no. 5 (1956), pp. 373–374 through interlibrary loan. Outcome: the article is produced, in which case a seventy-year-old citation is finally verified; or it is not, in which case the founding citation of the suppression literature is formally recorded as untraceable.”
Cold War Science Programs
Whether present-day secrecy orders suppress innovation at the rate measured for the wartime cohort.
Running Gross's secrecy-order design forward on the 6,543 orders now in force, comparing released filings against matched unreleased ones, to price the current instrument rather than the 1940s one.
“The cheapest high-value study is to run Gross's design forward on the standing secrecy orders. The wartime cohort was measurable because the orders were eventually released and the technology classes are on the record.”
Operation Paperclip
What the imported German scientists actually added to American technical capability, expressed as a number rather than a narrative.
Estimate the counterfactual: use the IWG record of names, arrival dates and assignments against continuous patent records and constructible comparison groups of American engineers in the same classes.
“The single highest-value study in this subject is the one nobody has run: estimate the counterfactual. The ingredients exist. The IWG record supplies names, arrival dates and assignments.”
Historical Space Colonization Concepts
Whether a sealed ecological life-support system can hold its atmosphere over years once the hidden chemical sinks that broke Biosphere 2 are accounted for.
A re-run of the Biosphere 2 closure with the chemistry instrumented, carrying a full material inventory of the structure itself: every sink, every surface, every curing reaction.
“The experiment that would matter most now is a re-run with the chemistry instrumented. The 1991 failure was invisible in the carbon dioxide record because the carbon went into the concrete.”
Soviet Frontier Science
What Lysenkoism actually cost Soviet agriculture, as a measured number rather than seventy years of adjectives.
A difference-in-differences on Soviet crop-yield series across crops and regions differentially exposed to mandated Lysenkoist agronomic practices, with the collectivisation famines separated in time from post-1948 policy.
“The most valuable single study is the one the Lysenko literature has declined to attempt: estimate the agricultural cost. Soviet crop-yield series by region and crop exist; the timing of Lysenkoist agronomic mandates is documented; the collectivisation famines are separable in time from post-1948 policy.”
Scientific Revolutions
Whether taxonomic incommensurability shows up in the record, that is, whether successor lexicons cross-classify the kinds of the theory they replace.
Build the taxonomic test Kuhn's late position implies: take a discipline with documented terminological turnover, reconstruct what its kind terms picked out before and after, and check the no-overlap structure directly.
“If successor lexicons cross-classify the incumbent's kinds, the strongest surviving version of incommensurability has its first empirical support. If they nest cleanly, it does not. This is the one test in the subject that could come out either way on evidence, and nobody has built it.”
Historical Megaprojects
Whether cost overruns have really worsened over time, tested on the kinds of project the popular argument is actually about.
A comparable cost-outturn series for buildings, dams and canals spanning the same period as the 2002 transport sample, extending the F-test outside rail, roads and fixed links.
“One: extend the F-test outside transport. The 2002 sample is rail, roads and fixed links. The three examples that carry the popular argument are a building, a dam and a canal. A comparable outturn series for buildings, dams and canals spanning the same period would test the claim on the projects the claim is actually made about”
Frontier Aerospace Programs
Whether the collapse in experimental-aircraft cadence is explained by rising real cost per demonstrator airframe or by something else.
A deflated real cost-per-demonstrator-airframe series from the X-1 to the X-59, with programme duration and flight count alongside, built from appropriations and audit data that already exist.
“The cheapest decisive study is the unit-cost series. Deflated real cost per demonstrator airframe from the X-1 to the X-59, with programme duration and flight count alongside, using appropriations and audit data that already exist.”
Philosophy of Science
What practising scientists actually treat as the criterion for classifying a claim as outside science.
A stratified survey of practising researchers across disciplines, presenting the published criteria and measuring endorsement, disagreement between fields, and the gap between endorsement and application on worked cases.
“The highest-value missing study in this subject is a survey, and it is cheap. Ask a stratified sample of practising researchers across disciplines what would make them classify a claim as outside science”
Technology Forecasting
Whether the correlated error in technology forecasts is a property of one technology, one model family, or the whole practice.
A second, disinterested scoring of the 2,905 integrated-assessment-model projections of solar cost decline, which are dated, numerical, published and already resolved by outturn.
“The cheapest valuable experiment is to score the forecasts that already exist, and the highest-value corpus is sitting in the open. The 2,905 integrated-assessment-model projections of solar cost decline are dated, numerical, published, and resolved by an outturn nobody disputes.”
Innovation History
Whether the measured decline in disruptive science is real or an artefact of reference-list growth.
Recompute the CD disruption series across the same six datasets with the correction Petersen and colleagues named and quantified.
“Petersen and colleagues did not merely allege a bias in the disruption index; they named the mechanism and quantified it. Recomputing the CD series across the same six datasets with the correction applied would resolve the most-cited quantitative claim in the field one way or the other, and it requires no new data.”
Scientific Institutions Through History
Whether a rejection rate flat between 8.5% and 13.5% is a property of learned-society refereeing or an artefact of one institution.
Run the Royal Society manuscript analysis again on the editorial archives of a second learned society: a continental academy, a national society in a different tradition, or a long-lived commercial journal.
“The single highest-value study in this subject is a second manuscript series. The Royal Society records give rejection rate, referee count, participation share and workload distribution across a century.”
Grand Challenges of Humanity
Whether the leap in DARPA Grand Challenge performance from 2004 to 2005 was a prize framing effect or learning, consolidation and falling sensor prices.
Decompose DARPA 2004 to 2005 using team-level data on continuity, spending, personnel and component costs across the two years.
“Four: decompose DARPA 2004 to 2005. Team-level data on continuity, spending, personnel and component costs across the two years would separate the framing effect from learning, consolidation and falling sensor prices.”
Technology Trees of Civilization
Whether directed prerequisite structure, which technology must come before which, is recoverable at all from the data we have.
Take the normalised patent proximity network, add first-grant dates, and test whether any orientation rule recovers prerequisite relationships that domain experts endorse out of sample.
“Second, and this is the one that would settle the subject: try to orient the edges. Take the normalised patent proximity network, add first-grant dates, and test whether any orientation rule recovers prerequisite relationships that domain experts endorse out of sample.”
Civilization Timelines
Whether the named periods of world history survive a pre-registered comparative test, or are artefacts of the narrative that named them.
A pre-registered periodisation test on the named periods that have never had one, following the Axial Age template: state in advance what the comparative data would have to show for the period to survive, and publish the result either way.
“The cheapest valuable experiment is a pre-registered periodisation test, and it has been run roughly once. The Axial Age assessment is the template: take a named period, state in advance what the comparative data would have to show for it to survive, and publish the result either way.”
Science of Science
Does an AI research tool actually compress research stages in working laboratories without degrading output quality, when quality is scored by replication rather than citation?
Randomise access to an AI research tool at the laboratory level, log stage clocks across the pipeline, pre-register the quality metric, and score sampled outputs by independent replication or replication market two years on - the METR design scaled from code to bench.
“The decisive test is a stage-instrumented randomised rollout of an AI research tool across working laboratories, with quality endpoints scored by replication rather than by citation.”
Cloud Laboratories
Whether machine execution of a protocol actually makes an experiment reproducible, and if it does not, where the variance lives.
Implement one non-trivial protocol on two independent cloud laboratories, run it blinded on both with identical inputs, and compare the results against each other and against a skilled bench laboratory working from the written protocol.
“The decisive experiment is the cross-platform replication, and nobody has run it. Take one non-trivial protocol, implement it on two independent cloud laboratories, run it blinded on both with identical inputs, and compare the results against each other and against a skilled bench laboratory running the same protocol from its written form.”
Metrology Infrastructure
Do the uncertainties laboratories state for a traded but unstandardised measurand match the dispersion that appears when the same samples are measured independently?
A blinded interlaboratory round robin in durable carbon-removal quantification: identical homogeneity-tested samples to twenty or more laboratories, each reporting a value and its stated uncertainty under its normal procedure, with the reproducibility standard deviation published against those stated uncertainties.
“The decisive experiment is a blinded interlaboratory round robin in a market that already prices an unaudited number. Distribute identical, homogeneity-tested samples to twenty or more laboratories that currently quantify durable carbon removal, have each report a value and its stated uncertainty under its normal procedure without knowing the others, and publish the reproducibility standard deviation against the stated uncertainties.”
Research Front Detection
What sensitivity and false-alarm rate does any early-detection method actually achieve when it has to commit before the outcome is known?
A prospective register in which several teams publish, on the same frozen corpus and the same date, ranked lists of research fronts they predict will be consequential, with outcome measures and thresholds fixed in advance, held by an independent body and scored at five and ten years.
“The decisive experiment is a prospective register of dated emergence forecasts with resolution criteria fixed in advance. Several teams publish, on the same frozen corpus and the same date, a ranked list of research fronts they predict will be consequential, together with the outcome measures and thresholds by which they agree to be judged.”