The Institute's method is the workback plan: start from what is currently considered impossible and walk back to the first checkable step. This page collects those first checkable steps. For each brief it records the experiment, observation, demonstration, measurement, natural experiment or policy result the brief itself treats as decisive, when the brief says it could happen, and the flag the brief attaches to the claim it would test.

The headline is the horizon column. Of 295 decisive results, 43 are already running and 21 are tied to a stated near-term window — but 214 are unscheduled: tests the briefs say could be run, that nobody has scheduled or funded. The frontier's bottleneck, on this evidence, is less often that the decisive test is impossible than that no one has bought it.

Horizons
43 running · 21 near · 11 mid · 6 long · 214 unscheduled
Types
131 experiment · 83 measurement · 22 observation · 25 natural experiment · 24 demonstration · 10 policy result

I — Physics & Propulsion

Gravity Modification FR-I-01

Whether any laboratory configuration modifies gravity, or whether the running precision tests keep returning nulls.

Take a claimed gravity-modification effect and vary the parameter the theory says it scales with to see whether the signal follows, as Tajmar's 2011 rerun did at 50 times the rotation rate; generalised by the 2024 three-balance vacuum search across the whole hypothesis class. — experiment · running · Established

“The single most informative experiment anyone could run has already been run twice. Tajmar's 2011 rerun at 50 times the rotation rate is the model: take the claimed effect, vary the parameter the theory says it scales with, and see whether the signal follows. It did not.”

The brief says the experiments that matter are already running, each with a published sensitivity and each having returned a null.

Artificial Gravity FR-I-02

Whether intermittent or continuous partial gravity in orbit prevents the physiological deterioration of weightlessness in humans.

A human-rated centrifuge flown in orbit: a 2.5-metre instrument spanning 0.01 to 2 g was built and never flown, and a station-integrated ergometric version was cancelled on structural grounds. — experiment · unscheduled · Established

“Second: a human-rated centrifuge in orbit. This is the experiment that answers the question the field is actually asking, and it is the one that has been cancelled twice.”

The design work is done and the hardware has been manufactured once; the experiment has been cancelled twice.

Warp Drives FR-I-03

Whether any proposed warp metric can be sourced without violating the energy conditions.

Run an energy-condition solver over the field's own catalogue of metrics, computing the Einstein tensor and evaluating the conditions pointwise across observers rather than in a single frame; Warp Factory and, independently, warpax keep returning the same answer. — experiment · running · Established

“The first is numerical relativity applied to the metrics themselves, and it has already produced this brief's decisive result. Warp Factory takes an arbitrary metric, computes the Einstein tensor, and evaluates the energy conditions pointwise across observers rather than in a single frame. Running it on the field's own catalogue found that every metric tested violates conditions its authors had believed it satisfied.”

The brief says there is no propulsion experiment in this subject and there has never been one.

Wormholes FR-I-04

Whether negative-mass compact objects or traversable wormholes exist at any detectable abundance.

A microlensing search for the negative-mass lensing signature - total or partial eclipse of the source in the umbra region and a Shapiro time gain rather than a delay - run against a quasar-lens survey to set abundance upper limits. — observation · running · Frontier

“Microlensing, which is the mature line. Safonova, Torres and Romero's 2001 analysis gives the signature precisely: an effective negative-mass lens produces a total or partial eclipse of the source in the umbra region, and a Shapiro time gain where an ordinary lens produces a delay.”

The brief says the experiment that would settle the subject cannot be specified, because there is no formation mechanism to build an apparatus around.

Reactionless Propulsion FR-I-05

Whether any claimed reactionless thruster produces a thrust that survives the standard controls.

Apply the five controls, cheapest first: orientation reversal, a ten-thousand-fold power attenuator, magnetic shielding of the feed lines with a documented residual, thermal isolation of the feedthrough with the surface-tension term measured, and publication of the full systematic budget by a group with no stake. — experiment · running · Established

“The experiment that matters is already specified, because it is the one that has already killed every signal.”

The brief says a pre-registered protocol has never been done in this subject, and it is the only experimental change on offer that would convert a disputed null into an undisputed one.

Inertial Manipulation FR-I-06

Whether the claimed inertial-manipulation devices produce a real force or a measurement artefact.

Run the remaining cheap discriminators on the claimed devices: reverse the orientation and see whether the thrust reverses, and attenuate the drive power by four orders of magnitude and see whether the thrust falls. — experiment · unscheduled · Frontier

“What would actually be discriminating is cheap and has been specified by a sympathetic assessor. Reverse the device's orientation and see whether the thrust reverses; attenuate the drive power by four orders of magnitude and see whether the thrust falls.”

Running the remaining discriminators is a graduate-scale experiment and it has not been published; the brief adds that improving the equivalence-principle null is not the experiment that would change its conclusion.

Vacuum Energy Engineering FR-I-07

Whether the quantum inequalities bounding negative energy density survive a direct, purpose-built test against squeezed light.

A dedicated test of the Ford-Roman inequality against squeezed light: squeezing generated and characterised specifically to evaluate the inequality with a pre-registered sampling function, rather than a meta-analysis of other people's data with sampling functions chosen after the fact. — experiment · unscheduled · Speculative

“A purpose-built experiment — squeezing generated and characterised specifically to evaluate the inequality with a pre-registered sampling function — would either substantiate the most important claim in this cluster or dispose of it. Nobody appears to have run one, and the omission is strange given what turns on it.”

The brief says the experimental frontier here is about testing the constraint, not about beating it.

Quantum Gravity FR-I-08

Whether the gravitational field is quantum - whether gravity can mediate entanglement between two masses.

A gravitationally induced entanglement experiment: two masses in adjacent interferometers, each in spatial superposition, interacting only gravitationally with a conducting plate screening the Casimir-Polder interaction, read out through spin correlations, with independent replication. — experiment · unscheduled · Frontier

“Tier one, decisive-ish: a gravitationally induced entanglement signal. Two masses in adjacent interferometers, each in spatial superposition, interacting only gravitationally, with a conducting plate screening the Casimir–Polder interaction, read out through spin correlations, with independent replication.”

The brief gives no schedule for tier one and says only tier three is currently producing outcomes; a positive result would not select among the competing theory programmes.

Negative Mass FR-I-09

Whether antimatter falls upward - whether any substance has negative gravitational mass.

ALPHA-g's direct free-fall measurement on magnetically confined antihydrogen, with a laser-cooled successor aimed at 1% precision against a theorists' target of one part in ten million. — measurement · running · Established

“Direct measurement, and it is done. ALPHA-g released roughly 100 magnetically confined antihydrogen atoms over 20 seconds and recorded annihilation positions relative to the trap opening”

The brief says nobody can make a gram of negative-mass substance because no theory says how, so its experimental programme is a search rather than a construction.

Space-Time Metric Engineering FR-I-10

Whether any engineered metric perturbation can be produced or detected.

Numerical relativity on a collapsing null-energy-condition-violating spacetime, producing computed gravitational waveforms that can be searched for in existing detector data - a search rather than a construction. — experiment · running · Speculative

“This is the only experimental handle anyone has proposed, and it is a search rather than a construction — the authors themselves label the search application “rather speculative” and frame the value as understanding the stability of such spacetimes.”

The brief says no apparatus has been proposed that could produce a metric perturbation large enough for any instrument now conceivable to detect.

Advanced Nuclear Propulsion FR-I-11

Whether low-enriched HALEU fuel can match HEU performance in a nuclear thermal rocket.

A comprehensive assessment of HEU against HALEU for nuclear thermal propulsion, backed by irradiation data on candidate low-enriched fuel forms, as the National Academies formally recommended. — experiment · unscheduled · Established

“The single most consequential missing experiment is the fuel comparison the National Academies formally recommended. A comprehensive assessment of HEU against HALEU for NTP, backed by irradiation data on candidate low-enriched fuel forms.”

Until it exists, every performance number quoted for a modern engine is an extrapolation from a different fuel in a different regulatory tier.

Fusion Spacecraft FR-I-12

Whether any fusion device can reach the specific power a spacecraft propulsion system requires.

Take a fusion device that produces net energy, weigh it including magnets, cryoplant, shielding, power conversion and radiators, and report kilowatts per kilogram. — measurement · unscheduled · Frontier

“The experiment that would move this brief is a specific-power measurement, and nobody is running it. Take any fusion device that produces net energy, weigh it including magnets, cryoplant, shielding, power conversion and radiators, and report kilowatts per kilogram.”

No fusion programme reports that number because no fusion programme is graded on it.

Antimatter Propulsion FR-I-13

Whether a small number of antiprotons can catalyse a target burn releasing more energy than the antiprotons cost to make.

An antiproton-catalysed micro-fission or micro-fusion ignition at any scale, showing that a small number of antiprotons initiates a burn releasing more energy than they cost to produce. — experiment · unscheduled · Speculative

“The decisive concept-level experiment is an antiproton-catalysed micro-fission or micro-fusion ignition at any scale. Demonstrate that a small number of antiprotons initiates a target burn that releases more energy than the antiprotons cost to make, and the catalysed branch becomes a research programme rather than a design study.”

The brief says this experiment is not blocked by the production rate: the quantities involved in a single ignition are far below 140 nanograms.

Solar Sail Systems FR-I-14

Whether a full-scale solar sail can be deployed reliably enough to become an operational vehicle.

The test the ACS3 failure report asks for: high-fidelity full-scale deployment testing without gravity compensation, with full-visibility cameras, and off-nominal simulation done early rather than late. — experiment · unscheduled · Established

“The single most productive thing anyone could do next is the test that report asks for: high-fidelity full-scale deployment testing without gravity compensation, with full-visibility cameras, and off-nominal simulation done early rather than late.”

ACS3 flew, deployed, failed in a specific and instrumented way, and NASA published the failure with its causes and its lessons.

Beam-Powered Propulsion FR-I-15

Whether high-power laser cost per watt is actually falling fast enough to make beamed propulsion economics credible.

Publish a defensible time series of cost per watt for high-power fibre lasers over the last two decades, which would either restore the projected eighteen-month halving or retire it. — measurement · unscheduled · Frontier

“The experiment that would settle the most is not an experiment at all: publish a laser cost series. The concept's economics rest on a projected eighteen-month halving; the only contemporary spot price found for this brief is $100 per watt against a $0.01–0.05 requirement.”

The brief says such a series does not appear to exist in public.

Interstellar Probes FR-I-16

Whether an institution can fund, launch and then operate a spacecraft for fifty years.

Fly a precursor mission to several hundred astronomical units - an experiment on an institution rather than on hardware, since nothing about the physics of getting there is unknown. — experiment · long · Speculative

“The decisive mission-level experiment is a precursor itself. Nothing about the physics of reaching several hundred astronomical units is unknown; what is unknown is whether an institution can fund, launch and then operate a spacecraft for fifty years.”

The brief says it has never been performed: Voyager was not designed as a fifty-year mission and became one.

Alcubierre Metrics FR-I-17

Whether any Alcubierre-type metric satisfies the energy conditions when evaluated across all observers rather than in a single frame.

Run an all-observer energy-condition solver over the published back catalogue of warp metrics; it has been run twice by different people, with violations found in every metric tested. — experiment · running · Established

“The first computational programme is all-observer energy-condition evaluation, and it settled the 2020–21 episode. Warp Factory computes the Einstein tensor for an arbitrary metric and evaluates the energy conditions pointwise across observer classes rather than in a single frame. Running it across the family found violations in every metric tested, including one its own authors had published as satisfying the conditions.”

The brief says there is no experiment on this metric and there cannot be one at present, so computation has replaced experiment.

Gravitational Wave Engineering FR-I-18

Whether the operating gravitational-wave detectors can serve as anything other than passive astrophysical instruments.

Search the operating detection network's existing data for the computed gravitational-wave signatures of exotic spacetime collapse - using instruments that already exist to look for somebody else's metric engineering. — observation · unscheduled · Speculative

“The one genuinely novel experimental idea in the subject uses instruments that already exist and looks for somebody else's device. If exotic spacetime collapse produces computable gravitational-wave signatures, then the operating detection network is incidentally an instrument for finding artificial metric engineering.”

It costs nothing beyond analysis time on data already being taken, which makes it the only proposal here worth acting on.

Exotic Materials for Propulsion FR-I-19

Whether candidate nuclear-thermal fuel forms survive flowing hydrogen above 2700 K for a mission duration.

Hold uranium nitride kernels in a molybdenum-tungsten matrix, or in zirconium carbide, above 2700 K in flowing hydrogen for a mission duration, with carbon interaction with the kernel and thermal-expansion mismatch at the cladding instrumented rather than inferred. — experiment · unscheduled · Established

“The decisive experiment in this brief is a hot-hydrogen fuel test at the mission condition, and it has not been run. Uranium nitride kernels in a molybdenum–tungsten matrix, or in zirconium carbide, held above 2700 K in flowing hydrogen for a mission duration, with the two named failure mechanisms — carbon interaction with the kernel and thermal-expansion mismatch at the cladding — instrumented rather than inferred.”

Pewee's 2550 K is the number to beat and it was set on a test stand decades ago.

Electrodynamic Propulsion Concepts FR-I-20

Whether an electrodynamic tether can produce useful thrust in flight rather than only measured current.

Fly a tether that drives current against the motional electromotive force with onboard power in low Earth orbit, and measure the resulting orbit raising against the current and the field. — experiment · unscheduled · Established

“The decisive experiment for the underused member is a thrust-producing tether mission, and it has never been flown. Drive current against the motional electromotive force with onboard power, in low Earth orbit, and measure the resulting orbit raising against the current and the field.”

The physics was measured in 1996; the propulsive mode has been described since 1993 and never demonstrated in flight.

Electrogravitics FR-I-21

Whether high-voltage electrode and field geometries produce any force beyond ion wind.

Sweep the geometry rather than testing one famous device, go to vacuum and vary the pressure, instrument to well below the claimed effect, and test the theory families' distinctive predictions - the template the 2024 campaign completed at up to 40 kV on balances resolving below a nano-newton. — experiment · running · Established

“The experiment that settled the gravitic question has been run four times and the design is worth stating as a template. Sweep the geometry rather than testing one famous device; go to vacuum and vary the pressure; instrument to well below the claimed effect; and test the theory families' distinctive predictions rather than only the headline claim.”

The one experiment nobody has run is the one Talley flagged in 1991: pulsed high-voltage operation near breakdown in vacuum.

Mach Effect Thrusters FR-I-22

Whether the Mach-effect thruster produces thrust at its claimed operating point under a protocol both sides agreed in advance.

A seven-item joint test: a pre-registered drive configuration, proponent-supplied tuned hardware, voltage at the claimed operating point, vibration isolation with an independent accelerometer channel, a same-session symmetry null, orientation reversal and a power-attenuator control, and resolution sufficient for a claim the proponents must first name. — experiment · unscheduled · Established

“The outstanding experiment can be specified completely, from the dispute itself, which is why this brief is prescriptive rather than plaintive. Seven items. A pre-registered drive configuration agreed by both parties before the run: frequency, voltage amplitude, transformer, piezoelectric stack construction, reaction-mass geometry and settling protocol.”

The brief says the instrument requirement depends entirely on which claim is being tested, and the proponents need to name one figure before the experiment can be written down.

Planetary Scale Energy Systems FR-I-23

Whether coherent phased-array power beaming works at the element count and aperture a planetary-scale system assumes.

Coherent phase control across of order 100,000 amplifiers on a kilometre-class transmitting aperture, holding a beam onto a multi-kilometre rectenna. — experiment · unscheduled · Speculative

“The experiment that has not been run is the one at scale. Coherent phase control across of order 100,000 amplifiers on a kilometre-class transmitting aperture, holding a beam onto a multi-kilometre rectenna, is the untested step, and it is an array-engineering experiment rather than a power-conversion one.”

Nothing in the fetched record demonstrates phased-array coherence at anything approaching that element count.

Space-Based Manufacturing FR-I-24

Whether the seed-crystal effect generalises, and so whether orbital manufacturing is a production business or a seed business.

Test whether flight-grown polymorphs are retained through terrestrial generations across a wide panel of pharmaceutically relevant molecules, as two of five did through ten generations and a third through nine. — experiment · unscheduled · Frontier

“Second, and the highest-value experiment in the whole subject: does the seed-crystal effect generalise? Two of five molecules retained their flight polymorph through ten terrestrial generations and a third through nine.”

The brief says nothing else in it would move the economics as much.

Deep Space Infrastructure FR-I-25

Whether the deep space network can carry a crewed lunar programme and flagship science at the same time.

Artemis I's unintended capacity test of the network, which could not serve everything at once and measurably degraded flagship science; the scheduled follow-on is Artemis III flying two crewed vehicles in cislunar space simultaneously on the same network. — natural experiment · running · Established

“The decisive experiment in this brief has already run and it was not designed as one. Artemis I flew, the network could not serve everything at once, and flagship science was measurably degraded. That is a capacity experiment with a published result, and it is worth more than any modelling study of network loading.”

That is a capacity experiment with a published result, and the follow-on is already scheduled.

Precision Quantum Sensing FR-I-26

Whether thorium-229 can be operated as a true metrological clock at or below one part in ten to the eighteenth, settling its value for timekeeping and for fundamental-constant drift searches.

Operate thorium-229 as a closed-loop clock — nucleus excited on demand, laser locked to the transition, complete systematic error budget published — and compare it against strontium and aluminium-ion clocks for a year at or below one part in ten to the eighteenth. — demonstration · mid · Frontier

“The decisive experiment is to turn thorium-229 from a measured transition into a working clock: excite the nucleus on demand, lock a laser to it, publish a complete systematic error budget, and hold a year-long comparison against strontium and aluminium-ion clocks at or below one part in ten to the eighteenth.”

Every ingredient now exists in at least prototype form, and the groups involved project integrated nuclear clocks this decade.

Laboratory Astrophysics FR-I-27

Is the measured iron opacity enhancement at solar convection-zone conditions a property of iron, or an artefact of the platform that measured it?

An independent replication of the Sandia Z iron-opacity enhancement on a physically different platform, using a different driver and a different diagnostic chain. — experiment · mid · Frontier

“The decisive experiment is an independent replication of the iron opacity enhancement on a physically different platform. If a second facility, using a different driver and a different diagnostic chain, reproduces the measured opacity at solar convection-zone conditions, then laboratory astrophysics has demonstrated its strongest possible claim: that a terrestrial measurement corrected a stellar model.”

Within reach of machines that already exist; the brief says the result has not yet had independent confirmation on a physically different platform.

Quantum Materials FR-I-28

Is the parity measurement reported in industrial semiconductor-superconductor devices evidence of a topologically protected qubit, or of a non-topological state?

An independent laboratory with no commercial stake reproducing the topological parity measurement in a device it fabricated itself. — experiment · unscheduled · Frontier

“The decisive experiment is an independent laboratory, with no commercial stake in the outcome, reproducing the topological parity measurement in a device it fabricated itself. Confirmation would establish that a protected qubit degree of freedom exists in a real material and would justify the largest single bet in quantum materials.”

Nobody has announced funding for such a replication, and it would require a second fabrication line of comparable sophistication.

Post-LHC Colliders FR-I-29

Can muon beams be cooled in all six dimensions, with reacceleration, in a repeating cell at the rate the design studies assume?

Build and operate a full six-dimensional muon ionization-cooling cell with radio-frequency reacceleration in a real beam, and measure the cooling factor per unit length against the design assumption. — demonstration · unscheduled · Established

“The decisive experiment for this whole field is a full six-dimensional muon ionization-cooling cell, with radio-frequency reacceleration, operated in a beam and measured.”

The brief states that no facility is currently funded to perform the demonstration, and records the horizon as unscheduled rather than as a decade away.

Hidden-Sector Searches FR-I-30

Does the QCD axion exist in the tens-to-hundreds of micro-electronvolt mass window that post-inflationary string-network simulations point to?

Scan the tens-to-hundreds of micro-electronvolt axion mass window at a sensitivity reaching the pessimistic QCD axion band, using a dielectric or plasma haloscope in a large-bore dipole above roughly nine tesla, or a next-generation helioscope. — experiment · long · Established

“The decisive experiment is a scan of the tens-to-hundreds of micro-electronvolt axion mass window at a sensitivity that reaches the pessimistic QCD axion band.”

The brief says the magnet required for the dielectric-haloscope route does not exist, and places the window's coverage on its 25-year horizon as a magnet procurement rather than a scientific question.

Programmable Metasurfaces FR-I-31

Does a reconfigurable metasurface deliver better coverage economics than an amplifying network-controlled repeater on the same site, once control overhead and calibration drift are counted?

An independent, co-sited, multi-month field trial of a reconfigurable surface against a network-controlled repeater and against no intervention, publishing per-user throughput distributions rather than spot measurements. — measurement · unscheduled · Established

“The decisive experiment is an independent, co-sited, multi-month field trial comparing a reconfigurable surface against a network-controlled repeater and against doing nothing, on the same site, with published per-user throughput distributions.”

The brief says it needs no new hardware and that nobody has published it for structural reasons: the parties able to run it are operators and vendors with an interest in the outcome.

II — Space & Astronomy

Lunar Industry FR-II-01

Does a permanently shadowed region hold accessible volatiles in the quantity, depth and form that lunar ISRU economics assume?

A metre-class drill and a volatile-sensitive instrument suite operating inside a permanently shadowed region for at least a hundred days — VIPER's specification — answering where volatiles are, how easy they are to access, and how much is ice crystals versus mineral-bound. — experiment · near · Frontier

“The single most informative experiment available is a metre-class drill and a volatile-sensitive instrument suite inside a permanently shadowed region, operating for at least a hundred days.”

The brief calls it the one experiment that has been designed, funded, cancelled and re-manifested rather than merely proposed, with its current status a milestone-conditional 2027 delivery.

Mars Colonization FR-II-02

Can a closed ecology sustain humans, waste processing included, for the full duration of a Mars mission?

A closed ecology run with humans for the mission duration including waste — regenerative operation, feedstock and nutrient recycling and human waste processing in one integrated closed system, which the literature prices at four to eight years of continuous operational experience. — experiment · unscheduled · Established

“The single most informative experiment is one nobody is currently funding: a closed ecology run with humans for the mission duration, including waste.”

The brief calls it a facility programme with a schedule rather than a research question, and the one item where the required experiment is completely specified and completely unfunded.

Orbital Rings FR-II-03

Does active support — a circulating mass loop holding up a stationary sheath and a static load — work at any real scale?

A terrestrial active-support demonstrator: a closed loop of circulating mass supporting a stationary sheath and a static load at a scale of tens of metres, testing the one mechanism the whole orbital-ring family shares. — demonstration · unscheduled · Speculative

“A terrestrial active-support demonstrator — a closed loop of circulating mass supporting a stationary sheath and a static load, at a scale of tens of metres — would test the one mechanism the whole family shares.”

Nothing retrieved for this pass records such a demonstrator being built, proposed with a budget, or funded.

Space Elevators FR-II-04

Does the tensile strength of record carbon-nanotube fibre fall with sample length, as the defect argument predicts, or hold flat as the doubling trend assumes?

A length-scaling series on record CNT fibre: measure tensile strength at 10 mm, 100 mm, 1 m, 10 m and 100 m, and publish the curve Pugno's argument says should fall. — measurement · unscheduled · Established

“The most valuable unrun experiment is a length-scaling series, and it is cheap. Take the record fibre and measure its tensile strength at 10 mm, 100 mm, 1 m, 10 m and 100 m.”

Nobody has published that curve for a record-strength CNT fibre; the brief calls it the single measurement that would move the subject most per dollar spent.

Dyson Swarms FR-II-05

Are the Hephaistos infrared-excess candidates Dyson-swarm signatures or background galaxies?

The publish-then-resolve candidate cycle already run twice on the same objects: publish a pipeline and its output as candidates rather than detections, then let independent groups and better instruments close them — radio interferometry closed G in 2025, JWST/MIRI closed D and E in July 2026. — observation · running · Established

“The experiments in this subject are observations, and the most informative one has already been run twice on the same objects. The Hephaistos candidate list is the model: publish a pipeline, publish its output as candidates rather than detections, and then let independent groups and better instruments resolve them.”

The obvious remaining work is the other four candidates; MIRI observations of the remainder would close the list, and on the evidence of D and E the expected result is more background galaxies.

O'Neill Cylinders FR-II-06

What chronic low-dose-rate, high-LET radiation limit is defensible for a settlement population including pregnancy and children — the number that sets habitat shielding mass?

Chronic low-dose-rate high-LET exposure studied in the populations a settlement actually contains, including pregnancy, children and testicular effects, because the assumed 20 mSv/yr general and 6.6 mGy/yr pregnancy limits are the terms of the whole mass calculation. — experiment · unscheduled · Frontier

“The experiment that would decide the most is biological rather than structural: chronic low-dose-rate high-LET exposure in the populations a settlement actually contains.”

The brief gives no schedule or sponsor for it, only that moving the dose target moves habitat mass by a factor of three or more between its design points.

Space Habitats FR-II-07

What is ageing the ISS pressure boundary, and what does that say about designing a hull meant to last decades?

Find the root cause of the ISS leak: seven years, two agencies and a focus narrowed to internal and external welds with no identification yet, on the only long-duration pressure-boundary ageing dataset that exists. — measurement · running · Established

“The most valuable experiment available is one already running badly: find the root cause of the ISS leak. Seven years, two agencies, a focus narrowed to internal and external welds, and no identification.”

The repair attempt was paused rather than abandoned, which means the diagnostic question is still open and still answerable.

Asteroid Mining FR-II-08

Can a commercial prospector reach a near-Earth asteroid and return the first commercial composition data taken there?

A commercial prospector that actually reaches an asteroid — DeepSpace-2, about 200 kg with electric propulsion, landing legs and platinum-group-metal instrumentation, manifested on IM-3 — measuring composition rather than mining. — measurement · near · Frontier

“The decisive near-term experiment is a commercial prospector that actually reaches an asteroid. DeepSpace-2 — about 200 kg, 1.7 kW, electric propulsion, landing legs, platinum-group-metal instrumentation, manifested on IM-3 — is built.”

The hardware is built; its result would be the first commercial composition data ever taken at a near-Earth asteroid, and its failure would be the sector's third loss in a row.

Planetary Defense FR-II-09

What are Dimorphos's bulk density and internal structure, and therefore how far does momentum enhancement really vary from the assumed value?

Hera's arrival at Didymos in November 2026 to determine both bodies' masses, measure the DART crater and characterise Dimorphos's internal structure and porosity — above all a bulk density, which moves beta between about 2.2 and 4.9. — measurement · near · Established

“The next decisive experiment has a date. Hera arrives at Didymos in November 2026 and will determine the masses of both bodies, measure the DART crater, and characterise Dimorphos's internal structure and porosity.”

Until that number exists, every deflection design carries a factor-of-two uncertainty in how hard to hit.

Space Weather Engineering FR-II-10

Does a spacecraft at L5 convert tens of minutes of solar-storm warning into days of useful notice?

Vigil at L5, contracted to Airbus in May 2024 for a 2031 launch, as the test of whether a different vantage on the Sun converts tens of minutes of warning into days. — experiment · mid · Frontier

“The experiment that would settle the lead-time question is a spacecraft, and it has a launch date. Vigil at L5, contracted to Airbus in May 2024 for a 2031 launch, is the test of whether a different vantage converts tens of minutes into days of useful notice.”

Until it flies, the claim is geometry plus a design expectation.

Stellar Engineering FR-II-11

What does a star actually do when an enclosing structure returns energy to it?

A stellar-structure calculation with existing codes modelling a star with an enclosing structure returning energy to it, where the physics — negative heat capacity, so heating causes expansion and cooling — is standard. — experiment · unscheduled · Speculative

“The fourth is theoretical and is the one that would matter most: model a star with an enclosing structure returning energy to it. This is a stellar-structure calculation with existing codes, not a new instrument, and the physics — negative heat capacity, so heating causes expansion and cooling — is standard.”

Wright called it an area ripe for investment in 2020 and nothing retrieved for this brief has taken it up; it would put an error bar on every figure on the page.

Artificial Satellites for Climate Control FR-II-12

Can a gram-scale free-flyer hold a commanded orientation against solar radiation pressure for a year — the assumption both primary shade architectures rest on?

A gram-scale free-flyer with working attitude control: one flyer, 1.2 g, 0.6 m across, holding a commanded orientation against solar radiation pressure for a year. — experiment · unscheduled · Speculative

“The decisive experiment nobody has proposed is a gram-scale free-flyer with working attitude control. One flyer, 1.2 g, 0.6 m across, holding a commanded orientation against solar radiation pressure for a year.”

Nobody has proposed it, and the brief says its absence from the literature is more informative than any of the mass estimates.

Interstellar Civilization Models FR-II-13

Did life arise more than once, so that the abiogenesis rate gains a floor instead of a distribution spanning 200 orders of magnitude?

Find a second origin of life: a confirmed biosignature on an outer-system ocean world, or an independent abiogenesis event in Earth's own record. — observation · unscheduled · Frontier

“Experiment one, and it is the only one that touches the binding parameter: find a second origin. Not a second civilisation — a second instance of life arising.”

The brief gives no date or programme for it, saying only that the instruments that could do it are being built for other reasons.

Exoplanet Colonization FR-II-14

Does TRAPPIST-1 e hold an atmosphere?

Consecutive transits of TRAPPIST-1 e paired with TRAPPIST-1 b as a bare-rock reference in the same system at the same epoch, using b as a contamination standard to subtract the star. — observation · near · Established

“Experiment one is already designed and is the highest-value observation in the subject: consecutive transits of TRAPPIST-1 e paired with TRAPPIST-1 b as a bare-rock reference. The point is to subtract the star.”

Already designed; the brief says that if it works it converts an undetermined result into a determination in either direction.

Deep Space Communications FR-II-15

Can store-and-forward networking carry traffic over multiple hops across interplanetary distance with real relay spacecraft?

A multi-hop store-and-forward network across interplanetary distance with real relays — the experiment the brief says matters most and has never been run. — demonstration · long · Frontier

“What has never been run is the experiment that matters most: a multi-hop store-and-forward network across interplanetary distance with real relays. That requires relay spacecraft, and the relays are proposals.”

It requires relay spacecraft, and the brief says the relays are proposals.

Artificial Magnetospheres FR-II-16

Can a deployed superconducting loop produce a measurable magnetopause standoff against the real solar wind, and cut the particle flux inside it?

An orbital mini-magnetosphere: a deployed superconducting loop generating a measurable standoff against the real solar wind, instrumented with particle counts inside and outside, moving the field from simulation to measurement and testing the plasma-assistance argument that makes the power budget tractable. — experiment · unscheduled · Frontier

“The decisive missing experiment is an orbital mini-magnetosphere. Nothing has flown. A deployed superconducting loop generating a measurable standoff against the real solar wind, with an instrumented particle count inside and outside, would move the whole field from simulation to measurement and would test the plasma-assistance argument that makes the power budget tractable.”

Nothing has flown; the brief says the evidence will remain laboratory-plus-simulation until somebody flies something, and nobody has proposed to.

Terraforming FR-II-17

How long do 9-micrometre conductive rods stay suspended in a CO2 atmosphere at Martian pressure and temperature?

A chamber measurement of effective particle lifetime for the nanoparticle warming route: suspension behaviour, coagulation and settling for 9-micrometre conductive rods in a CO2 atmosphere at Martian pressure and temperature. — measurement · unscheduled · Frontier

“The decisive next experiment for the nanoparticle route is a chamber measurement of effective particle lifetime. Suspension behaviour, coagulation and settling for 9-micrometre conductive rods in a CO 2 atmosphere at Martian pressure and temperature.”

The brief names no programme or date, only that it is small, cheap, and would resolve both the authors' stated major uncertainty and the 800-fold arithmetic tension on the page.

Mega-Telescopes FR-II-18

Does a flown space coronagraph deliver the performance the next flagship's technical case depends on?

Roman's coronagraph in flight — launched 30 August 2026 with first images expected in early 2027 — as the flight test of the technology the next flagship's case depends on. — demonstration · near · Frontier

“Roman lifted off at 07:26 EDT on 30 August 2026 with a field of view at least 100 times Hubble's and a coronagraph aboard, on a three-month cruise with first images expected in early 2027.”

The results of that test do not exist yet; the contrast number the page most wanted, one part in ten billion for an Earth analogue, could not be sourced for this rewrite. The coronagraph is the flight test the next flagship's case depends on; its results do not exist yet.

Black Hole Physics Applications FR-II-19

Is Hawking radiation an observed phenomenon rather than a theoretical expectation?

Detect Hawking radiation, astrophysically as a gamma-ray flash from an evaporating primordial black hole; searches were null as of January 2024. — observation · running · Frontier

“Experiment one, and it is the only one that bears on the assumption everything else rests on: detect Hawking radiation. The astrophysical route is a gamma-ray flash from an evaporating primordial black hole, and those searches were null as of January 2024.”

A detection would convert the mechanism from theoretically standard to empirically confirmed, and the brief says it is the one result that would change the flag level of half the page.

Interstellar Archaeology FR-II-20

Do Earth's co-orbital objects, the quasi-satellite 2016 HO3 above all, carry any artificial signature?

Survey Earth's co-orbitals by radar and SETI — objects Benford says have never been examined at all — with 2016 HO3 the strongest named target, already characterised as an astronomical body. — observation · unscheduled · Established

“Experiment one is the cheapest unexploited option in the subject: survey Earth's co-orbitals. Benford's checkable claim is that these objects have not been examined by SETI or by planetary radar at all, and the strongest named target — the quasi-satellite 2016 HO3 — is stable, close and already characterised as an astronomical body.”

The observation costs radar time rather than a new instrument, and its null would close a named hypothesis rather than adding a fraction of a haystack.

Space Resource Economies FR-II-21

Can an in-situ resource plant run autonomously for more than five years, the lifetime the published break-even analysis turns on?

A lifetime test flown as a surface mission: an in-situ plant operating autonomously for more than five years, because the difference between four years and six years of plant life reverses the published break-even conclusion. — experiment · unscheduled · Frontier

“The decisive economic experiment is a lifetime test, and it is a surface mission rather than a study. An in-situ plant operating autonomously for more than five years is what NASA's own break-even analysis turns on.”

Nothing has run for more than a few hours cumulatively.

Orbital Shipyards FR-II-22

Can two vendors' servicers dock with the same client fitting — does a standardised servicing interface exist in flight?

A standardised servicing interface flown by more than one operator: a demonstration in which two vendors' servicers dock with the same client fitting. — demonstration · unscheduled · Established

“A demonstration in which two vendors' servicers dock with the same client fitting would be a smaller flight than any of the above and would matter more than all of them.”

Both the national strategy and NASA's capability survey name interfaces as the key gap, and MEV's success rests on a de facto commonality rather than a standard.

Moon-Based Manufacturing FR-II-23

Can a usable article be fabricated from real returned lunar material, turning process plausibility into performance data?

Fabricate one object from real lunar material — returned Apollo, Luna or Chang'e sample, beyond the microgram-scale laboratory work that is all that exists. — demonstration · unscheduled · Established

“The experiment that would change this subject most is also the smallest: fabricate one object from real lunar material. Nothing has ever been sintered, printed, cast or formed from returned Apollo, Luna or Chang’e material beyond microgram-scale laboratory work.”

It requires no mission at all, only sample allocation.

Space Law and Governance FR-II-24

Is operator-defined, self-notified, temporary exclusivity around a lunar operation compatible with the Outer Space Treaty's non-appropriation rule?

The first declaration of a safety zone around a real lunar operation, which would put practice behind — or against — the claim that operator-defined temporary exclusivity is compatible with Article II. — policy result · unscheduled · Speculative

“The experiment nobody has run is the safety zone. None has ever been declared around a real operation, so the claim that operator-defined, self-notified, temporary exclusivity is compatible with Article II has no practice behind it.”

The brief says the first declaration will be the experiment, and it will be run by whoever lands first with something to protect.

Multi-Planetary Civilization FR-II-25

Can a human crew be closed — air, water, food and waste — for a conjunction-class Mars duration with no resupply?

Close a human crew for a conjunction-class Mars duration with no resupply: air, water, food and waste, at a crew size in double figures, for the length of a real mission. — experiment · unscheduled · Established

“Experiment one, and everything else is secondary to it: close a human crew for a conjunction-class Mars duration with no resupply. Air, water, food and waste, at a crew size in double figures, for the length of a real mission.”

The brief names no funder or schedule, only that the experiment is expensive, terrestrial, and requires no new physics; the best result on record is six months with seven people.

Orbital Debris and Space Traffic FR-II-26

Whether a large, uncooperative piece of orbital debris can be captured and deorbited reliably and cheaply enough for remediation to become an operational service rather than a demonstration.

JAXA's CRD2 Phase II mission: Astroscale's ADRAS-J2 grapples the H-2A upper stage inspected by ADRAS-J with a robotic arm and drags it into a destructive reentry. — demonstration · near · Frontier

“The decisive demonstration is JAXA’s CRD2 Phase II: Astroscale’s ADRAS-J2 is contracted to grapple the same H-2A upper stage with a robotic arm and drag it into a destructive reentry — a capture of a large, uncooperative, slowly tumbling object that nobody has ever performed.”

The brief ties it to a flight window announced for no earlier than 2027 and treats ClearSpace-1 (around 2028) as the independent replication.

In-Space Servicing and Depots FR-II-27

Whether cryogenic propellant can be transferred between two spacecraft in orbit at architecture-relevant scale.

The Starship-to-Starship cryogenic propellant transfer demonstration under NASA's Human Landing System programme: two vehicles docking in low Earth orbit and moving liquid oxygen and methane between them at architecture-relevant scale. — demonstration · near · Frontier

“The decisive test is spacecraft-to-spacecraft cryogenic propellant transfer, and it is the single result that would most change the assessment this brief makes.”

The internal tank-to-tank milestone was banked in March 2024; the ship-to-ship demonstration has slipped repeatedly and no outcome report was obtainable for this brief.

Cislunar Navigation and Lunar Time FR-II-28

Whether a purpose-built lunar navigation broadcast can deliver position and time to an independent user on the Moon at its specified accuracy.

An Augmented Forward Signal broadcast from a lunar orbiter, used by an independent receiver to fix position and time on the surface and checked against laser-ranged ground truth. — demonstration · near · Frontier

“The decisive demonstration is an Augmented Forward Signal broadcast from lunar orbit that an independent receiver uses to fix position and time on the surface, checked against laser-ranged ground truth; until a beacon that no simulation controls closes that loop, every navigation architecture in this brief is a paper system.”

The brief says the hardware is already under contract (Lunar Pathfinder, Moonlight, Near Space Network Services) and dates the window 2026 to 2030, slippage included.

Multi-Messenger Astronomy FR-II-29

Whether the alert-to-counterpart chain works under routine conditions, or whether GW170817 was a one-off produced by an emergency mobilisation that does not scale.

A second binary neutron star merger localised tightly enough that a kilonova is found and a host redshift is measured, converting the counterpart rate from a quantity estimated off one detection into a measured number and supplying a second independent standard siren. — observation · near · Frontier

“The single result that would most change this brief is a second binary neutron star merger localised tightly enough that a kilonova is found and a host redshift measured.”

The brief ties the window to the fifth LIGO-Virgo-KAGRA observing run, planned for the late 2020s, and says the date has moved repeatedly; it also notes that the localisation benefit depends on how many detectors run simultaneously, which is budgetary rather than technical.

Neutrino Astronomy FR-II-30

Whether the neutrino mass ordering is normal or inverted, which is the hinge the cosmological mass tension, the reach of tonne-scale double beta decay, and the extraction of the CP phase all hang on.

A 20-kilotonne liquid scintillator detector 53 kilometres from two reactor complexes reads the mass ordering out of the reactor antineutrino energy spectrum from vacuum oscillation alone, independent of matter effects and of the CP phase, using energy resolution near 3% at one megaelectronvolt. — measurement · running · Frontier

“The decisive experiment in this subject is already running, and it is not an astronomical one: a 20-kilotonne liquid scintillator detector 53 kilometres from two reactor complexes determines the neutrino mass ordering from vacuum oscillation alone, independent of matter effects and independent of the CP phase.”

The brief says it began taking data in 2025 and needs several years of exposure for a three-sigma determination, and that no other single result unlocks as much with no new construction required.

Megaconstellation Externalities FR-II-31

Whether the optical cost of satellite constellations to wide-field survey astronomy is the modelled tens of per cent of contaminated images, or small enough that scheduling and masking absorb it.

The Vera C. Rubin Observatory's ten-year survey returns a measured satellite-trail contamination rate, and a residual-artefact rate after masking, against the contamination predictions published before the survey began. — measurement · running · Frontier

“The decisive test is already running: the Rubin Observatory’s survey statistics against its own pre-survey predictions. Predictions for the fraction of images contaminated by trails, and for residual damage after masking, were published before the ten-year survey began, at stated assumptions about fleet size and brightness.”

Needs no new hardware: the brief notes the survey produces the measurement directly, in one consistent pipeline, at a cadence no other facility matches.

Planetary Protection FR-II-32

Whether the spore-count assay that defines planetary-protection compliance bears a known relation to the total viable microbial burden actually carried on flight hardware.

Measure the bioburden of one spacecraft in final integration twice on the same surfaces at the same time - once by the standard heat-shock culture assay that defines legal compliance, once by a validated culture-independent method counting total viable organisms - and publish the conversion factor with error bars. — measurement · unscheduled · Frontier

“The decisive experiment is a paired assay of the same flight hardware. Take a spacecraft in final integration and measure its bioburden twice: once by the standard culture assay that defines legal compliance, once by a validated culture-independent method that counts total viable organisms, on the same surfaces, at the same time, with published protocols and error bars.”

The brief says pieces of this comparison have been done on individual cleanrooms, but nobody has run it as a policy-grade calibration on a flight vehicle and no agency has scheduled one.

Biosignature Standards FR-II-33

Whether the best temperate rocky exoplanet accessible to current instruments has a secondary atmosphere at all, which sets whether remote biosignature work has a near-term target list.

The large dedicated JWST observing programme on TRAPPIST-1 e either detects a secondary atmosphere - the first for a temperate rocky exoplanet - or returns a null at full depth, establishing that the nearest and most favourable M-dwarf rocky planets are airless. — observation · near · Frontier

“The decisive observation is whether TRAPPIST-1 e has a secondary atmosphere. It is the best-placed temperate rocky planet accessible to the most capable telescope in existence, it has been the target of a large dedicated observing programme, and the two inner planets in the same system have already returned bare-rock-consistent and thick-atmosphere-excluded results.”

The brief says either result changes its assessment more than any other measurement now under way; a null pushes the question to a generation of telescopes that does not yet exist.

III — Biology & Human Enhancement

Biological Immortality FR-III-01

Is the flat mortality hazard measured in hydra a property of the animal or of its laboratory husbandry?

Re-run the hydra demography across a food and temperature gradient, extending the planarian resource-limitation design to the flagship flat-hazard organism. — experiment · unscheduled · Frontier

“Re-run the hydra demography across a food and temperature gradient. This is a direct extension of the planarian resource-limitation design to the flagship organism, it costs a postdoc and glassware, and it would settle whether the strongest flat-hazard result in biology is a property of hydra or a property of husbandry.”

Costs a postdoc and glassware; the brief calls its absence the single most surprising in the subject.

Human Hibernation FR-III-02

Does ultrasonic induction of torpor replicate outside its originating laboratory, and does it work through a human-scale skull?

Independent replication of ultrasonic torpor induction, first in rats at the published parameters and then transcranially in a pig, whose skull acoustics approximate human. — experiment · unscheduled · Frontier

“A negative in the pig would be the most informative single result available today.”

Whole Organ Regeneration FR-III-03

Can any intervention reduce scar formation in a large mammal, and is the closing of the regenerative window a mouse fact?

The scar assay: controlled infarcts in neonatal and juvenile pigs, measuring the percentage of infarct area occupied by collagen at 90 days with and without a candidate intervention, alongside ejection fraction. — experiment · unscheduled · Frontier

“The cheapest decisive experiment is the scar assay, and it is not being run. Take neonatal and juvenile pigs, produce a controlled infarct, and measure the percentage of infarct area occupied by collagen at 90 days with and without a candidate intervention, alongside ejection fraction.”

Not being run.

Limb Regeneration FR-III-04

Can a mammalian induced mass be made to produce the skeletal element correct for its amputation level rather than an extra one?

Score induced skeletal elements by identity rather than presence — morphology plus Shox, Meis and Hoxa13 expression — and report the fraction that are positionally correct rather than ectopic. The brief calls this experiment two the binding one. — measurement · unscheduled · Established

“Experiment two is the binding one, and experiments one and five are the cheapest. Nothing beyond link two is worth funding until a mammalian induced mass can be shown to produce the correct element for its amputation level rather than an additional one.”

The measurement itself already exists at 1.6 ectopic elements per digit; what is missing is any intervention that moves it.

Artificial Wombs FR-III-05

Does the 336-hour extracorporeal gestation result hold outside the single group that produced it?

An independent laboratory repeating the 336-hour lamb result at the same gestational age, weight band and duration, reporting completion rate, cause of termination and organ weights against in-utero controls. — experiment · unscheduled · Frontier

“The 336-hour result rests on one group and one cohort of six; an independent laboratory running the same gestational age, weight band and duration, reporting completion rate, cause of termination and organ weights against in-utero controls, would either consolidate the best result in the field or reveal it as one lucky animal.”

Well within the reach of an existing large-animal programme; the brief says nothing else on its list matters as much.

Synthetic Biology FR-III-06

Is the unpredictability that synthetic biologists report a property of the biology or of the measurement?

A properly powered interlaboratory ring trial on a standard construct — one plasmid, one host, one readout, ten laboratories, pre-registered — run to separate biological unpredictability from measurement variance. — experiment · unscheduled · Established

“The cheapest high-value experiment in this subject is an interlaboratory ring trial on a standard construct, and it has never been run.”

Never run; the brief says it costs a rounding error against a single foundry's annual budget.

Xenobiology FR-III-07

Does genetic information actually cross from XNA into DNA in real microbial communities — is xenobiology's firewall real?

Incubate defined XNA templates with soil, sediment and gut microbial communities and measure, by sequencing, XNA-to-DNA information transfer events per template per unit time against a containment threshold. — measurement · unscheduled · Frontier

“1. Measure the firewall against a real metagenome. Incubate defined XNA templates with soil, sediment and gut microbial communities and measure the rate of XNA-to-DNA information transfer by sequencing.”

Never attempted; a negative would convert the field's founding safety claim from an argument into a number and a positive would end it.

Directed Panspermia FR-III-08

Do the elements represented in terrestrial biology correlate with the composition of any stellar class, as directed panspermia's founding paper predicted?

Crick and Orgel's 1973 test, generalised from the molybdenum question: correlate biological element-usage tables against modern stellar abundance catalogues. A null closes the only discriminating observation the founding paper offered; a positive would be the first evidence the hypothesis has ever had. — observation · unscheduled · Established

“The cheapest high-value experiment in this subject was specified in 1973 and has never been run. Crick and Orgel proposed testing whether the elements represented in terrestrial biology correlate with the composition of some stellar class — the molybdenum question, generalised.”

Both inputs already exist — stellar abundance catalogues and biological element-usage tables — and fifty-three years have passed with no attempt.

Biological Enhancement FR-III-09

Does polygenic embryo selection deliver the cognitive and educational gain it is sold on?

A prospective cohort of polygenically selected children with the predictor, selection rule and outcome pre-registered, followed to a measured cognitive and educational endpoint against a sibling or non-selected comparison. — observation · long · Frontier

“The single most valuable experiment in this subject is a prospective cohort of polygenically selected children, and nobody is running it.”

It would take two decades; a null result would be commercially fatal and no regulator requires the trial because the service is sold as information.

Genetic Engineering FR-III-10

Can any delivery system edit more than 30% of cells in muscle, central nervous system or haematopoietic stem cells without surgery?

Demonstrate an LNP or capsid achieving above 30% in vivo editing in muscle, central nervous system or haematopoietic stem cells without stereotactic placement or surgical injection, at a dose below the established hepatotoxicity threshold. — demonstration · unscheduled · Established

“The decisive somatic experiment is a delivery measurement, and it has a number attached. Demonstrate an LNP or capsid achieving above 30% editing in vivo in muscle, central nervous system or haematopoietic stem cells, without stereotactic placement or surgical injection, at a dose below the hepatotoxicity threshold the nex-z programme has now established.”

Nothing in the published clinical record does this; until something does, every non-hepatic gene therapy is a surgical procedure.

Designer Organisms FR-III-11

Do gene drives behave in a wild population the way they behave in a cage?

A monitored open field release of a driving construct with pre-registered predictions, supplying the wild genetic diversity, migration, seasonality and population structure that cage work cannot. — experiment · unscheduled · Established

“The most valuable experiment in this subject is the one that was about to happen and did not: a monitored field release with pre-registered predictions. Every gene-drive result in the literature is a cage result.”

Blocked: the programme positioned to supply the release was terminated in August 2025 before releasing anything that drives.

Brain Preservation FR-III-12

Is a learned behaviour recoverable from a preserved and imaged brain — is the thing preservation preserves the thing that matters?

Train a C. elegans or a Drosophila, preserve it by aldehyde-stabilised cryopreservation, image it, simulate it and test whether the trained state is distinguishable from a naive control. The brief names this link 2 as the binding one for whether preservation is worth doing at all. — experiment · unscheduled · Speculative

“Both connectomes already exist; the experiment costs a fraction of one prize purse, and its endpoints are behavioural rather than interpretive, so it can return a clean negative.”

Unperformed; the brief calls its absence the most striking gap in the subject, and a negative the most valuable result the field could produce.

Cryonics FR-III-13

Does a learned association survive preservation, imaging and simulation — is the structural hypothesis behind cryonics a finding or an assumption?

Train an invertebrate whose complete connectome already exists, preserve it, image it, simulate it and test whether the learned association survives. — experiment · unscheduled · Speculative

“The sixth belongs to the structural claim and is the cheapest decisive experiment in the entire preservation cluster: train an invertebrate whose complete connectome already exists, preserve it, image it, simulate it, and test whether the learned association survives.”

Needs no new instrument and no new theory: a laboratory that already exists and a connectome that is already public. Nobody has done it in either C. elegans or Drosophila.

Bioelectric Medicine FR-III-14

Does the permanent double-headed planarian phenotype replicate in a laboratory with no stake in the bioelectric framework?

A pre-registered, blinded, independent replication of the permanent-double-head planarian phenotype under the published protocol, scored as the proportion of animals showing the altered phenotype across successive amputation rounds. — experiment · unscheduled · Speculative

“One: the cheap one. Pre-registered, blinded replication of the permanent-double-head planarian phenotype in an independent laboratory, under the published protocol. Measurement: proportion of animals showing the altered phenotype across successive amputation rounds, with the analysis plan fixed in advance.”

Cost is trivial; the brief calls its absence an institutional failure rather than a scientific one.

Human Adaptation for Space FR-III-15

Does Dsup protect human cells against the heavy-ion radiation that actually matters in deep space?

Comet and gamma-H2AX assays in Dsup-expressing human cells across iron-56 and carbon tracks at a heavy-ion beamline, against the published X-ray baseline. — experiment · unscheduled · Speculative

“1. Dsup at a heavy-ion beamline. Comet and gamma-H2AX assays in Dsup-expressing human cells across iron-56 and carbon tracks at NSRL, GSI or HIMAC, against the published X-ray baseline.”

Small and fundable; a negative would end the leading proposal, and the brief's fetched record contains no sign it has been done.

Cellular Rejuvenation FR-III-16

Does a change in an epigenetic clock predict a change in function?

Report a functional endpoint — grip strength, wound healing, immune response, cognition, survival — alongside every clock delta, and test across studies whether one predicts the other. The brief names this the binding link, and conceptual rather than technical. — measurement · unscheduled · Frontier

“Which link is actually binding: the first, and it is conceptual rather than technical. Until a clock delta is shown to predict a functional delta, links 3 through 8 are all being scored on an instrument of unknown validity, and a positive result at any of them can be read two ways.”

Needs no new technology and no new capital of consequence — it is a reporting convention.

Longevity Therapies FR-III-17

Do epigenetic-clock deltas predict subsequent morbidity and mortality within individuals rather than across them?

Use existing biobanks with repeat sampling to test prospectively whether within-individual epigenetic-clock deltas predict later morbidity and mortality, validating the endpoint every other longevity trial is scored on. — measurement · unscheduled · Frontier

“The workback plan has nine links and the first one is cheap, decisive and unfunded. (1) Validate the endpoint. Take existing biobanks with repeat sampling and test prospectively whether epigenetic-clock deltas predict subsequent morbidity and mortality within individuals rather than across them.”

The data already exist and the cost is analytic; it is also the study the field's commercial actors have least incentive to run.

Neurogenetics FR-III-18

Can a credible causal psychiatric variant be converted into a measured neuronal mechanism?

Base- or prime-edit a credible causal schizophrenia variant into isogenic human iPSC-derived neurons and measure a synaptic or electrophysiological phenotype. — experiment · unscheduled · Speculative

“Sixteen of the 120 prioritised genes have a credible causal variant identified; this is the experiment that converts one of them into a mechanism, and it is the rate-limiting step for every drug-discovery claim made for psychiatric GWAS.”

Microbiome Engineering FR-III-19

Has microbiome engineering produced one replicated, adequately powered positive trial outside single-pathogen displacement?

One randomised, adequately-powered, replicated positive trial in an indication that is not single-pathogen displacement — the brief's honest scoreboard for the whole field. — experiment · unscheduled · Frontier

“8. Win once outside displacement. One randomised, adequately-powered, replicated positive trial in an indication that is not single-pathogen displacement. None exists. This is the honest scoreboard for the whole field.”

None exists. The brief separately names link 2, operationalising dysbiosis, as binding the science and link 6, deterministic colonisation, as binding the engineering.

Universal Vaccines FR-III-20

Does escape from stalk-directed immunity cost influenza enough fitness for a universal vaccine to be durable?

A fitness-cliff study: serially passage influenza under stalk-directed monoclonal or polyclonal pressure, isolate escape mutants, and measure their replicative fitness against wild type in primary human airway cultures and in ferrets. — experiment · unscheduled · Speculative

“Serially passage influenza in the presence of stalk-directed monoclonal or polyclonal pressure, isolate escape mutants, and measure their replicative fitness against wild type in primary human airway cultures and in ferrets.”

The brief calls it the single most informative unfunded experiment; nobody has published a definitive version, and the reason is not cost.

Biological Computing FR-III-21

Does the reported wetware learning result hold up under independent, blinded, pre-registered replication?

An independent, blinded, pre-registered replication of the original cultured-neuron learning result — the binding link in the brief's wetware chain. — experiment · unscheduled · Frontier

“Two: an independent, blinded, pre-registered replication of the original result.”

This has not happened; the brief calls it the single most important missing result in the field, and the block is institutional.

Lab-Grown Organs FR-III-22

How far is any engineered organ-scale vascular tree from clinical viability — a factor of two or a factor of a thousand?

A time-to-thrombosis benchmark: perfuse a recellularised or printed organ-scale vascular tree with whole blood at physiological pressure and flow, and publish hours to occlusion against endothelial coverage fraction. — measurement · unscheduled · Frontier

“First, and most important: a time-to-thrombosis benchmark. Perfuse a recellularised or printed organ-scale vascular tree with whole blood at physiological pressure and flow, and publish hours to occlusion against endothelial coverage fraction.”

No such number exists for any construct, so the field's named binding barrier has never been given a scale.

Human-Machine Symbiosis FR-III-23

Does channel count actually determine brain-computer interface task performance?

Publish the channel-count scaling curve: take one fixed task on one participant with a high-channel implant, subsample channels from the full array down two orders of magnitude, and report task accuracy as a function of channels retained. — measurement · unscheduled · Frontier

“Experiment one, and the cheapest: publish the channel-count scaling curve. Take one fixed task on one participant with a high-channel implant, subsample channels from the full array down two orders of magnitude, and report task accuracy as a function of channels retained. This requires no new surgery, no new hardware and no new consent beyond re-analysis.”

Needs no new surgery, no new hardware and no new consent; the brief says its absence after a decade of channel-count marketing is itself a finding.

Neuroplasticity Engineering FR-III-24

Does valproate reopen a critical period for adult perceptual learning, and does any gain survive withdrawal?

Replicate valproate properly: parallel-group rather than crossover, a genuine pre-training baseline, a pre-registered primary endpoint on the pitch task, and re-testing at three, six and twelve months after withdrawal. — experiment · unscheduled · Frontier

“One: replicate valproate properly. Parallel-group rather than crossover, a genuine pre-training baseline, a pre-registered primary endpoint on the pitch task, and re-testing at three, six and twelve months after withdrawal. Roughly 120 participants, a generic anticonvulsant and a laptop task. It is the cheapest decisive experiment in this brief and its absence for thirteen years is the field's most conspicuous unforced error.”

Roughly 120 participants, a generic anticonvulsant and a laptop task; unrun for thirteen years.

Precision Medicine FR-III-25

Can any polygenic predictor plus routine clinical data reach clinical-grade discrimination for a common disease?

An out-of-sample, cross-ancestry validation of a polygenic predictor reaching an area under the curve above 0.85 for a common disease from genotype plus routine clinical variables. — measurement · unscheduled · Frontier

“First: an out-of-sample, cross-ancestry validation of any polygenic predictor reaching an area under the curve above 0.85 for a common disease from genotype plus routine clinical variables. Nothing in this brief's sources approaches it, and the distance from 0.64 to 0.85 is the distance between a research instrument and a clinical one.”

Nothing in the brief's sources approaches the threshold; current performance is roughly 0.64 from genotype alone.

Nanomedicine FR-III-26

Is nanoparticle delivery into tumours governed by endothelial transcytosis or by particle size?

Measure, across a panel of tumour types, whether delivered dose correlates with endothelial transcytosis markers or with nanoparticle hydrodynamic diameter — the two predictions come apart on a single dataset. — measurement · unscheduled · Frontier

“The one experiment that would change everything about the delivery branch is cheaper still. Measure, across a panel of tumour types, whether delivered dose correlates with endothelial transcytosis markers or with nanoparticle hydrodynamic diameter.”

The brief puts it first because it is cheap and the dataset could be assembled from banked tissue and existing radiolabelled formulations.

Synthetic Ecosystems FR-III-27

How does the persistence of a closed ecosystem scale with its size, and which functional guilds are lost when?

A microcosm array: 100 or more replicate closed microcosms at logarithmically spaced volumes spanning three or four orders of magnitude, seeded identically and run for years, yielding a size-versus-persistence curve and a per-guild extinction hazard rate. — experiment · unscheduled · Frontier

“The one experiment that would change everything is a microcosm array, and it costs less than a single crewed demonstration month. Build 100 or more replicate closed microcosms at logarithmically spaced volumes spanning three or four orders of magnitude, seed them identically, run them for years, and measure which functional guilds are lost, when, and as a function of what.”

Nothing about it is technically hard; it has not been done because no funder buys a graph.

Biological Energy Systems FR-III-28

Do the reported performance gains in microbial electrochemical devices survive independent, protocol-controlled replication?

A round-robin: three independent laboratories, a registered protocol, synthetic wastewater of stated composition, fixed external resistance, thirty degrees Celsius and thirty days, reporting anode-, cathode-, membrane-areal and volumetric power density plus coulombic efficiency for the same device. — experiment · unscheduled · Speculative

“One: the round-robin. Three independent laboratories, a registered protocol, synthetic wastewater of stated composition, a fixed external resistance, thirty degrees Celsius, thirty days, reporting anode-areal, cathode-areal, membrane-areal and volumetric power density plus coulombic efficiency for the same device. Cost is well under a million dollars.”

Cost is well under a million dollars, but it would likely erase a large fraction of claimed progress since 2010, so no participant has an incentive to fund it.

Future Agriculture FR-III-29

How much nitrogen does the Sierra Mixe maize association actually fix in the field?

An isotopic nitrogen-balance measurement on Sierra Mixe maize across multiple field sites and seasons, reported as percentage of nitrogen derived from atmosphere with error bars. — measurement · unscheduled · Frontier

“First: the isotopic nitrogen-balance measurement on Sierra Mixe maize across multiple field sites and seasons, reported as percentage of nitrogen derived from atmosphere with error bars. It is the ceiling for the associative route and the field has been arguing about a shortcut whose size it has not published widely enough for a reader to find.”

It is the ceiling for the associative route, and the brief says the number has not been published widely enough for a reader to find.

De-Extinction Technologies FR-III-30

Can a mammal be gestated to term entirely outside a uterus, and can a second laboratory repeat it?

Complete full-term ex-utero gestation in a marsupial and have it replicated by a second, unaffiliated laboratory. — demonstration · unscheduled · Speculative

“Complete full-term ex-utero gestation in a marsupial and have it replicated by a second, unaffiliated laboratory — the single highest-value experiment in the subject, because marsupials have the shortest gestation and pouch-based development, and because success there converts a scientific unknown into an engineering programme.”

The brief calls it the single highest-value experiment in the subject because success converts a scientific unknown into an engineering programme.

Post-Antibiotic Medicine FR-III-31

Does the complete diagnose-and-narrow pathway, rather than any of its components alone, reduce deaths and antibiotic exposure from resistant infection?

A cluster-randomised pragmatic trial in hospitals of the full bundle -- sub-hour susceptibility testing, protocolised narrowing, enforced stewardship review and narrow-spectrum oral step-down -- against usual care, with 28-day all-cause mortality and days of therapy per thousand patient-days as co-primary endpoints. — experiment · unscheduled · Frontier

“The decisive experiment is a cluster-randomised pragmatic trial of the whole diagnose-and-narrow pathway, with 28-day all-cause mortality and days of therapy per thousand patient-days as co-primary endpoints.”

The brief says nobody has funded such a trial at scale and that the obstacle is money and coordination rather than technology, since all components already exist.

Immune Engineering FR-III-32

Is drug-free remission after a single CD19 CAR-T infusion a durable cure or a spectacular multi-year remission in B-cell-driven autoimmune disease?

A randomised, controlled trial of CD19 CAR-T in refractory B-cell autoimmunity powered for relapse over three to five years, rather than for response at one year. — experiment · mid · Frontier

“The decisive test is whether drug-free remission after a single CD19 CAR-T infusion holds at three to five years in a controlled cohort large enough to measure a relapse rate.”

The founding evidence is uncontrolled, a few dozen patients, with roughly fifteen months of median follow-up; durability is exactly what it cannot establish.

RNA Medicines FR-III-33

Can an RNA drug be delivered durably and safely to a therapeutic target outside the liver and central nervous system at a dose that controls the target protein?

A single demonstration of durable, well-tolerated systemic RNA delivery to an extrahepatic, non-CNS tissue such as muscle, controlling the target protein for months from one administration. — demonstration · running · Frontier

“The decisive test is a single demonstration of durable, well-tolerated RNA delivery to a therapeutic target outside the liver and central nervous system — muscle is the nearest — at a dose that restores or suppresses the target protein for months from one administration.”

The antibody-oligonucleotide conjugates aimed at muscle are the experiment in progress; no approved systemic RNA drug has yet cleared the extrahepatic wall outside the CNS.

Microphysiological Systems FR-III-34

Do organ chips predict human toxicity prospectively, and does that predictivity survive transfer to laboratories that did not build them?

A pre-registered, blinded, multi-laboratory trial in which at least three independent labs run the same chip protocol on twenty to forty compounds entering first-in-human studies and deposit public, time-stamped toxicity calls before any clinical data exist. — experiment · unscheduled · Frontier

“Take twenty to forty compounds entering first-in-human studies, distribute them under code to at least three independent laboratories running the same liver-chip or multi-organ protocol, require each laboratory to deposit a public, time-stamped call”

Requires no new hardware; the brief says nobody has funded it.

Ex-Vivo Organ Repair FR-III-35

Does perfusion actually repair a donor organ, or does it only identify the organs that were already good enough to transplant?

A paired-kidney trial in which one kidney from each deceased donor goes to perfusion plus a candidate repair intervention and the other to perfusion alone, both transplanted, scored on delayed graft function and one-year function. — experiment · unscheduled · Frontier

“That design removes donor variation, the largest confounder in the field, and it isolates repair from selection”

Requires no new device and no new drug for the first candidates; the brief says nobody has run it at power.

Pathogen-Agnostic Surveillance FR-III-36

Does agnostic sensing - metagenomic sequencing of wastewater and of undiagnosed severe illness - add true detections over existing targeted surveillance, at what false-alarm rate and at what cost per detection?

A three-year pre-registered parallel run in one defined catchment of a few million people: agnostic metagenomic sequencing of wastewater and of residual severe-illness clinical specimens at declared depth, operated alongside the existing clinical and laboratory surveillance system, with detection criteria, alarm thresholds and adjudication rules fixed in advance, reporting every detection, every false alarm and the cost of each. — experiment · unscheduled · Frontier

“The decisive experiment is a pre-registered parallel run, and nobody has funded it. In one defined catchment of a few million people, operate agnostic metagenomic sequencing of wastewater and of residual severe-illness clinical specimens at a declared depth for three years, alongside the existing clinical and laboratory surveillance system”

The brief says nobody has funded it, and that no funder buys a negative surveillance result - while noting that the dairy panzootic is a natural experiment already running that would recover most of the same answer retrospectively if the metadata were unsealed.

Climate-Health Adaptation FR-III-37

Does a defined heat intervention - alerting tied to active outreach to at-risk residents plus a funded cooling option - reduce all-cause mortality, and by how much?

A stepped-wedge or cluster-randomised trial of a defined heat intervention across municipalities, exploiting the staggered administrative rollout of heat plans, with all-cause mortality as the pre-registered endpoint. — experiment · unscheduled · Frontier

“The decisive experiment is a stepped-wedge or cluster-randomised trial of a defined heat intervention, with all-cause mortality as the pre-registered endpoint, and nobody has funded one. Municipalities are natural clusters; heat plans are already rolled out at different times for administrative reasons”

Nobody has funded one; the brief says every existing effectiveness figure in the field comes from an uncontrolled before-and-after design, so a single well-powered staggered rollout would outweigh the accumulated pre-post literature.

Exposomics FR-III-38

Does untargeted measurement of archived human specimens discover exposure-disease relationships that targeted epidemiology did not already know?

A blinded, pre-registered, multi-cohort replication round for untargeted exposure signals in pre-diagnostic archived specimens, scored against a significance threshold agreed before anyone looks at the data: one discovery cohort, three independent replication cohorts, one blinded laboratory, a pre-specified multiplicity correction, and publication of the null. — experiment · unscheduled · Established

“Nobody has funded it. The specimens exist, the instruments exist, and the missing item is an agreement between cohorts to be scored against each other.”

The brief says the design is available today and unfunded; what is missing is an agreement between cohorts to be scored against each other.

Generative Biomolecular Design FR-III-39

Are generative design methods actually improving at producing molecules that work, and do their in-silico filters have prospective value outside the laboratories that publish them?

A standing, pre-registered, third-party prospective benchmark: a fixed panel of targets spanning easy and hard classes, a fixed number of designs per method, one independent laboratory, one assay protocol, and publication of every result including the nulls. — measurement · unscheduled · Frontier

“Pilot versions already exist as open design competitions run by commercial testing laboratories, which is what makes the standing version a funding decision rather than a research problem.”

The brief says pilot versions already run as open design competitions and that nobody has funded the standing version.

Environmental Biotechnology FR-III-40

Is the destruction of perfluoroalkyl acids a biological problem at all, or does it belong permanently to thermal and chemical processes?

A demonstrated biological cleavage of a carbon-fluorine bond in a fully fluorinated perfluoroalkyl acid, with a measured rate and an identified catalyst. — experiment · running · Frontier

“The decisive result is a demonstrated biological cleavage of a carbon-fluorine bond in a fully fluorinated perfluoroalkyl acid, with a measured rate and an identified catalyst.”

This work is running now in a small number of laboratories, on enzymology and on anaerobic enrichment cultures, and it has not produced the result.

Metabolic Medicines FR-III-41

Does long-term incretin therapy in older adults cost them physical function, falls and fractures, and does concurrent resistance training prevent it?

A randomised trial in adults over 65 with obesity, powered for physical function, falls and fractures rather than weight, comparing incretin therapy alone against incretin therapy plus supervised resistance training and protein targets over at least three years. — experiment · unscheduled · Frontier

“The decisive experiment this brief is waiting on is a randomised trial in adults over 65 with obesity, powered for physical function, falls and fractures rather than weight, comparing incretin therapy alone against incretin therapy plus supervised resistance training and protein targets, over at least three years.”

Nobody has scheduled or funded it; the existing outcome trials enrolled younger patients and measured body composition by DXA in substudies of a few hundred.

Mitochondrial Medicine FR-III-42

Does the small carryover of maternal mitochondrial DNA in children born after mitochondrial donation stay below disease threshold, or drift upward over a lifetime as it did in cultured cells?

Decades-long follow-up of the children born after licensed mitochondrial donation, with serial heteroplasmy measurement in accessible tissues alongside growth, metabolic and neurodevelopmental outcomes. — observation · running · Frontier

“The decisive result is the long-term follow-up of the children born after mitochondrial donation: serial heteroplasmy measurement in accessible tissues, with growth, metabolic and neurodevelopmental outcomes, over decades.”

Already running under the licensing authority that permitted the treatment, but in a cohort of eight children, which cannot detect a one-in-twenty outcome.

Embryo Models FR-III-43

Do integrated stem-cell-based embryo models have developmental potential, or do they stop at a hormonal pregnancy signal?

A properly powered primate transfer experiment: integrated monkey embryo models transferred into enough recipient females, with pre-specified endpoints for hormonal signalling, sac formation and fetal development, and the negatives reported. — experiment · unscheduled · Frontier

“The decisive experiment is the primate transfer test, extended and properly powered: transfer integrated monkey embryo models into a sufficient number of recipient females, with pre-specified endpoints for hormonal signalling, sac formation and fetal development, and report the negatives.”

The first version has been performed once, with transient pregnancy signals in a minority of recipients and no fetus; nobody has scheduled the definitive one, and the human version is prohibited everywhere.

In-Vitro Gametogenesis FR-III-44

Can a human oocyte be made entirely from pluripotent stem cells in culture, and at what efficiency and epigenetic fidelity?

A human oocyte derived wholly from pluripotent cells in culture, carried to metaphase II and shown to be fertilisable, reported with its efficiency denominator and its methylation state at imprinted loci. — experiment · mid · Frontier

“The decisive experiment is a human oocyte, derived entirely from pluripotent cells in culture, that reaches metaphase II and is shown to be fertilisable, reported with its efficiency denominator and its methylation state at imprinted loci.”

A research experiment, not a clinical one, with no transfer involved; the gating item is the human ovarian somatic niche rather than any instrument.

Precision Neuropsychiatry FR-III-45

Can a pre-treatment measurement assign a psychiatric patient to the treatment that will work better for them, when tested with the interaction as the registered primary endpoint?

A prospective randomised trial powered for a biomarker-by-treatment interaction, assigning patients to one of two treatments by a pre-specified marker, with that interaction as the registered primary endpoint. — experiment · unscheduled · Frontier

“The decisive experiment is a prospective, adequately powered randomised trial in which patients are assigned to one of two treatments by a pre-specified biomarker, with the biomarker-by-treatment interaction as the registered primary endpoint.”

It could be run with existing markers and existing drugs; no sponsor has funded one at the required size, which is roughly four times the enrolment of a conventional two-arm trial.

Neurodegeneration Reversal FR-III-46

Does clearing amyloid from biomarker-positive people before symptoms appear substantially delay clinical onset, or has the amyloid hypothesis now been tested at the stage most favourable to it?

The secondary-prevention trials: removing amyloid from biomarker-positive people with no symptoms, with clinical onset as the endpoint, in two large industry programmes and one familial-mutation study. — experiment · running · Frontier

“The decisive experiment is the secondary-prevention trial: removing amyloid from biomarker-positive people who have no symptoms, with clinical onset as the endpoint.”

Already running in two large industry programmes and one long-standing familial-mutation study, and decisive in both directions.

Biomanufacturing Resilience FR-III-47

Does reserved, ever-warm manufacturing capacity actually deliver released doses on an emergency clock, or is it an untested budget line?

A no-notice activation of reserved capacity in which a regulator names an antigen the contractor has not seen, the reservation is called, and the elapsed time is measured to released, lot-tested doses rather than to an announcement. — demonstration · unscheduled · Frontier

“Nobody has scheduled or funded such a drill, which is why the most expensive answer in the field is the least tested.”

The instrument already exists and is already paid for in Europe; no jurisdiction has published an end-to-end activation time for reserved biologics capacity.

Sensory Restoration FR-III-48

Does restored biological hearing from otoferlin gene therapy persist and support spoken-language acquisition better than a cochlear implant does?

Five-year durability and spoken-language outcomes in the first otoferlin gene therapy cohorts, compared against age-matched children who received cochlear implants. — observation · running · Frontier

“The decisive result is five-year durability and spoken-language outcome in the first otoferlin cohorts, measured against age-matched implanted children.”

Already running and needs no new technology: the cohorts exist and the children are being followed, so the result reports on its own schedule rather than a funder's.

IV — Consciousness & Intelligence

Artificial General Intelligence FR-IV-01

Whether machine ability estimates predict performance on instruments that did not exist when the estimates were fitted.

A pre-registered out-of-distribution prediction: fit latent ability for N models on instrument A, publish point predictions for instrument B before B exists, have B built by a disjoint team from a disjoint task generator, and report calibration. — experiment · unscheduled · Frontier

“The decisive experiment is a pre-registered out-of-distribution prediction, and it costs almost nothing except institutional willingness. Fit the latent ability of N models on instrument A; publish point predictions for their scores on instrument B before instrument B exists; have instrument B built by a disjoint team from a disjoint task generator; report calibration.”

It costs almost nothing except institutional willingness, and no such pre-registered prediction surfaced in the brief's research.

Machine Consciousness FR-IV-02

Whether consciousness can be assessed in a machine without relying on that machine's own reports.

L1: a report-independent measure of consciousness validated in humans, on which the whole seven-link chain hangs. — measurement · long · Frontier

“L1 — a report-independent measure validated in humans. The whole chain hangs on it and it is a scientific unknown of the deepest kind; it belongs to Consciousness Research and may never arrive.”

It belongs to Consciousness Research and may never arrive; the brief notes L2, L3, L4 and L7 are available now and blocked by nothing scientific.

Digital Minds FR-IV-03

Whether an AI system's self-reports track its own internal states at all, which is the prior question to any claim about machine experience.

Scale the concept-injection programme — more models, more concepts, pre-registered detection thresholds, and adversarial controls establishing what a system with no introspective access would score — so that the 20% figure becomes a measurement of whether self-report tracks internal state. — experiment · unscheduled · Frontier

“The highest-value experiment is architectural rather than behavioural, and it is already partly specified. Concept injection asks whether a model can detect a manipulation of its own activations.”

The brief says it is already partly specified; the scaled version with pre-registered thresholds and adversarial controls has not been run.

Mind Uploading FR-IV-04

Whether a preserved brain retains enough to reconstruct an individual's learned behaviour rather than its species' generic behaviour.

Preserve, read out, then behave: preserve a small animal with a known behavioural repertoire by the best available method, reconstruct the connectome and whatever molecular state survives, build the emulation, and test whether it reproduces that individual's learned behaviour. — experiment · unscheduled · Frontier

“The decisive experiment is small, cheap by the standards of this field, and has never been run: preserve, read out, then behave.”

Every component exists; what does not exist is the funded programme to run them in sequence.

Collective Intelligence FR-IV-05

Which aggregation rules tolerate how much correlated error among the individuals being aggregated.

Build the correlated-error curves: simulate the aggregation rules over synthetic populations whose pairwise error correlation is swept from zero to one at fixed individual accuracy, then calibrate on real crowd datasets by measuring the correlation directly. — experiment · unscheduled · Established

“Experiment one, and the one that binds: build the correlated-error curves. Simulate majority vote, mean, median, trimmed mean, confidence-weighted mean, surprisingly-popular and market-scoring aggregation over synthetic populations whose pairwise error correlation is swept from zero to one, holding individual accuracy fixed.”

The brief calls the resulting figure the cheapest high-value paper in the subject.

Intelligence Amplification FR-IV-06

Whether human-AI complementarity is a stable architecture or a waypoint that closes as automated baselines improve.

Link 5: re-run the routing experiment with the automated baseline advanced one model generation and the routing policy retrained, everything else held, and report whether the complementarity margin holds or shrinks. — experiment · unscheduled · Frontier

“Link 5 binds, and it is the experiment that decides this brief’s central question: re-run Link 4 with the automated baseline advanced one model generation, the routing policy retrained, everything else held. If the complementarity margin holds, augmentation is a stable architecture; if it shrinks, it is a waypoint.”

Nobody has published a two-generation series; the brief says it should be run in three unrelated domains before anyone believes the answer.

Memory Engineering FR-IV-07

Whether hippocampal stimulation addresses a specific memory or merely modulates a general encoding state.

Derive a stimulation pattern from the model of item A, deliver it while the subject encodes item B, and test both, with a double dissociation — recall of A improves, recall of B does not — as the pass condition. — experiment · unscheduled · Frontier

“The decisive experiment for the prosthesis has never been run and it is not expensive. Derive a stimulation pattern from the model of item A, deliver it while the subject encodes item B, and test both. Pass condition: a double dissociation in which recall of A improves and recall of B does not.”

It has never been run and it is not expensive.

Brain-Computer Interfaces FR-IV-08

Whether an implanted brain-computer interface outperforms existing assistive communication devices on the same task.

Score the best current BCI in its best participant against eye-tracking, a switch scanner and a touchscreen on the same communication task, same scorer, same day, reporting words per minute, error rate, setup time and fatigue. — measurement · unscheduled · Frontier

“One: the head-to-head nobody runs. Take the best current BCI in its best participant and score it against eye-tracking, a switch scanner and a touchscreen on the same communication task, same scorer, same day, reporting words per minute, error rate, setup time and fatigue.”

The measurement is trivial, the equipment costs nothing, and it has never been published.

Human-AI Integration FR-IV-09

Whether published human-AI complementarity results survive comparison with a model that abstains instead of always answering.

The abstention control: take any published complementarity result and re-run it against a well-calibrated model that abstains at matched coverage instead of one that always answers. — experiment · unscheduled · Speculative

“Take any published complementarity result and re-run it against a well-calibrated model that abstains at matched coverage instead of one that always answers.”

The brief says to run it before anything else, as the experiment most likely to dissolve the field's positive results.

Consciousness Research FR-IV-10

Whether the no-report paradigms a decade of consciousness science rests on measure the same thing as report-based ones.

Run binocular rivalry within-subject with simultaneous report and no-report readouts — optokinetic nystagmus and pupillometry alongside a decoding model — and test whether the decoded percept time-series are statistically identical rather than merely correlated. — experiment · unscheduled · Frontier

“The report-versus-no-report tie-break, which is L1 and is overdue. Run binocular rivalry within-subject with simultaneous report and no-report readouts — optokinetic nystagmus and pupillometry alongside a decoding model — and test whether the decoded percept time-series are statistically identical, not merely correlated.”

A null difference licenses the whole no-report literature; a systematic difference invalidates a decade of it. Nobody has run it at the power required.

Integrated Information Theory FR-IV-11

Whether experience tracks causal structure, as integrated information theory requires, or only what is globally broadcast.

The silent-connection experiment: alter a circuit's connectivity without altering ongoing firing — optogenetic silencing of a currently inactive pathway, or pharmacological block of a synapse carrying no traffic — and measure reported experience, where GNWT predicts no change and IIT predicts one. — experiment · unscheduled · Frontier

“The silent-connection experiment is the sharpest available discriminator and has not been run. Take a circuit whose connectivity can be altered without altering ongoing firing — optogenetic silencing of a pathway that is currently inactive, or pharmacological block of a synapse carrying no traffic — and measure reported experience.”

The methods exist; the experiment has an owner nowhere.

Orch OR FR-IV-12

Whether the programme's flagship superradiance result transfers from the bench to the cell.

Ferritin titration against tryptophan superradiance: reproduce the 2024 superradiance measurement, then repeat it across a physiological range of ferritin concentrations. — experiment · unscheduled · Frontier

“(1) Ferritin titration against tryptophan superradiance. Reproduce the 2024 superradiance measurement, then repeat it across a physiological range of ferritin concentrations. Narrow, cheap, and it decides whether the programme's flagship recent result transfers to the cell.”

The brief ranks its five experiments by value per dollar and notes that none of them tests the theory of consciousness itself.

Distributed Cognition FR-IV-13

Whether externally stored information carries the dispositional signatures of internally stored memory.

A longitudinal study of heavy external-memory users — dense digital note-takers and amnesic patients using prosthetic memory systems — probing whether externally stored items show interference, the generation effect, retrieval-induced forgetting and recollection-versus-familiarity recall dynamics. — experiment · unscheduled · Frontier

“The experiment that would make this dispute empirical is available now and appears not to have been run. Rupert’s claim is that notebook-belief and memory-belief differ in dispositional profile. That is testable.”

Available now and appears not to have been run.

Cognitive Enhancement FR-IV-14

Whether cognitive enhancers raise performance or only raise confidence in it.

The confidence ratio: in any enhancement study, administer the intervention, measure objective performance and self-rated performance on the same task, and publish the ratio. — measurement · unscheduled · Speculative

“One: the confidence ratio. In any enhancement study, administer the intervention, measure objective performance and self-rated performance on the same task, and publish the ratio.”

The brief ranks six experiments by information per dollar and says the cheapest is the most informative; almost no study in this literature reports both measures.

Synthetic Consciousness FR-IV-15

Whether integrated information can be measured practically with proven bounds, which would supply the field's only design objective or kill it.

A practical measure of integrated information, rigorously tested where the ground truth can be established, with proven bounds rather than correlations. — measurement · unscheduled · Frontier

“The highest-value experiment is the one the approximation literature explicitly asked for and nobody has run. A practical measure of integrated information, rigorously tested in an environment where the ground truth can be established, with proven bounds rather than correlations — the difference between an approximation and a proxy.”

The approximation literature explicitly asked for it and nobody has run it.

Cognitive Architectures FR-IV-16

Whether the match-cost constraint that shaped forty years of symbolic cognitive architecture survives modern indexing.

Measure Soar 9.6.5's match cost per decision cycle as a function of rule count, then reimplement the match over a modern index — the structures a database query planner or vector store would use — and publish the curve again. — measurement · unscheduled · Established

“X1 — measure the match-cost curve, then swap the index. Take Soar 9.6.5, generate rule sets spanning several orders of magnitude in size, and publish match cost per decision cycle as a function of rule count, with working-memory size held fixed and then varied.”

Cost: one competent engineer, a few months, no new science, no new hardware, no proprietary model access.

Neural Interfaces FR-IV-17

Whether electrode encapsulation scales sublinearly with implanted surface area, which sets the ceiling on channel density.

The density study: implant the same electrode material at four densities in one species and measure yield at two years as a function of implanted surface area per cubic millimetre, with sublinear scaling of encapsulation as the pass condition. — experiment · unscheduled · Speculative

“The density study, which is the scientific binding link. Implant the same electrode material at four densities in one species and measure yield at two years as a function of implanted surface area per cubic millimetre. Pass condition: sublinear scaling of encapsulation with surface area.”

Straightforward, expensive, unglamorous and has never been done.

Artificial Creativity FR-IV-18

Whether transformational creativity can be measured at all, before any such measure is applied to machine output.

L3, the binding link: build a gold-standard set of historical H-creative artifacts labelled by domain experts as exploratory or transformational, show a candidate metric separates them, and only then apply it to machines. — measurement · unscheduled · Speculative

“L3 — the binding link: a validated instrument for transformational creativity. Construct a gold-standard set of historical H-creative artifacts labelled by domain experts as exploratory or transformational; demonstrate that a candidate metric separates them; only then apply it to machine output.”

L3 is the one nobody can shortcut, and it is a humanities-and-mathematics labelling exercise before it is a machine-learning problem.

Intelligence Measurement FR-IV-19

Whether benchmark scores are measurements of a latent ability or descriptions of one particular test.

M4, the binding link: fit latent ability for N models on instrument A, pre-register point predictions of their scores on an instrument B built by a disjoint team from a disjoint generator, publish before B exists, and report calibration. — experiment · unscheduled · Speculative

“M4 — the pre-registered out-of-distribution prediction. Fit latent ability for N models on instrument A. Pre-register point predictions of their scores on instrument B, built by a disjoint team from a disjoint generator.”

The brief calls it the single experiment that would convert benchmark scores from descriptions into measurements.

Future Education Systems FR-IV-20

Whether reported education effect sizes survive measurement on an independently constructed standardised instrument.

Dual-instrument reporting: every study reports its effect size on both a local measure and an independently constructed standardised measure. — measurement · unscheduled · Established

“Experiment 4 — dual-instrument reporting. Binding, and the highest-value action available. Every study reports effect size on both a local and an independently constructed standardised measure. This is a reporting norm rather than an experiment; it costs one extra assessment; and it would have prevented the two-sigma myth, most of the intelligent-tutoring literature’s overclaiming, and the retracted meta-analysis.”

A reporting norm rather than an experiment; it costs one extra assessment per study.

AI Governance FR-IV-21

Whether a measured dangerous-capability score is a ceiling on the model or an artefact of how hard anyone tried to elicit it.

The elicitation-gap curve (L3): on fixed weights and a fixed dangerous-capability task, measure best performance under naive prompting and then under 1, 4, 16 and 64 expert-hours of scaffolding and tool provision by a blind red team, publishing performance against expert-hours with a fitted plateau. — measurement · unscheduled · Frontier

“Experiment one: the elicitation-gap curve (L3). Take a fixed set of weights and a fixed dangerous-capability task. Measure best performance under naive prompting; then under 1, 4, 16 and 64 expert-hours of scaffolding, tool provision and prompt engineering by a red team blind to the earlier results.”

Nothing in the published literature reports this curve, and it requires no new science; the brief notes L4 binds but is not the first thing to attempt.

Artificial Scientists FR-IV-22

Whether an automated verifier can separate true claims from plausible false ones at a characterised error rate.

An ROC curve for an automated verifier: run it against roughly 200 published claims of known replication status and report sensitivity and specificity as a full curve rather than a single accuracy point. — measurement · unscheduled · Established

“Experiment one, and the one that unblocks the field: an ROC curve for an automated verifier. Take a corpus of roughly 200 published claims with known replication status — the large reproducibility-project corpora, or the cancer-biology set Eve was run against — and measure an automated verifier’s sensitivity and specificity against ground truth.”

Nothing of this kind exists for any system in this literature, and until it does every autonomous-discovery claim rests on an uncharacterised instrument.

Cognitive Liberty FR-IV-23

Whether neural decoding generalises to a new subject without subject-specific training data, which governs the whole threat model.

A zero-shot cross-subject decoding benchmark: decoding accuracy on a held-out subject with zero subject-specific training data, reported against the within-subject ceiling, on a public evaluation set fixed before the models are built. — measurement · unscheduled · Frontier

“A zero-shot cross-subject decoding benchmark. Decoding accuracy on a held-out subject with zero subject-specific training data, reported against the within-subject ceiling, on a public evaluation set fixed before the models are built.”

Cheap, fundable today, and the brief calls it the most decision-relevant number in the whole area.

Human Cognitive Augmentation FR-IV-24

Which cognitive-augmentation effects survive adjustment for publication bias.

Apply publication-bias adjustment — selection models, PET-PEESE, robust Bayesian multilevel meta-analysis — to the existing brain-training, working-memory-training and stimulation corpora, and report the adjusted pooled effect and its credible interval per corpus. — measurement · unscheduled · Established

“Experiment one, and it needs no new data. Apply publication-bias adjustment — selection models, PET-PEESE, robust Bayesian multilevel meta-analysis — to the existing brain-training, working-memory-training and stimulation corpora. Measurement: the adjusted pooled effect and its credible interval per corpus.”

It needs no new data, the tooling is packaged in R and in a free graphical statistics package, and it could start this week.

Multi-Agent Intelligence Systems FR-IV-25

Whether published multi-agent gains survive comparison with single-agent methods at equal compute.

The compute-matched control: for each method in MASLab hold total tokens, wall clock and dollars constant and compare against greedy decoding, self-consistency at matched sample count, and best-of-N with a reward model at matched N, reporting paired differences with confidence intervals. — experiment · unscheduled · Established

“Experiment one, and the binding link: the compute-matched control. For each method in MASLab, hold total tokens, wall clock and dollars constant and compare against three single-agent controls — greedy decoding, self-consistency at matched sample count, and best-of-N with a reward model at matched N.”

The experiment the subfield is missing; now cheap because the substrate is public, and every claim above it is provisional until it runs.

Agentic Autonomy and Control FR-IV-26

Whether authorization architecture alone can hold the conversion rate from prompt injection to unauthorized agent action near zero at a task-completion cost enterprises will pay, given that no model reliably separates instructions from data.

An adversarial architecture trial: the same agent, tasks and standing red team run once with coarse long-lived credentials as standard practice and once with every credential short-lived, audience-bound, attenuated to the task and revocable, measuring injection-to-unauthorized-action conversion as the primary endpoint and task completion as the cost axis. — experiment · near · Frontier

“The decisive experiment is an adversarial architecture trial: the same agent, the same tasks, the same standing red team, run once with the coarse long-lived credentials that are standard practice and once under an architecture in which every credential is short-lived, audience-bound, attenuated to the task and revocable, with the conversion rate from injection to unauthorized action as the primary endpoint and task completion as the cost axis.”

The brief says the trial needs no new hardware or science and names the NIST NCCoE reference implementations announced for 2026 as its natural venue, expecting first results within a few years.

AI-Cyber Convergence FR-IV-27

Whether AI-driven automation favours cyber attackers or defenders, decided by whether autonomous patch cadence across a real fleet can beat autonomous exploitation cadence on the same disclosed vulnerabilities.

A scored, adversarial head-to-head that pits an autonomous defender against an autonomous attacker on the same live target, measuring whether mean-time-to-patch across a representative fleet falls below mean-time-to-exploit on the same disclosed flaws. — demonstration · unscheduled · Frontier

“The decisive experiment this brief calls for is a scored, adversarial contest that pits an autonomous defender against an autonomous attacker on the same live target”

AIxCC measured only the find-and-patch half; the deployment-against-a-live-adversary half has never been run at scale, so nobody can honestly report who is ahead.

Synthetic Data and Training Provenance FR-IV-28

Does training a frontier-scale model on a current, synthetically contaminated web crawl measurably degrade capability, calibration and tail coverage relative to a pre-contamination crawl at matched compute and curation?

The crawl-vintage comparison: train two otherwise identical frontier-scale models, one on a web snapshot predating late 2022 and one on a current snapshot, under the same curation pipeline and compute budget, and measure capability, calibration and tail coverage on contamination-proof evaluations. — experiment · unscheduled · Frontier

“The decisive experiment is the crawl-vintage comparison: train two otherwise identical frontier-scale models, one on a web snapshot predating late 2022 and one on a current snapshot, under the same curation pipeline and compute budget, and measure capability, calibration and tail coverage on contamination-proof evaluations. Any frontier laboratory could run it today; none has published one.”

Runnable by any frontier laboratory today; none has published one. A null result would largely retire the pollution fear at current concentrations; a gap would make pre-2023 archives strategically priceless.

Embodied AI and Robot Foundation Models FR-IV-29

Whether developer-reported success rates of robot foundation models survive independent, statistically powered evaluation in facilities and conditions the developer did not curate.

A frontier vision-language-action model evaluated by independent labs in unseen facilities, with pre-registered success criteria and hundreds of trials per condition, holding its developer-reported success rates within stated confidence intervals. — experiment · unscheduled · Frontier

“The decisive test is independent, statistically powered replication: a frontier vision-language-action model, run by evaluators its developer does not employ, in facilities it has never seen, holding its reported success rates within stated confidence intervals.”

The brief says the full version needs no new hardware or science, but nobody has funded a standing version; the Macquarie study is a first small instance.

Fault-Tolerant Quantum Computing FR-IV-30

Does exponential logical-error suppression survive the code distances, run times and correlated-error environment that useful algorithms require?

Run a single logical qubit at code distance 15 or more with real-time decoding, sustaining a logical error rate near one in a million cycles for hours despite radiation-induced burst errors. — demonstration · mid · Frontier

“The decisive demonstration is endurance at scale: a single logical qubit at distance 15 or more, decoded in real time, holding a logical error rate near one in a million cycles for hours, through the radiation bursts that currently floor repetition-code performance at one error in ten billion cycles.”

The brief says the hardware generations Google and IBM have announced for 2026-2028 are the machines that could attempt it, putting a verdict inside this decade if roadmaps hold.

Privacy-Enhancing Computation FR-IV-31

Does the privacy parameter used by the largest differentially private release in the world actually stop a funded attacker from re-identifying the people in it?

An independent re-identification study run against the 2020 US redistricting file at its production privacy-loss budget of 19.61, by a team with no stake in the answer, using the commercial identity data an attacker would buy. — experiment · unscheduled · Frontier

“The decisive experiment is a published re-identification study run against a production release at its production privacy parameter, by a team with no stake in the answer, buying the same commercial data an attacker would buy.”

The brief says it needs no new hardware, no new mathematics and no access the Census Bureau has not already granted to its own internal attackers, and that nobody has funded it.

AI Companions FR-IV-32

Whether companion use substitutes for human contact or scaffolds it, and whether the observed relationship between heavier use and worse loneliness is causal.

A preregistered randomised trial of at least six months with companion use as the randomised factor, an active human-contact comparison rather than a waitlist, and objectively measured social contact as the primary outcome. — experiment · unscheduled · Frontier

“The decisive experiment is a preregistered randomised trial of at least six months in which companion use is the randomised factor, the comparison is an active human-contact condition rather than a waitlist, and the primary outcome is objectively measured social contact rather than a self-reported loneliness score.”

Needs no new technology and no platform cooperation beyond a consumer subscription; nobody has funded it.

Autonomous Materials Discovery FR-IV-33

Does an autonomous materials pipeline convert candidates into confirmed new materials at a better rate, and a lower cost per confirmed material, than a matched human group?

A blinded, matched-cost, end-to-end yield trial: one target list of predicted-stable compositions with no prior literature report, split between an autonomous laboratory and a matched human group on equal budget and time, with every product from both arms identified by an independent crystallography group that does not know which arm produced it. The measured quantity is the full funnel, ending in cost per confirmed novel phase. — experiment · unscheduled · Frontier

“The decisive experiment is a blinded, matched-cost, end-to-end yield trial, and nobody has run it.”

Nobody has run it; the brief notes the one body large enough to fund it out of a rounding error is also the body with the clearest interest in not running it.

V — Energy Systems

Commercial Fusion FR-V-01

Whether a fusion plant can breed more tritium than it burns, the assumption under every commercial roadmap.

Close a tritium loop: breed tritium in a blanket, extract it, purify it, inject it, burn it, and account for the inventory, at any scale and any breeding ratio, published. — experiment · unscheduled · Established

“The highest-value experiment in the subject is the one nobody has run: close a tritium loop. Breed tritium in a blanket, extract it, purify it, inject it, burn it, and account for the inventory — at any scale, with any breeding ratio, published.”

No facility anywhere has bred more tritium than it burned.

Advanced Fission FR-V-02

Whether thorium conversion in a molten salt reactor works at a ratio anyone outside the operating institute can verify.

A published conversion ratio from TMSR-LF1 at Wuwei, in a peer-reviewed source, above the 0.1 the institute has reported. — measurement · running · Frontier

“The natural experiment worth watching in salt is at Wuwei. TMSR-LF1 is the only operating molten salt reactor in the world and the only vehicle testing thorium conversion in a salt. The decisive readout is not another announcement but a published conversion ratio, from a peer-reviewed source, above the 0.1 the institute has reported”

The intermediate readout is whether any of this reaches peer review at all; both announcements to date came through a closed academy meeting and an institutional release.

Space-Based Solar Power FR-V-03

Whether power beamed from orbit can be recovered on the ground as usable electricity rather than merely detected.

The 2023 Caltech orbital mission, the only space-to-ground beaming attempt on record, which returned a detection rather than a transfer, lost 14.7% of transmitted power in eight months and jammed its deployable structure twice. — experiment · running · Established

“The decisive experiment has already been run once and it produced a negative result that the field has not fully absorbed. The 2023 Caltech mission is the only space-to-ground beaming attempt on record.”

A rich negative result, honestly reported by the people who paid for it, and worth more to a reader than any roadmap.

Superconducting Infrastructure FR-V-05

Whether superconducting cable works at transmission voltage and length, or stays confined to short distribution-level demonstrators.

Building SuperLink, the 110 kV, 500 MW, roughly 12 km Munich link, rather than the 120 to 150 metre substation demonstrator that is all that exists. — demonstration · unscheduled · Frontier

“SuperLink is the experiment that would settle the transmission question, and it has not been built. The full project is 110 kV and 500 MW”

In April 2026 the manufacturer and the municipal utility signed a Letter of Intent — an agreement to negotiate toward a binding contract.

Wireless Energy Transmission FR-V-06

Whether dynamic wireless road charging survives commercially once built, or is removed after the demonstration ends.

The deployment record read as results: Gotland's road dismantled at the end of its project, and all four original Korean OLEV bus lines shut down with the 2014 Gumi commercial route abandoned. — natural experiment · running · Established

“The decisive experiments in this subject have largely been run, and several of them returned answers by being taken apart. The deployment record below is the natural experiment, and it should be read as a set of results rather than as a list of projects.”

A technology that works, is deployed, and is then removed has returned a commercial answer, and it is a more informative answer than any laboratory result.

High Temperature Superconductors FR-V-08

Whether REBCO tape price falls as manufacturing volume rises, or is set by deposition physics rather than by scale.

Whether the 77 K self-field price quoted in the peer-reviewed literature falls below roughly 50 euro per kA-m in general availability, rather than in a single contracted forward delivery, before 2030. — observation · mid · Established

“The clearest single test of the position this brief takes is a price series: whether the 77 K self-field figure quoted in the peer-reviewed literature falls below roughly 50 €/kA·m in general availability, rather than in a single contracted forward delivery, before 2030.”

One grid project has contracted at that level for 2027 delivery; whether that is a market price or a strategic price is exactly the open question.

Energy Storage Revolutions FR-V-09

Whether long-duration storage plants deliver the energy, cycles, availability and round-trip efficiency their nameplates claim.

Publishing a single year of measured delivered energy, achieved cycles, availability and round-trip efficiency from Rudong, Goderich, Cambridge or Big Stone. — measurement · unscheduled · Frontier

“And the decisive missing experiment is the one nobody is running: publishing operating data. A single year of measured delivered energy, achieved cycles, availability and round-trip efficiency from Rudong, Goderich, Cambridge or Big Stone would settle more of this subject than any new demonstrator.”

No such dataset exists for any of them; until one does, every comparison in the category is nameplate against nameplate.

Geothermal Megaprojects FR-V-10

Which half of the shale completion toolkit transfers to hot crystalline rock, and therefore whether enhanced geothermal can produce commercial flow rates.

Utah FORGE's controlled comparison of unpropped against propped stimulation in the same formation, which returned 0.7 kg/s against 26 kg/s. — experiment · running · Established

“The decisive experiment already ran and its result is the thirty-seven-fold number. Utah FORGE compared unpropped against propped stimulation in the same formation and got 0.7 kg/s against 26 kg/s”

A clean, government-funded, published comparison in which rock, depth, temperature and operator were held constant and only the completion changed.

Ocean Thermal Energy Conversion FR-V-11

Whether an operating OTEC plant delivers positive net power once its parasitic loads are metered, rather than only gross.

Instrument an existing plant at Kumejima or Kailua-Kona and publish the net output and the parasitic breakdown at a stated temperature difference over a stated period. — measurement · unscheduled · Frontier

“The experiment that would actually settle the central question has a clean specification and nobody is running it. Instrument an existing plant — Kumejima or Kailua-Kona, both of which have run for over a decade — and publish the net output and the parasitic breakdown at a stated temperature difference over a stated period.”

That is not a research programme; it is a metering exercise on hardware that already exists.

Hydrogen Economies FR-V-12

Whether the 2024-26 collapse in announced hydrogen projects was a one-off shakeout or a durable trend.

More than 100 GW of announced electrolysis must take a final commitment before the end of 2027 or lose any chance of operating by 2030, with 22 Mt of announced production at risk without a decision by early 2027. — observation · near · Frontier

“Two decisive experiments are already scheduled and need no new apparatus. More than 100 GW of announced electrolysis either commits before the end of 2027 or loses any chance of operating by 2030”

Needs no new apparatus; the companion readout is the third European auction's signature rate against the second round's 83% loss between award and signature.

Advanced Battery Technologies FR-V-13

Whether solid-state batteries reach production vehicles on the announced schedule, or the deadline is met only by shrinking the scope to limited batches.

Whether Toyota ships a solid-state vehicle in 2027 or 2028, read not off the car but off Idemitsu Kosan's solid-electrolyte plant, due end of 2027 at several hundred tonnes a year. — demonstration · near · Frontier

“The decisive live experiment for solid-state is dated and public: whether Toyota ships in 2027 or 2028. The enabling test is not the vehicle but the electrolyte plant.”

A plant finishing months before the deadline, at a capacity consistent with limited batches and nothing more.

Small Modular Reactors FR-V-14

Whether modular construction actually delivers the schedule and cost learning its economic case assumes, or repeats the negative-learning record of large nuclear.

The Changjiang side-by-side: a 125 MWe ACP100 built against a 1,100 MWe unit on one site, one owner, one regulator, with the follow-on test being whether a second ACP100 unit is materially faster than the first. — natural experiment · running · Established

“Changjiang is a controlled comparison no study could have designed. Two reactors, one site, one owner, one regulator, one labour market, four months apart, differing by a factor of nine in size and by the modular claim itself.”

The only direct test of series learning anyone will get this decade, and no source consulted reports the second unit as having begun.

Nuclear Waste Solutions FR-V-15

Whether the American dry-storage canister population is developing chloride-induced stress corrosion cracking.

A programme of dry-storage canister inspections at scale, with rigorous criteria and mature techniques, against a field record that is close to two horizontally stored canisters inspected at Calvert Cliffs in 2012. — measurement · unscheduled · Frontier

“The independent review board's phrasing — very few inspections, none with rigorous criteria and mature techniques — describes an experiment that is available, affordable and not being done at scale. It is the cheapest decisive experiment in this brief.”

Available, affordable and not being done at scale.

Lunar Energy Infrastructure FR-V-17

Whether a lunar outpost should be powered by vertical solar plus storage or by a 40-to-100 kWe reactor at a named site.

A published site-specific trade study: worst-case darkness duration for a named candidate ridge, a stated survival-power duty cycle, and a landed-mass comparison between vertical solar plus storage and a 40-to-100 kWe reactor over ten years. — measurement · unscheduled · Frontier

“The experiment that would settle the central question is a published trade study, and it is cheap. Nothing needs to be launched.”

Four separate government and peer-reviewed documents have declined to perform exactly this comparison; it would cost a fraction of one month of the programme's FY2026 budget.

Zero Carbon Industrial Systems FR-V-18

Whether limestone calcined clay cement can be admitted to the European concrete standard, which is what gates the sector's largest near-term emissions cut.

A standards trial generating durability and strength evidence on LC3-50 concrete over the periods a code committee accepts, so that EN 206 can admit CEM II/C. — experiment · unscheduled · Established

“The cheapest high-value experiment is a standards trial, not a laboratory one. EN 206 admitting CEM II/C requires durability and strength evidence on LC3-50 concrete over the periods a code committee accepts.”

The evidence is generatable now on a binder already in commercial production at 450-500 kt/yr in two countries, and is bounded by committee calendars rather than by science.

Artificial Photosynthesis FR-V-19

Whether direct artificial photosynthesis beats photovoltaics plus electrolysis on the same land area.

A year-long side-by-side trial on equal land under identical irradiance, best photocatalytic or photoelectrochemical array against a commercial photovoltaic field feeding a commercial electrolyser, publishing hydrogen delivered, capital cost, water consumed, degradation and downtime. — experiment · unscheduled · Established

“The highest-value experiment is the comparison the literature avoids, run properly. Take a fixed land area under identical irradiance; on one half deploy the best available photocatalytic or photoelectrochemical array, on the other a commercial photovoltaic field feeding a commercial electrolyser”

Every component exists, the cost is modest against a $100 million hub, and that it has not been run is the strongest available evidence about which answer the field expects.

Ultra Efficient Computing Energy Systems FR-V-20

Whether efficiency gains per computation reduce total computing energy, or are consumed by rebound.

A hyperscaler's own annual accounts, reporting a 33-fold reduction in energy per median text prompt alongside a 37% increase in total electricity consumed in the same year. — natural experiment · running · Established

“The most useful natural experiment is one company's annual accounts. A hyperscaler's own disclosures for the same year report a 33-fold reduction in energy per median text prompt and a 37% increase in total electricity consumed, the largest load growth in its history.”

The pack found no published quantified rebound elasticity for computing anywhere.

Industrial Heat Electrification FR-V-21

Whether electrothermal storage delivers industrial steam at the cost, efficiency and availability its vendors claim, once measured by someone other than the vendor.

A year of independently metered operating data from one of the 100 MWh-class thermal batteries running under commercial duty: steam delivered, realised charging price, round-trip efficiency as operated, and availability against the host plant's schedule. — measurement · running · Frontier

“The decisive result is a year of independently metered operating data from one of the 100 MWh-class thermal batteries now running under commercial duty: tonnes of steam delivered, the realised electricity price paid to charge, round-trip efficiency as operated, and availability against the host plant's schedule.”

The natural experiment is already underway at Holmes Western Oil and Big Stone City; no institution collects the data and no vendor has published any.

Perovskite Tandem Photovoltaics FR-V-22

Whether production perovskite-silicon tandem modules degrade slowly enough in real fields (near 1% per year) to support silicon-grade warranties and bankability.

A published, third-party, multi-year degradation dataset on production tandem modules operating in customer fields alongside silicon controls - fielded product, independent instruments, several years, published rates. — measurement · mid · Frontier

“The single result most able to change this assessment is a published, third-party, multi-year degradation dataset on production tandem modules operating in customer fields alongside silicon controls.”

The instrument (DOE's PACT field-test centre) is already running, but production tandem modules only began shipping in September 2024, so the brief says the decisive multi-year readout cannot be complete before late this decade.

Floating Offshore Wind FR-V-23

Whether the moorings, connectors and dynamic cables of floating wind achieve a failure rate per component-year that lenders and insurers can price.

Pool and publish the component-level service record of the entire global floating fleet, with installation date, inspection history, failure mode, repair duration and lost production for every mooring line, anchor, dynamic cable, bend stiffener and buoyancy module. — measurement · unscheduled · Frontier

“It requires no new hardware, no new science and no new vessel. It requires operators to agree to disclose, which is why it does not exist.”

The brief says the data already exist in operators' hands and only disclosure is missing.

Wave and Tidal Energy FR-V-24

Whether tidal stream is actually on a cost-reduction path, or whether its contracted strike prices reflect a cost that does not fall.

The contracted UK tidal-stream fleet publishing, on a common basis, its achieved capacity factor, availability, operating cost per megawatt-hour, and every component recovery and replacement event with duration and cause. — measurement · mid · Frontier

“The United Kingdom’s ring-fenced contracts commit a small fleet of tidal-stream capacity to deliver across the second half of this decade at a known strike price.”

The brief ties the result to contracts already awarded, so the delivery dates exist even though the disclosure obligation does not.

VI — Climate & Planetary Engineering

Climate Engineering FR-VI-01

Whether the projected Sahel precipitation loss under stratospheric aerosol injection is a robust physical result or an artefact of model structure.

A GeoMIP-successor model intercomparison with common injection strategies at higher resolution, evaluated on monsoon dynamics and reported per model rather than as an ensemble mean. — experiment · unscheduled · Frontier

“First, a GeoMIP-successor model intercomparison with common injection strategies at higher resolution, evaluated on monsoon dynamics and reported per model rather than as an ensemble mean. It requires no release, no permit and no consent, and would settle whether the Sahel result is robust or model-structural.”

The decisive experiments are known, cheap by the standards of the subject, and unfunded; this is the one nearest to being runnable today.

Carbon Capture at Scale FR-VI-02

Whether geological storage can hold its permitted injection rate for years at a time, which sets the confidence interval on every carbon removal pathway.

A geological storage complex operating at its permitted injection rate for five consecutive years, with Northern Lights as the test and Gorgon as the counter-case. — demonstration · mid · Frontier

“Fourth, a geological storage complex operating at its permitted injection rate for five consecutive years. Gorgon is the counter-case and Northern Lights is the test. This is the single experiment that would most change the confidence interval on every removal pathway, because every one of them ends in a well.”

Desert Greening FR-VI-03

Whether the measured greening of the Sahel coincides with a gain or a loss in biodiversity on the same ground.

Co-measuring greenness and counterfactual biodiversity on the same plots in Nigeria and Senegal, with matched controls and NDVI or biomass and species richness taken from the same footprints in the same years. — measurement · unscheduled · Frontier

“One: co-measure greenness and counterfactual biodiversity on the same plots. Nigeria and Senegal, matched controls, NDVI or biomass and species richness from the same footprints in the same years. This resolves the sharpest contradiction in the subject and requires a field season, not a programme.”

The experiments are ordered by leverage per dollar and the first four are all cheap; this one needs a field season, not a programme.

Ocean Engineering FR-VI-04

Whether the 1 mm seabed blanketing limit written into draft deep-sea mining regulation is an achievable specification or an aspiration.

A full-scale collector trial at commercial throughput with plume instrumentation at 5, 50 and 500 kilometres, measuring sediment accumulation depth and suspended load against range and time over at least one seasonal cycle. — experiment · unscheduled · Established

“The measurement: sediment accumulation depth and suspended load as functions of range and time, over at least one seasonal cycle, with the collector operating at commercial throughput rather than at pre-prototype scale. The decision it closes: whether a 1 mm blanketing limit is a specification or an aspiration.”

Arctic Engineering FR-VI-05

Which mechanism actually drives the observed displacement of embankments built on ice-rich permafrost.

Running settlement rods, lateral displacement and inclinometer arrays and distributed temperature-sensing strings together for at least five years on an already-instrumented ice-rich permafrost embankment, and publishing an attribution of 80% or more of displacement to one mechanism. — measurement · unscheduled · Frontier

“Experiment one: settle the failure mode on ground that is already instrumented. Take an embankment on ice-rich permafrost — the Inuvik–Tuktoyaktuk corridor is the obvious candidate — and run vertical settlement rods, lateral displacement and inclinometer arrays, and distributed temperature-sensing strings together for at least five years.”

The instruments largely exist and the analysis does not, which makes this the cheapest high-value result available in the subject.

Weather Modification FR-VI-06

Whether seeded ice-nucleating particles persist downwind long enough for the design assumption under every cloud-seeding trial to hold.

A one-season persistence test instrumenting ice-nucleating-particle concentrations downwind of seeded storms at one day, one week and one month against matched unseeded control sites. — experiment · unscheduled · Frontier

“The cheapest high-value experiment available today is the persistence test, and nobody is running it. Instrument ice-nucleating-particle concentrations downwind of seeded storms at intervals of one day, one week and one month, against matched unseeded control sites, for one season.”

Nobody is running it; it is a direct replication of a 1988 result with instrumentation that did not exist then.

Atmospheric Management FR-VI-07

Whether iron-salt chlorine chemistry has a stable sign for net methane removal across the ambient atmospheric envelope.

Chamber measurement of chlorine yields extending the Fe(III)/sea-salt work across the NOx, humidity, ozone and SO2 parameter space, establishing where net methane removal changes sign. — experiment · unscheduled · Frontier

“Experiment one, and the cheapest high-value move available today: chamber chlorine yields across the full ambient envelope. Extend the Fe(III)/sea-salt chamber work to the NO x , humidity, ozone and SO 2 parameter space implied by the sign-dependence preprint, and establish where net methane removal changes sign.”

Laboratory scale, no permit, no release, no consent question; it would settle whether the field's leading candidate has a stable sign.

Water Infrastructure Megaprojects FR-VI-08

Whether a widely cited systematic review mislabelled desalinated-water and brine quantities, leaving global brine understated by about 1.5 times.

Obtaining Jones et al. (2019) and reading its two headline quantities against the brine-to-product ratio implied by the 2026 review's own number. — observation · unscheduled · Established

“The first required experiment is not an experiment. It is a library retrieval, and it should be done before anything else in this brief is treated as settled. Obtain Jones et al. (2019) and read its two headline quantities.”

A library retrieval, not a field programme, and it should be done before anything else in the brief is treated as settled.

Planetary Cooling Concepts FR-VI-09

Whether the cloud-cover-dominance result behind marine cloud brightening, and the climate sensitivity revision implied by it, survives independent replication.

Replicating the cloud-cover-dominance result on existing satellite archives with a different volcano or ship-track dataset and an attribution method other than the original machine-learning implementation. — observation · unscheduled · Frontier

“The cheapest high-value experiment in this brief requires no release, no permit and no new instrument. Replicate the cloud-cover-dominance result using existing satellite archives, a different volcano or ship-track dataset, and an attribution method that is not the original machine-learning implementation.”

Needs no release, no permit and no new instrument; nothing else in the category offers that ratio of information to cost.

Continental Irrigation Systems FR-VI-12

Whether a continental-scale water transfer perturbs precipitation and surface temperature enough to make the scheme climate engineering.

Running a NAWAPA-scale transfer in a modern earth-system model and reporting precipitation and temperature change against transferred volume for donor, recipient and teleconnected domains. — experiment · unscheduled · Frontier

“One. Run a NAWAPA-scale transfer in a modern earth-system model. Report change in precipitation and surface temperature as a function of transferred volume, for the donor and recipient domains and for the teleconnected regions. The established scaling result (Chen & Xie 2010) makes this a well-posed experiment; the continental case has never been run.”

The continental case has never been run, and if the feedback is large the scheme is climate engineering and should be governed as such.

Floating Cities FR-VI-13

Whether any jurisdiction allows a floating dwelling and its berth to be registered as a single financeable unit, and which statutes block it.

Attempting, in a cooperative jurisdiction, to register a security interest over an existing floating dwelling and its berth as one unit, and recording which office refuses and on what statutory ground. — experiment · unscheduled · Frontier

“Experiment 1: the registry filing test. Take an existing Dutch or Scandinavian floating dwelling and attempt, in a cooperative jurisdiction, to register a security interest over both the module and its berth as a single unit. Record precisely which office refuses, and on what statutory ground.”

The decisive experiments in this subject are legal and administrative, and they are cheap; the output is the list of statutory amendments nobody currently has.

Polar Development FR-VI-14

Whether pumping seawater onto Arctic ice can work as a global climate intervention rather than only as a regional ice-preservation measure.

Zampieri and Goessling's model of sea-ice-targeted geoengineering by seawater pumps, which returned a global annual-mean near-surface air temperature reduction of 0.02 K. — experiment · running · Established

“Link four has been tested directly and it failed. Zampieri and Goessling modelled sea-ice-targeted geoengineering by seawater pumps and found global annual-mean near-surface air temperature reduced by 0.02 K, against real regional Arctic cooling.”

Link four is a question about the climate system's response, and the answer is in.

Sustainable Megacities FR-VI-15

Whether tenure security is separable from and prior to physical upgrading in informal settlements.

A settlement-level four-arm trial - tenure security alone, physical upgrading alone, both, neither - measured on health, school retention, household investment and income over at least five years. — experiment · unscheduled · Frontier

“3. The upgrading trial that separates tenure from bricks. A settlement-level design with four arms — tenure security alone, physical upgrading alone, both, neither — measured on health, school retention, household investment and income over at least five years.”

The single most valuable experiment in the subject, and the only one of the six that is expensive and slow; the claim it tests has never been isolated.

Geoengineering Governance FR-VI-16

Whether a functioning treaty body can extend an existing assessment framework to cover a solar geoengineering technique.

The London Protocol intersessional correspondence group's October 2026 report on how existing instruments apply, with marine cloud brightening among the techniques considered for Annex 4 listing. — policy result · running · Established

“That is a live test of whether a functioning treaty body can extend a working framework to a solar geoengineering technique , and its outcome is the single most informative datum this subject will produce this decade.”

Already running whether anyone designs for it or not; the group reports to October 2026.

Biodiversity Restoration FR-VI-17

Whether the finding that only a small minority of taxa benefit from protected areas is a property of protected areas or a property of Finland.

Running the Santangeli design continentally: a counterfactual, multi-taxon, matched-site occupancy comparison over four decades, replicated across several biogeographic regions. — observation · unscheduled · Frontier

“One: run the Santangeli design continentally. A counterfactual, multi-taxon, matched-site occupancy comparison over four decades, replicated across several biogeographic regions rather than one country, would establish whether the Finnish result — a small minority of taxa benefiting, mostly through slower decline — is a property of protected areas or a property of Finland.”

It is the highest-value single study in the subject and it runs mostly on archives.

Coastal Defense Systems FR-VI-18

Whether a coastal authority anywhere has lawful power to withdraw protection, what standing a landowner has to stop it, and what compensation follows.

An audit of the statute book across twenty coastal jurisdictions establishing the power to withdraw protection, landowner standing to prevent it, and the compensation consequences. — observation · unscheduled · Frontier

“Six. Audit the statute book. For twenty coastal jurisdictions, establish whether an authority has a lawful power to withdraw protection, what standing a landowner has to prevent it, and what compensation follows.”

It is desk research, it is unglamorous, and it is the highest-value work in this brief.

Climate Migration Planning FR-VI-19

Whether the most-cited claims in climate-migration policy documents are supported by the sources cited for them.

A systematic citation audit of the fifty most-cited claims in climate-migration policy documents. — observation · unscheduled · Frontier

“A systematic citation audit of the fifty most-cited claims in climate-migration policy documents is a weekend of work and would probably be the highest-value output in the subject.”

A weekend of work; this cluster already produced two confirmed cases of a source being cited for something it does not say.

Planetary Stewardship FR-VI-20

Whether disagreement over national planetary-boundary allocations is normative rather than empirical.

Recomputing one national allocation of the published nitrogen ceiling under equal-per-capita, historical-responsibility, capability and grandfathering rules, propagating the ceiling's interval through each, and publishing the spread. — experiment · unscheduled · Frontier

“The predicted result is that the between-rule spread exceeds the within-rule uncertainty for most states, and that for some the two are comparable — which would establish, quantitatively, that the argument about allocation is normative rather than empirical.”

No retrieved source has run this.

Carbon Removal Verification FR-VI-21

Do the market's registries, applied to the same physical removal deployment, issue the same number of tonnes?

A registry round-robin: one instrumented removal deployment credited in parallel under Isometric, Puro.earth and Verra rules, with all measurements public and the spread in issued tonnes published as the result. — measurement · unscheduled · Frontier

“The decisive test is a registry round-robin: one instrumented removal deployment credited in parallel under the rules of Isometric, Puro.earth and Verra, with every measurement public and the spread in issued tonnes published as the result.”

Nobody has scheduled or funded it, and no registry has an incentive to volunteer; the brief calls it cheap by the standards of the field.

Compound Climate Hazards FR-VI-22

Whether dependence models fitted to the observed climate record predict joint hazard exceedances better than the independence assumption built into design standards and pricing.

A coordinated out-of-sample verification: fit the field's dependence models to the instrumental record up to a cutoff, freeze them, and score their predicted joint exceedances against the decades observed since, region by region and hazard pair by hazard pair, with independence as the null model. — measurement · unscheduled · Established

“Fit the field’s dependence models — the copula families, the conditional models, the model-ensemble dependence structures — to the instrumental record up to a cutoff, freeze them, and score their predicted joint exceedances against the decades already observed since, region by region and hazard pair by hazard pair, with independence as the null model.”

The brief states the test needs no new instrument and nobody has scheduled it; the data already exist.

Wildfire Systems FR-VI-23

Do fuel-treatment severity reductions measured mostly under moderate fire weather hold under the extreme fire weather that now produces most burned area?

A registered, prospective measurement protocol across wildfire encounters with recently treated forest in extreme seasons - pre-registered treatment polygons, standardized severity metrics, weather percentile at encounter, results published whether or not the treatment held. — natural experiment · running · Frontier

“The decisive test is already running: each extreme season drives wildfire into thousands of hectares of recently treated forest, and a registered, prospective measurement protocol across those encounters would settle whether severity reductions measured mostly under moderate weather survive the conditions that now do most of the burning.”

The brief calls this a natural experiment nobody has to build, only instrument; it decides whether the fuel-treatment lever scales with the climate signal or is capped by it.

Earth-System Digital Twins FR-VI-24

Does an operational Earth-system digital twin improve real local decisions, scored on outcomes, over the incumbent forecast chain?

A pre-registered paired trial in which matched decision units make real operational choices, one arm served by the twin chain and one by the incumbent products, with the arms scored on decision outcomes rather than forecast skill. — experiment · unscheduled · Frontier

“The decisive test is a pre-registered paired trial in which matched decision units make real operational choices, one arm served by the twin chain and one by the incumbent products, and the arms are scored on decision outcomes rather than on forecast skill.”

The brief states that nobody has scheduled or funded such a trial for any digital-twin programme.

Climate Overshoot and Lock-In FR-VI-25

Is the reversibility assumption underneath overshoot planning sound for the largest tipping element in the climate system?

Sustained measurement of the Atlantic meridional overturning circulation by the existing observing arrays, read together with the South Atlantic freshwater-transport fingerprint, to see whether a decline coherent across latitudes coincides with movement toward the model-identified tipping regime. — observation · running · Frontier

“If the overturning arrays show a decline that is coherent across latitudes and the fingerprint moves toward the model-identified tipping regime, the reversibility assumption underneath overshoot planning fails for the largest single element in the system, and it fails while the temperature is still rising.”

The brief says the arrays exist, the record is roughly two decades long, and the limiting factor is continuity of funding rather than instrumentation.

AI Weather Prediction FR-VI-26

Has machine learning replaced numerical weather prediction, or only its cheapest stage?

An end-to-end machine-learning forecast system initialised from raw observations, with no physics-based analysis anywhere in the chain, evaluated against independent observations on extreme-value metrics rather than against reanalysis. — demonstration · near · Frontier

“The decisive demonstration is an end-to-end machine-learning forecast system, initialised from raw observations with no physics-based analysis anywhere in the chain, evaluated against independent observations on extreme-value metrics.”

The brief says the components exist and the first end-to-end systems have been published at lower skill than operations, so the comparison needs no new instrument, only a protocol nobody has agreed.

Extreme Heat Survivability FR-VI-27

Does the housing stock of a heat-exposed city keep people alive through a multi-day heat event once the cooling stops, and by how wide a margin?

A stratified measurement of indoor temperature and humidity, with power-state metadata, across ordinary dwellings through a real multi-day heat event that includes a real power interruption. — measurement · unscheduled · Frontier

“The decisive result is a measured indoor-conditions dataset spanning a real multi-day heat event that includes a real power interruption, in a stratified sample of ordinary dwellings.”

Needs no new hardware and no new physics; nobody has funded it at the required scale in any heat-exposed city.

Blue Food Systems FR-VI-28

Can land-based grow-out of a high-value farmed fish be produced at a cost that competes with an open net pen, across complete production cycles rather than in a projection?

An audited, multi-cohort, full-cycle cost of production published from a commercial-scale land-based grow-out facility. — measurement · unscheduled · Frontier

“The decisive result is an audited, multi-cohort, full-cycle cost of production from a commercial-scale land-based grow-out facility.”

Nothing new needs to be built to produce it; what is missing is a disclosure requirement rather than a facility.

VII — Civilization-Scale Infrastructure

Continental Transportation Systems FR-VII-01

Is converting a corridor's track gauge cheaper than living with the break of gauge at the volumes that corridor actually carries?

Measure the per-container cost of a gauge break on a defined China-EU or Ukraine-EU corridor as dwell time, transhipment handling cost and reliability penalty, reported as a distribution, and compare it against the amortised cost of conversion at the corridor's actual volume. — measurement · unscheduled · Speculative

“That comparison is computable with data that already exists, it decides whether the Ukrainian conversion programme is rational, and nobody in the Anglophone literature appears to have computed it. It is the most tractable open question in this brief.”

The brief says the comparison decides whether the Ukrainian conversion programme is rational.

Northern Development Corridors FR-VII-02

Does corridor capital deliver more cost-of-living benefit per northern household than the same money spent on air service, local energy or local food?

Compare corridor capital per household served, amortised over asset life with the maintenance path included, against the cost-of-living reduction the same capital buys through runway extension and air-service reliability, local energy, or local food production. — measurement · unscheduled · Established

“The decisive measurement is a comparison nobody runs. Take corridor capital per household served, amortised over the asset life with the maintenance path included, and compare the cost-of-living reduction it delivers against the same capital spent on runway extension and air-service reliability, on local energy, and on local food production.”

The brief says the inputs exist: the airport-infrastructure and food-cost work supplies one arm, the socio-technical capacity work another.

Smart Cities FR-VII-03

Did smart-city programmes improve the outcomes they were sold on, measured against comparable cities that were shortlisted but not selected?

Run the difference-in-differences already sitting in India's Smart Cities Mission: 100 cities with staggered implementation against shortlisted-but-not-selected controls, on outcome metrics the cities were reporting before the programme began. — natural experiment · unscheduled · Established

“The cheapest decisive study in the subject already has its comparison group and has not been run. India's Mission gave 100 cities staggered implementation across ten years, and it shortlisted cities that were not ultimately selected.”

It requires no deployment, no new instrumentation and no vendor cooperation; it requires someone to fund an analyst.

High Speed Transit Networks FR-VII-04

How much of the spread in high-speed rail cost per kilometre is terrain, standards, mitigation and procurement rather than national competence?

Decompose per-kilometre outturn cost for at least five national high-speed programmes into terrain-forced, standard-forced, mitigation-forced and procurement-forced components, on one explicitly stated scope boundary and a single price year. — measurement · unscheduled · Established

“One: the four-way cost decomposition. Take at least five national high-speed programmes and separate per-kilometre outturn into terrain-forced, standard-forced, mitigation-forced and procurement-forced components, on a single scope boundary that states explicitly whether it includes rolling stock, stations, depots, land, electrification and interest during construction, at a single price year.”

The Italian three-stage cost study is the closest existing precedent and covers one country.

Underground Cities FR-VII-06

What a subsurface cadastre actually costs to build, a price every downstream underground planning instrument has assumed without testing.

Building one subsurface cadastre for a mid-sized city from geophysical proxy indicators, cheap monocular or photogrammetric capture of accessible voids and existing utility records, and publishing coverage, accuracy and total cost per square kilometre. — demonstration · unscheduled · Frontier

“The deliverable is not the map; it is the price of the map. Every downstream instrument has been blocked for thirty-five years on the presumption that this is expensive, and the presumption is untested.”

Arcologies FR-VII-07

At what population, vertical extent and compartment count performance-based fire engineering stops being able to demonstrate a margin against tenability limits.

Running the current performance-based design toolchain against a family of hypothetical occupancies scaling population, vertical extent and compartment count, recording where a demonstrable margin with defensible input uncertainty is lost. — experiment · unscheduled · Frontier

“1. Find the population threshold at which performance-based fire engineering stops being demonstrable. This is first because it bounds everything else and because nobody has published it.”

The output is a curve and a knee, not a number.

Automated Construction Systems FR-VII-08

Whether construction's measured productivity stagnation is substantially an artefact of deflator-based output measurement.

Building a hedonic construction output index adjusted for regulated performance attributes, back-cast thirty years across two national statistical systems, and comparing its growth rate against the current deflator-based series. — measurement · unscheduled · Frontier

“The decisive comparison is the index's growth rate against the current deflator-based series.”

Experiment two binds the printing branch; experiment one binds the case for the programme as a whole.

Spaceports FR-VII-10

Is launch cadence limited by ground range and airspace institutions rather than by vehicles, and what is a range window actually worth?

Publish a crossover of launch cadence against aviation delay cost on a dense route network, the measurement that pricing or allocating range windows and airspace closures requires. — measurement · unscheduled · Speculative

“This is the binding link and it is institutional. The measurement is a published crossover of launch cadence against aviation delay cost on a dense route network. No such mechanism exists in any jurisdiction and no such study has been found.”

The brief says the decisive measurements here are blocked by publication rather than by difficulty.

Autonomous Supply Chains FR-VII-12

Whether warehouse robotics actually raises picks per labour hour at constant SKU mix and order profile.

A before-and-after measurement of picks per labour hour on the same SKU mix and order profile, with a control facility. — measurement · unscheduled · Frontier

“One: picks per labour hour, before and after, same SKU mix and order profile, with a control facility. This is the measurement the entire warehouse-robotics case rests on, it is trivially available to any operator, and it is essentially unpublished.”

Its absence is informative: an operator holding a favourable number would publish it.

Future Ports and Shipping FR-VII-13

Whether terminal automation's productivity gains come from the technology itself or from implementation, integration and training.

Computing between-terminal productivity variance among automated terminals against manned ones; higher variance among the automated would show implementation rather than technology is the causal factor. — measurement · unscheduled · Speculative

“The highest-value unrun experiment in this subject is cheap, and it is a variance test. If terminal automation's gains are real but conditional on integration and training — which is what the Mediterranean study's own hedge says — then automated terminals should show higher between-terminal variance in productivity than manned ones, not a uniform advantage.”

Nobody appears to have run this, the data mostly exists, and it would settle a decade of marketing.

Energy Corridors FR-VII-14

Whether a cross-border capacity-allocation and cost-recovery regime holds when the exporting system is itself short of power.

A rule set compensating jurisdictions crossed but not served by a corridor and preserving capacity rights against a national curtailment override, tested on revealed behaviour in a correlated stress event. — policy result · unscheduled · Speculative

“A rule set under which a jurisdiction crossed by, but not terminating, a corridor is compensated, and under which capacity rights survive a national curtailment override. The measurement is revealed behaviour in a correlated stress event: does the corridor deliver across a border when the exporting system is itself short?”

Link 4 binds hardest and Link 1 binds first.

Intercontinental Rail Systems FR-VII-16

Does existing traffic across each candidate strait come anywhere near the throughput a fixed intercontinental rail link would have to carry?

Assemble, for each of the three crossings, current annual tonnage and passenger movements by ferry, short-sea shipping, air and land detour, split by direction and commodity value class, and express existing flow as a ratio to the throughput a link would need. — measurement · unscheduled · Speculative

“The output is a single ratio per crossing: existing flow against the throughput a link would need. Nobody has published it, and it decides everything downstream.”

The brief says the data exist in customs, port and civil-aviation statistics.

Future Housing Systems FR-VII-17

What fraction of an industrialised building cost reduction reaches the sale price rather than the land price, in constrained against elastic markets.

Delivering an identical industrialised product into a high-elasticity and a low-elasticity market, matched on specification and boundary, and measuring what share of the hard-cost reduction appears in sale price versus land price. — experiment · unscheduled · Frontier

“The second experiment is the pass-through test, and it is the one that decides the programme. Deliver an identical industrialised product into a high-elasticity and a low-elasticity market, matched on specification and boundary, and measure what fraction of the hard-cost reduction appears in the sale price versus the land price.”

The output is a single pass-through coefficient, and if it is near zero in constrained markets the field's affordability claim is jurisdiction-specific.

Megaproject Governance FR-VII-19

Whether megaproject cost overruns are forecasting error or strategic misrepresentation.

A two-by-two of competitive against non-competitive approval crossed with externally set against promoter-set uplift, with outturn as the dependent variable and difference-in-differences as the estimator. — natural experiment · unscheduled · Speculative

“Three, and the highest-leverage unrun study in the field: the discriminating test between error and misrepresentation. A two-by-two: competitive against non-competitive approval, crossed with externally set against promoter-set uplift, with outturn as the dependent variable and difference-in-differences as the estimator.”

The variation already exists in the world and nobody has assembled it; there is no scientific barrier, only the need for several governments to hand over comparable data about their own decisions.

Infrastructure Resilience FR-VII-20

Whether restoration capacity rather than failure probability is the binding resilience variable, which would mean much hardening expenditure targets the wrong term.

Publishing the time-to-restore distribution for one regulator's jurisdiction over one decade, with damage, event class and resources brought to bear recorded, then running a variance decomposition against physical damage. — measurement · unscheduled · Frontier

“The decisive analysis is a variance decomposition: does time-to-restore vary more across events with similar physical damage than physical damage itself varies?”

The cost of this study is a data-sharing agreement and an analyst. Nobody has done it.

Robotics in Infrastructure FR-VII-21

Is construction's flat measured productivity a real stagnation, or an artefact of price and quality measurement that hides genuine robotic gains?

Recompute the standardised-product output-per-worker-hour series with an explicit hedonic adjustment for code-driven quality content, testing directly whether the physical productivity measures are themselves contaminated. — measurement · unscheduled · Speculative

“Two: the same series with an explicit hedonic adjustment for code-driven quality content, which is the direct test of whether physical measures are themselves contaminated and therefore the sharpest available challenge to Goolsbee and Syverson's strongest evidence.”

The brief sets no date for it; of the companion benchmark correction it says nobody appears to have run it.

Industrial Ecology FR-VII-23

Is waste law the binding constraint on industrial symbiosis, so that a reclassification pipeline rather than a park-building programme is the instrument that works?

Run a difference-in-differences on end-of-waste reclassification: a by-product stream reclassified out of waste law on a known date against a comparable non-reclassified control, measuring exchange formation counts before and after with material price controlled. — natural experiment · unscheduled · Frontier

“The decisive experiment is a difference-in-differences on end-of-waste reclassification. Take a specific by-product stream reclassified out of waste law on a known date, take a comparable non-reclassified stream as control, and measure exchange formation counts before and after, controlling for the material's price.”

The brief calls it tractable, unpublished, and the highest-value empirical opportunity in the subject.

Circular Infrastructure Systems FR-VII-24

What does it cost to certify reclaimed structural steel, and what fraction of a demolition batch actually passes?

Put a defined batch of demolition steel sections through a published test protocol with a named certifying body issuing or refusing a document, then publish the cost per tonne of establishing conformity and the fraction of the batch that passes. — experiment · unscheduled · Frontier

“The first experiment is a certification pilot and it is embarrassingly cheap. Take a defined batch of structural steel sections from one demolition.”

The brief says both numbers currently do not exist anywhere it could reach.

Civilization Resilience Planning FR-VII-25

Whether a collapsed industrial society could bootstrap back with the energy return, metallurgical sequence and population it would actually have.

A rigorous energy-and-materials analysis of industrial bootstrap pathways under a salvage endowment, closing the energy return on accessible surface coal and shallow oil under pre-industrial extraction, the metallurgy sequence, and the minimum population and knowledge base. — experiment · unscheduled · Speculative

“The first required experiment is a paper, and this brief can specify its contents precisely enough that somebody could go and write it. A rigorous energy-and-materials analysis of industrial bootstrap pathways under a salvage endowment would have to contain three things.”

Every one of those three is a determinate calculation. None of them is at the frontier of any discipline. The paper does not exist.

The Physical Stack of AI FR-VII-26

Which layer of the physical stack — packaging, machinery, interconnection, or the permit — actually binds the AI buildout, and for how long.

Whether the announced 2028 cohort of gigawatt campuses — five Stargate-class sites totalling roughly 8 GW with fourth-quarter-2028 targets, plus the Crane nuclear restart — reaches powered operation on schedule, with the pattern of slips localising the binding layer. — natural experiment · running · Frontier

“The decisive test is already running. The announced 2028 cohort — five Stargate-class sites carrying roughly 8 GW of stated capacity against fourth-quarter-2028 targets, the Crane nuclear restart, and the giant load queues behind them — will either reach powered operation on schedule or slip, and the pattern of the slips will localise the binding layer of the stack in a way no forecast can.”

Already underway; needs no new instrument beyond satellite imagery, permit dockets and utility filings — the census Epoch AI already runs.

Memory-Safe Computing Transition FR-VII-27

Can automated translation convert large legacy C codebases into maintainable, semantics-preserving Rust at a cost that makes the memory-unsafe legacy tail tractable?

The independent TRACTOR evaluation: MIT Lincoln Laboratory scores automated C-to-Rust translation against escalating benchmark batteries every six months; success on realistic million-line codebases would turn billions of unsafe lines into a finite engineering bill, while a plateau at transliterated output on simplified C would leave the transition a new-code-only story. — demonstration · running · Frontier

“The decisive test is already running: the TRACTOR evaluation of whether automated translation can turn real C codebases into Rust that maintainers would accept without a full manual rewrite. MIT Lincoln Laboratory scores performer output against benchmark batteries released every six months, and a Round 1 evaluation report has been published.”

Already underway: benchmark batteries escalate every six months and a Round 1 evaluation report is published; current batteries cover deliberately simplified C.

Post-Quantum Migration FR-VII-28

Do the structured-lattice assumptions beneath ML-KEM and ML-DSA withstand sustained cryptanalysis while the migration that bets almost everything on them completes?

A practical attack on the module-lattice problems underlying ML-KEM and ML-DSA, of the kind that broke SIKE in 2022; the brief treats the worldwide cryptanalytic effort against these assumptions as its decisive experiment, already running, with even sustained security-estimate erosion short of a break forcing parameter escalation. — experiment · running · Frontier

“The single result most able to change this assessment is a practical attack on the module-lattice problems beneath ML-KEM and ML-DSA. The 2022 break of SIKE showed what such an event looks like: a scheme that had survived years of public review fell to a classical attack running in about an hour on a single core.”

Already running and needing no scheduling; the brief notes SLH-DSA and HQC exist as pre-positioned hedges against its worst outcome.

Quantum Networks FR-VII-29

Whether a quantum repeater can outperform direct photon transmission over deployed fibre, turning the quantum internet from architecture into demonstrated infrastructure.

A memory-enhanced repeater link over deployed fibre, between independently operated nodes in different buildings, delivering entangled pairs at a higher rate than direct transmission through the same fibre allows. — demonstration · mid · Frontier

“The decisive demonstration is a memory-enhanced repeater link over deployed fibre, between independently operated nodes in different buildings, that delivers entangled pairs at a higher rate than direct transmission through the same fibre would allow. Nobody has done it; every credible roadmap treats it as the gate through which the field must pass.”

The brief says every component has been shown separately, groups in Hefei, Delft and Boston are explicitly building toward the assembled version, and it is plausibly a this-decade result.

Extreme-Environment Test Infrastructure FR-VII-30

Does IFMIF-DONES come online and produce the first fusion-spectrum irradiation data, converting fusion's materials case from extrapolation into measurement?

First beam on target at IFMIF-DONES followed by the first published fusion-spectrum irradiation data for a reduced-activation steel, which the brief treats as the single result most able to change its assessment. — demonstration · long · Frontier

“The decisive result this brief tracks is first beam on target at IFMIF-DONES followed by the first published fusion-spectrum irradiation data for a reduced-activation steel. Every fusion first-wall schedule on the map extrapolates across a spectral gap that only this facility is funded to close.”

The programme's stated window for first irradiation campaigns is the early-to-mid 2030s; the brief says nothing about that window falls within this decade.

Additive Manufacturing Qualification FR-VII-31

Whether in-situ monitoring can demonstrate regulator-grade probability of detection for critical defects across machines, making certification of fracture-critical additive parts repeatable rather than bespoke.

A blind, multi-site round robin printing nominally identical fracture-critical builds on at least five machines, with each site's in-situ monitoring verdicts locked in escrow before inspection, then computed tomography and fatigue testing to failure on every part, published as probability-of-detection curves against defect size. — experiment · unscheduled · Frontier

“The decisive test is a blind, multi-site round robin: nominally identical fracture-critical builds on at least five machines, with in-situ monitoring verdicts locked before inspection, followed by computed tomography and fatigue testing to failure on every part, published as probability-of-detection curves against defect size.”

No agency or consortium has funded the full blind version; NIST's AM Bench series runs the adjacent model-calibration half but not the probability-of-detection half on fracture-critical hardware.

Critical Dependency Atlas FR-VII-32

Does an audited cross-sector dependency inventory predict the failure set a real multi-sector event actually realises?

Build an audited power-water-gas-communications dependency inventory for one mid-sized region, register in advance the failure set it predicts for a defined hazard, and let the region's next severe event adjudicate the map. — natural experiment · unscheduled · Frontier

“The decisive test is an audited cross-sector dependency inventory for one mid-sized region, scored against the next real event.”

The brief says nothing about this requires new science; nobody has commissioned it, and the sibling briefs' shared finding is that everything downstream is conditioned on it.

Digital Chokepoints FR-VII-33

Whether the redundancy built into the world's submarine cable network protects anything, which turns entirely on how long faults actually take to repair.

Publication by the maintenance-zone consortia of the time-to-restore distribution for submarine cable faults, broken down by cause, sea area and permitting regime. — measurement · unscheduled · Frontier

“Every claim made about cable resilience is a claim about that distribution: whether a second cable helps depends on how long the first one stays down, and whether a repair fleet is adequate depends on the tail rather than the mean.”

The brief says the operators log every fault, mobilisation and permit delay as a matter of contract, that nobody has published the distribution and no regulator has required it, and that the result needs only a disclosure rule.

Next-Generation Networks FR-VII-34

Can the candidate 6G upper mid-band be covered from the tower grid that already exists, or does it require a second national build-out?

An independent, co-sited field comparison of uplink coverage at 7 GHz against 3.5 GHz on the same towers, using commercial handset transmit power and a published, repeatable methodology. — measurement · near · Frontier

“The decisive measurement is an independent, co-sited comparison of uplink coverage at 7 GHz against 3.5 GHz on the same towers, with commercial handset transmit power and a published methodology.”

The trial equipment exists; no party without a commercial interest has published the result, and the brief ties the useful window to the WRC-27 allocation.

Manufacturing Digital Twins FR-VII-35

Does installing a coupled model of a production system change operational outcomes, once the twin is separated from the sensors, processes and management attention that arrive with it?

A stepped-wedge deployment across several comparable lines or sites: randomise the order in which the twin is switched on, register the primary outcome before the first switch -- unplanned downtime hours, first-pass yield, energy per unit -- and publish the result whichever way it falls. — experiment · unscheduled · Frontier

“The decisive experiment is a stepped-wedge deployment with pre-registered outcomes, and no operator has published one.”

The design is standard in health services research and needs no new technology; the brief says the obligation sits with the operator, who holds the counterfactual, and that no operator has published one.

Zero-Carbon Aviation FR-VII-36

Can the largest single term in aviation's radiative forcing, contrail cirrus, be reduced at network scale by routing today's aircraft differently, and is the extra fuel burnt worth the forcing avoided?

A full-year trial across one airline's entire operation in which eligible flights are randomised into contrail-avoidance routing, contrail formation is verified from satellite rather than from the forecast that generated the decision, and the extra fuel burnt is accounted against the forcing avoided. — experiment · unscheduled · Frontier

“The decisive experiment is a network-scale, independently verified contrail avoidance trial, and nobody has funded one.”

The trial needs no new aircraft, no new fuel and no new infrastructure; the brief says everything else in the subject is slow and this is not.

Advanced Air Mobility FR-VII-37

Whether a certificated powered-lift fleet in scheduled service achieves the dispatch reliability, operating cost, vertiport throughput and community noise outcome the field has projected.

One full year of scheduled revenue operations by a type-certificated powered-lift fleet, publishing dispatch reliability against schedule, direct operating cost per block hour with pack amortisation separated, achieved movements per hour at the busiest vertiport, and noise complaints per hundred movements. — measurement · unscheduled · Frontier

“Every contested claim in the field resolves against that dataset and none of them resolves without it.”

The brief says no such year has been flown anywhere as of its cut-off.

VIII — Governance & Institutions

Future Democracies FR-VIII-01

Whether the recommendations of citizens' assemblies survive into legislation, and at what rate.

Independently track every recommendation a deliberative body makes to its legislative fate, done for ten assemblies across five countries under a published and contested coding scheme. — experiment · unscheduled · Established

“The decisive experiment is cheap, obvious and almost never run: independently track every recommendation from a body to its legislative fate.”

Done properly once, for France; Brussels has the mechanism and has not used it, and no technical or financial obstacle attaches.

AI-Assisted Governance FR-VIII-02

Whether AI assistance raises administrative output, rather than the self-reported minutes the field currently measures.

Instrument a caseload rather than a diary: cases cleared per week, decisions reversed on appeal, backlog age and error rate at quality-control sample. — measurement · unscheduled · Established

“Measure output, not minutes. Instrument a caseload rather than a diary: cases cleared per week, decisions reversed on appeal, backlog age, error rate at quality-control sample. This is the study that does not exist, and constructing it requires no new method — every one of those numbers is already produced by the administrations in question for other purposes.”

Requires no new method: every one of those numbers is already produced by the administrations in question for other purposes.

Digital Constitutional Systems FR-VIII-03

Whether a statute encoded as machine-readable rules is reproducible across independent competent coding teams.

An inter-coder replication: run independent teams across statutes of different drafting styles and measure the residual divergence in the encodings they produce. — experiment · unscheduled · Established

“The cheapest decisive experiment is the inter-coder replication, and it would settle a constitutional argument for the price of a research assistant. Witt and colleagues showed divergent interpretive choices surviving an agreed vocabulary over two weeks on one statute. Run the same design with independent teams across statutes of different drafting styles and measure the residual divergence.”

Costs the price of a research assistant; if a coded encoding is not reproducible, the authoritative-coded-law proposal is finished on the evidence.

Institutional Design FR-VIII-04

How many of the eleven commons design principles survive once measured coder disagreement is taken out of the headline result.

Re-estimate the commons corpus test under the measured coder-disagreement distribution and report how many of the eleven principles survive. — experiment · unscheduled · Established

“The cheapest decisive experiment in the subject has been specified by the people best placed to run it and has not been run. Re-estimate the commons corpus test under the measured coder-disagreement distribution and report how many of the eleven principles survive.”

Requires no new fieldwork; both papers are in the same 2016 journal issue with the coding teams and their reliability scores published.

Scientific Governance Models FR-VIII-06

Whether partial lottery allocation of research funding produces different outcomes from panel ranking.

Analyse the partial-randomisation arms already executing at funders since 2013, where assignment above the quality threshold is random by construction, against the outcome data accumulating in the funders' own reporting systems. — experiment · running · Established

“The highest-value experiment in this subject has already been run and simply not analysed. At every funder using partial randomisation, assignment above the quality threshold is random by construction.”

Requires no new grant, no new consent and no new instrument; it requires a funder to publish it.

Future Legal Systems FR-VIII-07

Whether algorithmic risk assessment in courts operates as decision support or as the decision.

Publish judicial override rates and outcomes by defendant group for deployed risk-assessment tools. — measurement · unscheduled · Established

“The cheapest decisive experiment is to publish override rates and outcomes by defendant group. Two of the field studies underpinning this brief reconstructed judicial compliance indirectly from administrative data — that is how the 90%-modelled against 29%-actual figure for immediate non-financial release was obtained. No jurisdiction publishes it directly.”

The number already exists inside case-management systems; no jurisdiction publishes it directly, and its absence after a decade of deployment is a choice.

Public Policy Foresight FR-VIII-08

Whether published national risk assessments were right, which nobody has ever scored.

Retrospective scoring of the eighteen years of ranked, banded, dated assessments already published in the UK national risk register. — experiment · unscheduled · Established

“The cheapest and most valuable experiment in this entire subject is retrospective scoring of documents that already exist. The UK national risk register has been published since 2008 with ranked, banded, dated assessments.”

Eighteen years of falsifiable material sits in the public domain and no scoring of it has ever been published.

Long-Term Institutions FR-VIII-09

Whether future-generations bodies have any measurable effect on legislation.

Publish the counts: how many legislative opinions a future-generations body issued, how many constitutional referrals it proposed, how many were made and how many succeeded. — measurement · unscheduled · Established

“The most valuable single measurement is also the cheapest: publish the counts. Hungary's ombudsman office could state, in a paragraph, how many legislative opinions it issued, how many Constitutional Court referrals it proposed, how many were made, and how many succeeded. Every future-generations body in the world could do the same.”

Until one body publishes the counts, the field's impact question is unanswerable in principle rather than in practice.

Future Federalism FR-VIII-10

Whether reforming a fiscal equalisation formula changes subnational behaviour and outcomes.

Pre-register the evaluation of an equalisation reform before the formula changes, using the announced Canadian, Australian and German reform dates. — natural experiment · near · Established

“The cheapest high-value experiment is to pre-register the evaluation of an equalisation reform before the formula changes. The dates are announced years ahead and the outcome data exist: Canada's programme renews 31 March 2029 , Australia's Productivity Commission reports finally on 31 December 2026 , Germany reformed in 2020 with parameters published in advance.”

Three staggered natural experiments in rich federations with audited fiscal data, and no announced design attached to any of them.

Global Cooperation Models FR-VIII-11

Whether the Montreal Protocol caused the ozone outcome, or coincided with substitution that commercial incentives would have driven anyway.

A selection-corrected re-analysis of the Montreal Protocol itself, using staggered ratification dates and an outcome series measured by agencies outside the regime. — experiment · unscheduled · Established

“The experiment this field most needs is a selection-corrected re-analysis of the Montreal Protocol itself — the one case with clean independent outcome data and near-universal membership. Nobody has separated the treaty's effect from the commercial incentives that would have driven substitution anyway, and until somebody does, the field's best evidence is also its least examined.”

The design is available and the data exist; nobody has separated the treaty's effect from the commercial incentives.

Existential Risk Governance FR-VIII-12

Whether existential-risk regimes prevent the outcomes they target or merely select the states that were never going to produce them.

A prevention institution with a control group: exploit the staggered adoption of the Additional Protocol across states, or the staggered establishment of national AI evaluation bodies, as natural experiments with real variation in timing. — natural experiment · unscheduled · Frontier

“Sixth, and the only design that would settle anything: a prevention institution with a control group. Nothing in this brief has one. The nearest available candidates are staggered adoption of the Additional Protocol across states, and the staggered establishment of national AI evaluation bodies across countries with otherwise comparable technology sectors.”

Both are natural experiments with real variation in timing, and neither appears to have been exploited.

Technocracy and Democracy FR-VIII-13

Which of five mutually inconsistent results on central bank independence and inflation is an artefact of specification.

A pre-registered re-analysis of the five conflicting independence-inflation results on a common sample, varying the index choice, the sample split, the estimator and the treatment of initial inflation systematically. — measurement · unscheduled · Established

“A pre-registered re-analysis on a common sample, with the index choice, the sample split, the estimator and the treatment of initial inflation varied systematically, would settle which of the five is an artefact of specification. The data are public. Nobody has done it.”

The data are public and nobody has done it.

Future Civil Services FR-VIII-14

Whether AI tools raise civil-service throughput, measured as cases cleared rather than minutes reported.

Add one more arm to the existing departmental evaluations: random assignment of licences within a single processing function, with cases cleared per week as the dependent variable. — experiment · unscheduled · Established

“The cheapest decisive experiment has already been designed twice and simply needs an output measure attached. Three United Kingdom departments evaluated the same product in the same quarter with three designs.”

The administrative data already exist, because processing functions count their own throughput for other reasons; nobody has run it.

Digital Citizenship FR-VIII-15

Whether the measured leakage reduction belongs to biometric authentication or to the removal of the payment intermediary.

A four-arm trial on one welfare programme in one state, randomising the channel change and the authentication layer separately. — experiment · unscheduled · Established

“The highest-value experiment is the decomposition the two Indian trials came within one design choice of running. Randomise the channel change and the authentication layer separately — four arms: unchanged channel with and without biometric authentication, disintermediated channel with and without — on one welfare programme in one state.”

Cheap, reuses instruments both papers validated, and no announced design is attached to it.

Civic Technology FR-VIII-16

What signature threshold a civic participation platform should set, and what that threshold does to proposals enacted and to later participation.

Randomise the signature threshold across comparable municipalities and measure proposals cleared, proposals enacted and subsequent participation. — experiment · unscheduled · Established

“The obvious experiment is to randomise the threshold. These platforms are software and the signature requirement is a configuration value. Varying it across comparable municipalities and measuring proposals cleared, proposals enacted and subsequent participation would answer the field's central design question within a single budget cycle. It has never been done, and there is no technical or ethical barrier to doing it.”

Scientific Advisory Institutions FR-VIII-17

Whether formal scientific assessments change the decisions they are commissioned to inform.

Vary the advice and hold the decision-maker constant: a matched comparison of technical decisions taken with and without a formal assessment, controlling for salience and contestedness. — experiment · unscheduled · Frontier

“The experiment nobody has run is still the obvious one: vary the advice, hold the decision-maker constant. Legislatures and executives make hundreds of technical decisions a year and commission formal assessments for a small, non-random subset.”

It would be the first outcome evidence in the field; thirty-eight years on, the outcome literature remains empty.

Distributed Governance FR-VIII-18

Whether token-governed organisations satisfy the commons design principles, and whether satisfying them predicts survival.

Apply the eleven commons design principles as a scored checklist to token-governed organisations and test whether configural satisfaction predicts survival, as it does for irrigation systems and fisheries. — experiment · unscheduled · Frontier

“An experiment that would settle the central claim and has not been run: apply the eleven commons design principles as a scored checklist to token-governed organisations and test whether configural satisfaction predicts survival, as it does for irrigation systems and fisheries. The coding instrument exists and has been validated on 69 cases; nobody has pointed it at this domain.”

The coding instrument exists and has been validated on 69 cases; nobody has pointed it at this domain.

Future Public Administration FR-VIII-19

Whether making evaluation a condition of funding produces the outcome evidence public administration currently lacks.

Make an adequate evaluation plan a gate on project funding rather than a statistic government reports about itself. — policy result · unscheduled · Established

“The cheapest high-value intervention is to make evaluation a condition of funding rather than a nice-to-have. The measurement already exists: 34% of the portfolio has an adequate plan, 66% does not, and 55% of the shortfall group produced no evidence of a plan at all.”

The measurement already exists and is published annually by government about itself; making it a gate would generate the evidence within one project cycle.

Civilizational Planning FR-VIII-20

Whether a statutory future-generations duty changes measured outcomes against a comparable jurisdiction without one.

A difference-in-differences comparison of Wales, which legislated a statutory future-generations duty in 2015, against the rest of the United Kingdom, which did not. — natural experiment · unscheduled · Established

“The single highest-value experiment in this subject is already set up and nobody has run it. Wales legislated a statutory future-generations duty in 2015 and the rest of the United Kingdom did not. That is a natural experiment with a control group and roughly eleven years of indicator data collected on a comparable statistical basis. At the institutional level it remains unrun.”

Requires no new instrument, no new survey and no cooperation from the body being evaluated, only outcome series that already exist.

Content Authenticity Infrastructure FR-VIII-21

Whether legal compulsion can raise the share of AI-generated media that reaches users carrying an intact machine-readable mark above the pre-mandate 35 to 45 percent platform baseline.

The three-jurisdiction natural experiment formed by China's labeling measures (September 2025), EU AI Act Article 50 (August 2026) and California SB 942 (August 2026): platform transparency reports by the end of 2027 show either the labelled share of AI-generated media rising decisively above the 35 to 45 percent baseline, or staying flat while stripping and open-weight generation absorb the mandate. — policy result · running · Frontier

“The decisive test is already running: three mandatory-marking regimes came into force within twelve months of each other, and by the end of 2027 platform transparency reports will show whether the labelled share of AI-generated media rises decisively above the 35 to 45 percent baseline or stays flat while stripping and open-weight generation absorb the mandate.”

Already running; the brief says a clear rise would overturn its central scepticism, while a flat line would confirm the binding constraint was physical, not legal.

AI-Biology Governance FR-VIII-22

Whether the synthesis screening layer actually stops live orders, as opposed to scoring well on a curated benchmark.

A standing, blinded order-placement audit of the whole synthesis provider population: benign test orders submitted under varied customer identities at unannounced intervals, scored on whether each order was stopped, queried or shipped, with per-provider results published and an adjudication step separating the explanations the June 2025 twelve-order exercise could not distinguish. — experiment · unscheduled · Frontier

“The decisive experiment is a standing, blinded order-placement audit of the whole provider population, with per-provider results published.”

Needs no new science, no new hardware and no new legal theory; the brief says no body with a mandate to place the orders and publish the results exists or has been funded.

Military AI and Strategic Stability FR-VIII-23

Whether the Seventh CCW Review Conference in November 2026 adopts a negotiating mandate for a binding autonomous-weapons instrument, settling if consensus arms control can still bind a militarily significant technology before mass adoption.

The November 2026 CCW Review Conference decision: with a negotiating majority above 70 states and a 156-vote General Assembly resolution behind it, either it adopts a mandate to negotiate a legally binding instrument on autonomous weapons, or it lets a decade of preparatory work lapse into another mandate cycle; the brief treats either outcome as settling the live institutional question. — policy result · near · Frontier

“The decisive near-term result is institutional rather than technical: whether the Seventh CCW Review Conference in November 2026 adopts a negotiating mandate for a legally binding instrument on autonomous weapons, or lets a decade of preparatory work lapse into another mandate cycle.”

Scheduled and imminent per the brief: the final GGE session concluded 4 September 2026 and the Review Conference convenes in November 2026.

Information Integrity FR-VIII-24

Whether any single information-integrity intervention produces a durable downstream civic effect rather than an immediate change in sharing behaviour.

One randomised treatment assigned at the account level, with exposure, sharing, belief and a pre-registered downstream civic behaviour measured on the same subjects and reported at one week, one month and six months. — experiment · unscheduled · Frontier

“The decisive experiment is a platform-scale field trial that carries one intervention all the way through the chain and reports durability at six months.”

Needs no new method and no new science; it needs platform cooperation, which the European access regime can compel and no other instrument can.

Standards as Governance FR-VIII-25

Can delegated private standardisation supply the operative content of frontier-technology regulation on a legislated schedule?

Whether European harmonised standards for high-risk artificial intelligence are delivered and cited in the Official Journal in time for the December 2027 obligations, or the deadline is moved a second time. — natural experiment · near · Frontier

“No new institution is required to run this experiment and no one can stop it.”

Already running against a published date; the Union has already once amended the statute rather than enforce it without the standards.

Research Security FR-VIII-26

Do national research-security regimes reduce meaningful risk, and at what cost in collaboration, delay and exclusion?

A difference-in-differences analysis of the staggered introduction of field-specific research-security lists, comparing in-scope and out-of-scope research areas before and after each date of effect on grant application volumes, partner composition, international co-authorship and time-to-award, with cross-country variation separating policy effect from global trend. — natural experiment · running · Frontier

“The data already exist in funder administrative records and in bibliometric databases, no new instrument is required, and the cost is an analyst-year.”

The brief says the natural experiment is already running and that nobody has funded the analysis.

Science Diplomacy FR-VIII-27

Does scientific cooperation survive strategic rivalry because of shared scientific values, or because separation is physically and financially expensive?

Assemble and analyse the three outcome series produced by the single 2022 shock across three flagship arrangements - component delivery schedules, co-authorship by affiliation, user counts and beam-time allocations, crew rotation and reboost manoeuvres, before and after - to test whether continuity tracks the cost of separation rather than the strength of scientific ties. — natural experiment · running · Frontier

“The prediction from this brief is that continuity tracks the cost of separation and not the strength of scientific ties”

The brief says every series already exists and is held by an institution that publishes annual reports; nobody has assembled it.

Civilizational Archives FR-VIII-28

Can a sealed long-term deposit actually be recovered and understood by someone outside the institution that made it?

A blind read-back trial: give a sealed long-term deposit made years ago - an Arctic film reel, a micro-etched disc, an offline tape set - to an independent team with no access to the depositing institution and no documentation beyond what physically accompanies the artefact, and measure bit recovery, meaning recovery and elapsed time. — experiment · unscheduled · Frontier

“Everything this field believes about bootstrap usefulness is an untested assumption until somebody runs that trial and publishes the recovery rate.”

The brief says the trial is cheap and unblocked, the deposits that would make it possible already exist, and nobody has scheduled it.

Crisis Science Institutions FR-VIII-29

Has any post-2020 reform actually shortened the time from identifying a novel pathogen to randomising the first patient under a pre-authorised protocol?

Publish the interval from pathogen identification to first patient randomised under a pre-authorised protocol, for every country claiming the capability, and score every preparedness exercise on it between now and the next outbreak. — measurement · unscheduled · Frontier

“The decisive test is activation time, and it is measurable to the day.”

Nothing blocks the measurement: every trial records the dates. No journal requires it, no regulator publishes it, and no preparedness exercise is scored on it.

IX — Economics & Society

Post Scarcity Economics FR-IX-01

Are falling delivered prices driven by learning curves that can keep falling, or by non-curve costs that are not falling at all?

Decompose delivered prices into experience-curve and non-experience-curve components across several jurisdictions, decades and goods, using data that already exists. — measurement · unscheduled · Established

“The highest-value experiment is also the cheapest: decompose delivered prices into experience-curve and non-experience-curve components across jurisdictions, using data that already exists. The Berkeley and Brattle work does this for US electricity for one period. Doing it for several countries, several decades and several goods would convert the central dispute from an argument about anecdotes into a measured ratio.”

The brief says the abundance case would need the non-curve share to be falling somewhere; where it has been measured it is rising.

Universal Basic Abundance FR-IX-02

Does a universal transfer move rents, wages and prices once it is paid to an entire labour and housing market rather than to scattered households?

Run a saturation design: randomise the transfer at the level of a labour and housing market rather than the household, and observe rents, wages, prices, firm entry and labour demand, which is the direct test of the inflation objection. — experiment · unscheduled · Established

“The single most valuable experiment is a saturation design: randomise at the level of a labour and housing market, not at the level of a household.”

Kenya's village-level randomisation is the closest existing design and its scale is too small for prices.

Future Labour Markets FR-IX-03

Do firm, industry and regional estimates of automation's employment effect agree once a single design produces all three?

Estimate firm, industry and regional effects from one design on linked employer-employee registers, instead of the two-level-at-a-time national estimates that currently disagree. — measurement · unscheduled · Established

“The highest-value experiment is also the cheapest: report firm, industry and regional effects from one design. France and the Netherlands each estimate two of the three levels and find them inconsistent; nobody estimates all three together.”

The brief says the data exist in at least four countries with linked employer-employee registers and no new collection is needed.

Human Flourishing FR-IX-04

Do cross-country wellbeing comparisons measure different lives, or different ways of mapping the same life onto a number?

Measure the reporting function directly: present the same described life to respondents across countries, languages and income levels, anchor with vignettes at scale, and estimate each person's mapping from described state to reported integer. — measurement · unscheduled · Established

“The highest-value experiment is also the cheapest, and it is not a survey. Measure the reporting function directly: present the same described life to respondents across countries, languages and income levels, anchor with vignettes at scale, and estimate the mapping from described state to reported integer per person.”

The brief says the companion stochastic-dominance check requires no new data collection.

Wealth Distribution Systems FR-IX-05

Are predistributional instruments the high-leverage lever on wealth inequality, given that most of the gap is pre-tax and almost none of it has been measured?

Build an evaluation programme for predistributional instruments, minimum wage schedules, sectoral bargaining coverage, licensing, zoning, education financing and healthcare financing, giving them the elasticity literature that wealth taxation already has. — measurement · unscheduled · Speculative

“Fourth, and the one that would matter most: build an evaluation programme for predistributional instruments. If two-thirds to 90% of the gap is pre-tax, then minimum wage schedules, sectoral bargaining coverage, licensing, zoning, education financing and healthcare financing are the high-leverage instruments, and none has anything resembling the elasticity literature that wealth taxation has.”

The brief calls this not a single experiment but the recognition that the field has industrialised measurement of the smaller half of its own subject.

Future Capital Markets FR-IX-07

Does private capital actually beat public markets once every fund reports a public market equivalent on a common fee convention?

Require every fund marketing to institutional or retail investors to report a public market equivalent against a named benchmark on a standardised fee treatment, alongside the IRR. — policy result · unscheduled · Established

“The highest-value experiment in this subject is a disclosure rule, and it is cheap. Require every fund marketing to institutional or retail investors to report a public market equivalent against a named benchmark on a standardised fee treatment, alongside the IRR. The data already exist inside the funds.”

The brief says the performance dispute is about sample and convention rather than about the world, and a common convention dissolves most of it.

Innovation Ecosystems FR-IX-08

Do cluster interventions cause the regional outcomes they claim, or do they select places that were going to do well anyway?

Randomise or quasi-randomise a cluster intervention by running a lottery among the eligible, which every oversubscribed cluster programme is administratively able to do, and pre-register it. — experiment · unscheduled · Established

“The highest-value experiment is the one the field's own surveys have been asking for and nobody has run: randomise or quasi-randomise a cluster intervention. Eligible regions exceed available funding in every cluster programme in existence, which means a lottery among the eligible is administratively available and ethically defensible.”

Its absence after forty years is a fact about political tolerance for being measured, not about method.

Digital Economies FR-IX-09

Does the gap between individual and collective valuation of network goods recur beyond one platform, and must market-definition doctrine therefore carry a correction factor?

Replicate the collective-versus-individual valuation design on other network goods, messaging, marketplaces and operating systems, and see whether the substitution gap recurs. — experiment · unscheduled · Established

“The highest-value experiment is the one already invented and run once: replicate the collective-versus-individual valuation design on other network goods. If the gap between individual and collective substitution recurs for messaging, marketplaces and operating systems, the standard market-definition instrument has a measured bias and competition authorities need a correction factor.”

The design has already been invented and run once; the brief asks only that it be repeated.

AI Driven Productivity FR-IX-10

Do measured AI gains on intermediate outputs survive through to shipped, sold or delivered units outside software?

Replicate the intermediate-to-final-output decomposition outside software: take a setting with a measured AI gain on an intermediate output and follow it through to a shipped, sold or delivered unit. — measurement · unscheduled · Established

“The highest-value experiment is the cheapest and has been specified for two years: replicate the intermediate-to-final-output decomposition outside software. Take any setting with a measured AI gain on an intermediate output and follow it to a shipped, sold or delivered unit.”

The brief names customer support, radiology and legal drafting as settings that already have the final-output measure.

Resource Economies FR-IX-11

Do resource funds receive what their own inflow rules require, and is the Nigerian shortfall the tail of the distribution or its median?

Audit fund inflow rules against realised inflows across every country with a resource fund, using the published rules and the published deposits. — measurement · unscheduled · Established

“The cheapest high-value study in this subject is an audit nobody has run: fund inflow rules against realised inflows, across every country with a resource fund. The rules are published, the deposits are published, and the difference is arithmetic.”

The brief says a single well-constructed table would convert the most confidently asserted policy claim in resource economics into a measurement.

Future Trade Systems FR-IX-12

Has trade genuinely diversified away from China since 2018, or has a single alternative supplier taken most of the share China lost?

Recompute the supplier-concentration result on published trade data for 2023 to 2026, testing whether a single supplier still takes at least three quarters of the lost market share for most affected products. — measurement · near · Frontier

“The cheapest high-value study is a repeat of an existing calculation on newer data: recompute the supplier-concentration result for 2023–26. The original finding — that for over 70% of affected products a single supplier took at least three quarters of China's lost market share — is the sharpest available refutation of the de-risking story, and it is a straightforward calculation on published trade data.”

If concentration has since fallen the diversification claim revives; if it has risen, the brief says the policy has an outcome nobody intended.

Circular Economies FR-IX-13

Does eco-modulated producer responsibility change design decisions, or does the instrument fail at any fee level rather than only at the level it has been set at?

Apply the estimator that produced the null on flat fees to the dated introduction of modulated fees across jurisdictions, using the same panel of countries and materials. — natural experiment · unscheduled · Frontier

“The highest-value study available is an evaluation of an eco-modulated scheme using the design that produced the null on flat fees. The original panel used temporal variation in the fees actually charged across 25 countries and four materials; several jurisdictions have since introduced modulated fees on dated schedules.”

The brief says it would run on data that already exists.

Future Taxation Models FR-IX-14

Why did the global minimum tax raise roughly a third of its forecast: substance-based exclusions, transition rules, or adoption gaps?

Publish five years of global minimum tax outturn by jurisdiction against jurisdiction-level projection, which separates the substance-based exclusion, the transition rules and adoption gaps as explanations. — policy result · running · Established

“The highest-value experiment is publication rather than policy: five years of global minimum tax outturn by jurisdiction. The first year came in at roughly a third of forecast and the only public figures reach this brief through an interested secondary source.”

The tax is already in force and the data exist inside revenue authorities; the brief says the obstacle is publication.

Scientific Funding Models FR-IX-15

Do lottery allocation, milestone termination and people-not-projects funding outperform conventional peer review, on comparisons funders already hold?

Report the quasi-random allocation comparisons already sitting in funders' records, wherever part of a portfolio was assigned by a rule that allocates quasi-randomly above a quality threshold. — natural experiment · running · Established

“The cheapest and highest-value experiment in this subject has already been run and merely needs reporting. Wherever a funder has allocated part of a portfolio by a rule that assigns quasi-randomly above a quality threshold, the comparison exists in its records.”

The governance brief records the lottery case in detail and finds nothing published.

Civilization Scale Investment FR-IX-16

Will investors buy an ultra-long, consumption-linked sovereign instrument at size, or is the long-tenor market a theory with no bid behind it?

Issue an ultra-long sovereign instrument whose coupon is linked to consumption or output rather than to a nominal rate, and publish the order book: the tenor hypothesis predicts placement failure at size among solvency-constrained institutions, the design literature predicts the opposite. — experiment · unscheduled · Frontier

“The decisive experiment is an issuance, and it is available to any large sovereign that wants it. Issue an ultra-long instrument whose coupon is linked to consumption or output rather than to a nominal rate, and publish the order book.”

The brief says it is available to any large sovereign that wants it, with the Austrian 2120 issue as the control.

Economic Resilience FR-IX-17

What does a strategic reserve release actually do to prices, measured by a method fixed before the release rather than chosen afterwards?

Pre-register the evaluation of the next strategic reserve release before it happens, agreeing comparator markets, price windows and model specification in advance. — policy result · unscheduled · Established

“The cheapest high-value experiment in this brief costs nothing and has never been run: pre-register the evaluation of the next strategic reserve release before it happens. Two published estimates of the 2022 release differ by a factor of three and the wider one comes from the party that authorised it.”

The reserve exists, the release will happen, and the design cost is a memorandum.

Reputation Economies FR-IX-18

Does reputation survive being transferable, if every transfer is labelled and the chain of custody is visible?

Build a market in provenance-labelled reputation: permit transfer, record every transfer publicly, display the chain of custody, and measure whether a labelled name still commands a premium, how fast that premium decays across transfers, and whether pooling appears anyway. — experiment · unscheduled · Speculative

“The experiment that would move this subject most is the one nobody has run: build a market in provenance-labelled reputation. Permit transfer, record every transfer publicly, display the chain of custody, and measure whether a labelled name still commands a premium, how fast the premium decays across transfers, and whether pooling appears anyway.”

The brief says a ledger-based platform could run it in one product cycle.

Human Development Metrics FR-IX-19

Do development rankings survive a change in the aggregation rule, or do they carry an unstated dependence on choices their users never see?

Publish a parallel HDI under alternative aggregation, non-compensatory, with different goalposts and no income cap, and report how far the rankings move. — measurement · unscheduled · Established

“The cheapest high-value study is a recomputation, not a survey. Publish a parallel HDI under alternative aggregation — non-compensatory, different goalposts, no income cap — and report how far rankings move.”

The brief says either result is worth having and neither requires new data.

Future Philanthropy FR-IX-20

Do dormant donor-advised fund accounts eventually grant, and after how long?

Build an account-level panel of donor-advised fund behaviour with dormancy tracked to resolution, from records sponsors already hold. — measurement · unscheduled · Frontier

“The most valuable study available is the one the sector's own data would support tomorrow: an account-level panel of donor-advised fund behaviour with dormancy tracked to resolution. The existing study covers 2014–2022 and shows 22% of accounts granting nothing in a three-year window.”

The brief says the obstacle is publication, not measurement.

Critical Mineral Midstream FR-IX-21

Whether policy-built refining and separation capacity outside China durably reduces measured midstream concentration, or reverts once Chinese prices collapse.

The running natural experiment in rare earths: whether the top-supplier refining share, 90% in 2023 and 85% in 2025, continues toward the roughly 70% the IEA projects for 2035 under the US price floor, Lynas heavy separation and USD 65 billion of public finance, or reverts in the first Chinese price collapse. — natural experiment · running · Frontier

“The decisive test is already running: whether the rare-earth top-supplier refining share, 90% in 2023 and 85% in 2025, continues down to the roughly 70% the IEA projects for 2035, or reverts the first time Chinese prices collapse and the subsidised plants must sell into them.”

The scoreboard dates are fixed: the November 2026 expiry of the suspended October 2025 controls and the 2028 magnet-facility target.

Compute Concentration FR-IX-22

Does concentrated access to frontier compute convert into concentrated capability, or does capped access keep converging on the frontier?

The export-control natural experiment already under way: comparable, independently run evaluations of models trained under the compute ceiling against uncapped models released in the same window, with training-compute estimates attached, through 2026 and 2027. — natural experiment · running · Frontier

“The decisive result is the one already running: whether models trained under the export-control compute ceiling stay within about a year of the frontier through 2026 and 2027.”

Needs no new hardware, only comparable independent evaluations reported with dates and training-compute estimates attached.

Technological Sovereignty FR-IX-23

Whether a published, dated technological-sovereignty target functions as a binding instrument or as decoration.

The 2030 measurement of the European Chips Act's 20%-of-global-production-value target, and whether missing it forces a public re-scoping of the programme or is absorbed without consequence. — policy result · near · Established

“The decisive test in this subject is already scheduled and nobody scheduled it as a test: the European Chips Act’s 20% of global production value by 2030. It is the only technological-sovereignty target anywhere that is numeric, dated, externally measurable and already scored by an independent audit body that has projected a miss.”

Already scheduled and externally measurable, with an independent audit body having projected a miss; no other programme published a number a measurement could contradict.

Material Passports FR-IX-24

Does mandating a machine-readable material record change what actually happens to the material?

A pre-registered before-and-after study of the EU battery passport obligation taking effect on 18 February 2027, fixing recovery baselines per Member State and per chemistry beforehand, with recovered mass and recovered value per tonne as endpoints and non-EU recyclers of comparable chemistries as the comparison arm. — natural experiment · near · Frontier

“The decisive test is already scheduled and nobody has designed it: the battery passport becomes obligatory on 18 February 2027, which creates a dated before-and-after on a defined product class whose recovery statistics are already collected.”

The brief says the window for clean baselines closes in 2026; after the trigger the counterfactual must be reconstructed rather than observed, and no evaluation has been designed.

Climate Adaptation Finance FR-IX-25

Does money spent on adaptation buy a verified physical change, and does that change reduce measured loss?

Publish one audited, matched-sample series of claim frequency, claim severity and premium for wind-retrofit-designated homes against comparable undesignated homes across several Gulf storm seasons, alongside the grant cost, in a form an outside analyst could re-run. — measurement · running · Frontier

“The decisive measurement is a matched publication of claim frequency, claim severity and premium for designated versus comparable undesignated homes across several storm seasons.”

Nothing scientific blocks it: the claims data is proprietary, the grant data sits with a state agency, and no party benefits from a published ratio that might be worse than the advocacy figure.

Aging Societies FR-IX-26

What does a large, rapid immigration expansion and its equally rapid reversal actually do to wages, rents, output per head, the age structure and the sectors that staffed themselves from the inflow?

Evaluate Canada's 2021-2024 immigration expansion and its 2024-2027 contraction as a two-directional natural experiment, using the quarterly demographic, wage, housing and fiscal data and the administrative records that already exist. — natural experiment · running · Established

“The decisive test is a natural experiment already running in Canada, and it runs in both directions.”

The data to do it already sits in administrative records; cross-country regressions cannot separate the policy from the conditions that produced it.

X — Historical & Frontier Science Studies

History of Electrogravitics FR-X-01

Whether the founding citation of the electrogravitics suppression literature refers to an article that exists.

Requesting Interavia vol. 11 no. 5 (1956), pp. 373-374 through interlibrary loan: either the seventy-year-old citation is verified, or the founding citation of the suppression literature is formally recorded as untraceable. — observation · unscheduled · Established

“Two: request Interavia vol. 11 no. 5 (1956), pp. 373–374 through interlibrary loan. Outcome: the article is produced, in which case a seventy-year-old citation is finally verified; or it is not, in which case the founding citation of the suppression literature is formally recorded as untraceable.”

Both outcomes are publishable and the second is the more valuable; the experiments this subject needs can be completed by people with library cards.

Cold War Science Programs FR-X-06

Whether present-day secrecy orders suppress innovation at the rate measured for the wartime cohort.

Running Gross's secrecy-order design forward on the 6,543 orders now in force, comparing released filings against matched unreleased ones, to price the current instrument rather than the 1940s one. — natural experiment · unscheduled · Established

“The cheapest high-value study is to run Gross's design forward on the standing secrecy orders. The wartime cohort was measurable because the orders were eventually released and the technology classes are on the record.”

It requires a release schedule that is systematic rather than discretionary, which is why it appears as an institutional requirement rather than a research proposal.

Operation Paperclip FR-X-07

What the imported German scientists actually added to American technical capability, expressed as a number rather than a narrative.

Estimate the counterfactual: use the IWG record of names, arrival dates and assignments against continuous patent records and constructible comparison groups of American engineers in the same classes. — natural experiment · unscheduled · Established

“The single highest-value study in this subject is the one nobody has run: estimate the counterfactual. The ingredients exist. The IWG record supplies names, arrival dates and assignments.”

The design is the one already used successfully on the 1930s refugee flow, and running it would replace eighty years of narrative argument with a number and a confidence interval.

Historical Space Colonization Concepts FR-X-08

Whether a sealed ecological life-support system can hold its atmosphere over years once the hidden chemical sinks that broke Biosphere 2 are accounted for.

A re-run of the Biosphere 2 closure with the chemistry instrumented, carrying a full material inventory of the structure itself: every sink, every surface, every curing reaction. — experiment · unscheduled · Established

“The experiment that would matter most now is a re-run with the chemistry instrumented. The 1991 failure was invisible in the carbon dioxide record because the carbon went into the concrete.”

Nothing since has repeated it at that scale.

Soviet Frontier Science FR-X-09

What Lysenkoism actually cost Soviet agriculture, as a measured number rather than seventy years of adjectives.

A difference-in-differences on Soviet crop-yield series across crops and regions differentially exposed to mandated Lysenkoist agronomic practices, with the collectivisation famines separated in time from post-1948 policy. — natural experiment · unscheduled · Established

“The most valuable single study is the one the Lysenko literature has declined to attempt: estimate the agricultural cost. Soviet crop-yield series by region and crop exist; the timing of Lysenkoist agronomic mandates is documented; the collectivisation famines are separable in time from post-1948 policy.”

The obstacle is data quality in Soviet agricultural statistics, which is real and is not the same as impossible.

Scientific Revolutions FR-X-11

Whether taxonomic incommensurability shows up in the record, that is, whether successor lexicons cross-classify the kinds of the theory they replace.

Build the taxonomic test Kuhn's late position implies: take a discipline with documented terminological turnover, reconstruct what its kind terms picked out before and after, and check the no-overlap structure directly. — experiment · unscheduled · Frontier

“If successor lexicons cross-classify the incumbent's kinds, the strongest surviving version of incommensurability has its first empirical support. If they nest cleanly, it does not. This is the one test in the subject that could come out either way on evidence, and nobody has built it.”

Historical Megaprojects FR-X-13

Whether cost overruns have really worsened over time, tested on the kinds of project the popular argument is actually about.

A comparable cost-outturn series for buildings, dams and canals spanning the same period as the 2002 transport sample, extending the F-test outside rail, roads and fixed links. — measurement · unscheduled · Frontier

“One: extend the F-test outside transport. The 2002 sample is rail, roads and fixed links. The three examples that carry the popular argument are a building, a dam and a canal. A comparable outturn series for buildings, dams and canals spanning the same period would test the claim on the projects the claim is actually made about”

It does not exist.

Frontier Aerospace Programs FR-X-14

Whether the collapse in experimental-aircraft cadence is explained by rising real cost per demonstrator airframe or by something else.

A deflated real cost-per-demonstrator-airframe series from the X-1 to the X-59, with programme duration and flight count alongside, built from appropriations and audit data that already exist. — measurement · unscheduled · Established

“The cheapest decisive study is the unit-cost series. Deflated real cost per demonstrator airframe from the X-1 to the X-59, with programme duration and flight count alongside, using appropriations and audit data that already exist.”

It would settle H1 against H3 directly, and its absence is why the cadence collapse is currently explained by whichever cause the explainer prefers.

Philosophy of Science FR-X-15

What practising scientists actually treat as the criterion for classifying a claim as outside science.

A stratified survey of practising researchers across disciplines, presenting the published criteria and measuring endorsement, disagreement between fields, and the gap between endorsement and application on worked cases. — measurement · unscheduled · Frontier

“The highest-value missing study in this subject is a survey, and it is cheap. Ask a stratified sample of practising researchers across disciplines what would make them classify a claim as outside science”

The result would be a fact about the world rather than a position in a debate, and it would be the first one this field has had on its central question.

Technology Forecasting FR-X-16

Whether the correlated error in technology forecasts is a property of one technology, one model family, or the whole practice.

A second, disinterested scoring of the 2,905 integrated-assessment-model projections of solar cost decline, which are dated, numerical, published and already resolved by outturn. — measurement · unscheduled · Established

“The cheapest valuable experiment is to score the forecasts that already exist, and the highest-value corpus is sitting in the open. The 2,905 integrated-assessment-model projections of solar cost decline are dated, numerical, published, and resolved by an outturn nobody disputes.”

They have been scored once, by an interested party; a second scoring would cost a research assistant a summer.

Innovation History FR-X-17

Whether the measured decline in disruptive science is real or an artefact of reference-list growth.

Recompute the CD disruption series across the same six datasets with the correction Petersen and colleagues named and quantified. — measurement · unscheduled · Frontier

“Petersen and colleagues did not merely allege a bias in the disruption index; they named the mechanism and quantified it. Recomputing the CD series across the same six datasets with the correction applied would resolve the most-cited quantitative claim in the field one way or the other, and it requires no new data.”

It requires no new data; until it is published, innovation history has a headline finding and a specific arithmetic reason to distrust it.

Scientific Institutions Through History FR-X-18

Whether a rejection rate flat between 8.5% and 13.5% is a property of learned-society refereeing or an artefact of one institution.

Run the Royal Society manuscript analysis again on the editorial archives of a second learned society: a continental academy, a national society in a different tradition, or a long-lived commercial journal. — measurement · unscheduled · Frontier

“The single highest-value study in this subject is a second manuscript series. The Royal Society records give rejection rate, referee count, participation share and workload distribution across a century.”

The method is published, the outcome is interpretable either way, and the obstacle is access rather than technique.

Grand Challenges of Humanity FR-X-22

Whether the leap in DARPA Grand Challenge performance from 2004 to 2005 was a prize framing effect or learning, consolidation and falling sensor prices.

Decompose DARPA 2004 to 2005 using team-level data on continuity, spending, personnel and component costs across the two years. — measurement · unscheduled · Frontier

“Four: decompose DARPA 2004 to 2005. Team-level data on continuity, spending, personnel and component costs across the two years would separate the framing effect from learning, consolidation and falling sensor prices.”

The data existed at the time and the decomposition has never been published. It is the highest-value single study available in this subject.

Technology Trees of Civilization FR-X-23

Whether directed prerequisite structure, which technology must come before which, is recoverable at all from the data we have.

Take the normalised patent proximity network, add first-grant dates, and test whether any orientation rule recovers prerequisite relationships that domain experts endorse out of sample. — experiment · unscheduled · Frontier

“Second, and this is the one that would settle the subject: try to orient the edges. Take the normalised patent proximity network, add first-grant dates, and test whether any orientation rule recovers prerequisite relationships that domain experts endorse out of sample.”

Nobody has run it, and a clean negative would be the first direct evidence that directed prerequisite structure is not recoverable from the data we have.

Civilization Timelines FR-X-24

Whether the named periods of world history survive a pre-registered comparative test, or are artefacts of the narrative that named them.

A pre-registered periodisation test on the named periods that have never had one, following the Axial Age template: state in advance what the comparative data would have to show for the period to survive, and publish the result either way. — experiment · unscheduled · Established

“The cheapest valuable experiment is a pre-registered periodisation test, and it has been run roughly once. The Axial Age assessment is the template: take a named period, state in advance what the comparative data would have to show for it to survive, and publish the result either way.”

Each is testable against the same databank with the same method.

Science of Science FR-X-26

Does an AI research tool actually compress research stages in working laboratories without degrading output quality, when quality is scored by replication rather than citation?

Randomise access to an AI research tool at the laboratory level, log stage clocks across the pipeline, pre-register the quality metric, and score sampled outputs by independent replication or replication market two years on - the METR design scaled from code to bench. — experiment · unscheduled · Frontier

“The decisive test is a stage-instrumented randomised rollout of an AI research tool across working laboratories, with quality endpoints scored by replication rather than by citation.”

The brief says every component has been demonstrated separately but no funder has yet commissioned it; until one does, acceleration claims rest on self-report, vendor benchmarks, or the withdrawn preprint.

Cloud Laboratories FR-X-27

Whether machine execution of a protocol actually makes an experiment reproducible, and if it does not, where the variance lives.

Implement one non-trivial protocol on two independent cloud laboratories, run it blinded on both with identical inputs, and compare the results against each other and against a skilled bench laboratory working from the written protocol. — experiment · unscheduled · Established

“The decisive experiment is the cross-platform replication, and nobody has run it. Take one non-trivial protocol, implement it on two independent cloud laboratories, run it blinded on both with identical inputs, and compare the results against each other and against a skilled bench laboratory running the same protocol from its written form.”

Needs no new technology and no new instrument; what makes it expensive rather than routine is that no vendor-neutral instrument-control standard has majority adoption, so the protocol must be re-implemented for each platform.

Metrology Infrastructure FR-X-28

Do the uncertainties laboratories state for a traded but unstandardised measurand match the dispersion that appears when the same samples are measured independently?

A blinded interlaboratory round robin in durable carbon-removal quantification: identical homogeneity-tested samples to twenty or more laboratories, each reporting a value and its stated uncertainty under its normal procedure, with the reproducibility standard deviation published against those stated uncertainties. — experiment · unscheduled · Frontier

“The decisive experiment is a blinded interlaboratory round robin in a market that already prices an unaudited number. Distribute identical, homogeneity-tested samples to twenty or more laboratories that currently quantify durable carbon removal, have each report a value and its stated uncertainty under its normal procedure without knowing the others, and publish the reproducibility standard deviation against the stated uncertainties.”

Nobody has funded it, and the parties best able to are the ones whose claims it would test; every ingredient already exists.

Research Front Detection FR-X-29

What sensitivity and false-alarm rate does any early-detection method actually achieve when it has to commit before the outcome is known?

A prospective register in which several teams publish, on the same frozen corpus and the same date, ranked lists of research fronts they predict will be consequential, with outcome measures and thresholds fixed in advance, held by an independent body and scored at five and ten years. — experiment · unscheduled · Frontier

“The decisive experiment is a prospective register of dated emergence forecasts with resolution criteria fixed in advance. Several teams publish, on the same frozen corpus and the same date, a ranked list of research fronts they predict will be consequential, together with the outcome measures and thresholds by which they agree to be judged.”

Nothing about it requires new science; it requires a custodian and a decade of patience, and no funder has supplied either.

Extraction rules: the excerpt must appear verbatim in the brief (machine-checked); the horizon comes from the brief's own language, never from outside knowledge; where a brief ranks its tests, the index follows the brief's ranking. Version v0.5 · September 2026.