1 · Concept overview
Governments run standing machinery for thinking about the long term. Horizon-scanning teams read widely and write up what they find; scenario units build sets of alternative futures and run policy teams through them; national Delphi surveys poll thousands of experts on when a technology will arrive; futures reports are laid before parliaments once an electoral term; and risk registers rank what might go wrong, with likelihood bands and impact scores attached. The machinery is real, it is decades old, several of the units are internationally famous, and no government of any size is without some version of it.
This brief was commissioned to assess whether it works. It cannot, and that is the finding. Not because nobody has looked — a bibliometric review counts 186 evaluation papers over thirty years — but because the field decided early that outcome measurement was the wrong criterion, never converged on a replacement, and now operates without one. Four official or peer-reviewed assessments of whether government foresight changes decisions have been located. Four of four report an absence of demonstrated impact. None found measured impact. None used a counterfactual. What follows is a map of that evidence gap, the places where measurement was attempted and what it returned, and — new in this version — the reasons to think that the obvious remedy would make things worse.
Where this brief stops, stated precisely, because four slots in this category touch the same material. This brief owns the machinery and the products: the units, their methods, their outputs, and the accuracy of forecasts wherever anyone has scored one. Civilizational Planning owns whole-society plans with dates — the Montreal Protocol, national five-year plans, the millennium goals — and the question of what social discount rate correctly values a person born in 2200. Long-Term Institutions owns durability: what holds any commitment in place once made, whatever its content. Scientific Advisory Institutions owns the advice relationship, where a body with expertise and no mandate writes for a body with a mandate and less expertise. The arbitration rule where they meet is the object of the verb. If the question is who anticipated what, how well, and did anyone act, it belongs here. If it is what was promised and did the promise hold, it belongs to the neighbours.
One consequence of that cut should be flagged immediately, because it is where this brief has most changed since its previous version. The forecasting-accuracy literature sits here, and it is the only genuinely well-measured part of the whole subject. Expert political judgement has been scored across tens of thousands of resolved questions. Megaproject cost and demand forecasts have a seventy-year series. A thirty-year run of energy outlooks has been checked ex post against deployment. Expert predictions of a randomised policy experiment were elicited in advance and then marked. That body of work exists, it is unflattering, and the government foresight literature almost never cites it — which is itself one of the more interesting facts in this brief.
2 · Current scientific position
Established The machinery exists and its shape is checkable. The UK function now sits in the Government Office for Science under the name Futures, Foresight and Emerging Technologies, with the foresight role shared with the Cabinet Office through a joint data and analysis centre — not the standalone “Foresight programme” most commentary still describes. Its published project collection runs from 2003 to 2026, roughly thirty-four to thirty-six projects, about 1.5 a year. Singapore runs two separate bodies that are routinely conflated: a Centre for Strategic Futures established in 2009 inside the Prime Minister's Office, and a Risk Assessment and Horizon Scanning programme launched in July 2005 inside the national security coordination secretariat. Finland has produced a Government Report on the Future since 1993, once per electoral term rather than annually. The European Commission has published strategic foresight reports in 2020, 2021, 2022, 2023 and 2025 — there was none in 2024.
Established Every institutional claimant disclaims the one property that could be measured. The UK's own guide states that “the future cannot be predicted with confidence”. The Cabinet Office told a parliamentary committee that horizon scanning “is not about predicting exactly what will happen”. A scholar writing in Singapore's own foresight journal proposes that such work be “judged more by the conversations it provokes, rather than how accurate it is ex post”. In place of accuracy the claimed benefits are resilience, earlier pattern-spotting, and policy that “stands the test of time” — none of which has an agreed measurement.
Established The most adverse verified finding is parliamentary, and it is twenty-two years into the programme it examined. A Commons select committee reported in April 2014 that the science office does work witnesses called world-leading — and that its “non-central location has limited its influence and horizon scanning remains… rarely used across much of government”. It found departmental practice varying from impressive to absent, described a newly created cross-government programme as a “quick fix” operating in near-total secrecy, and concluded that policy decisions are made for short-term reasons. The same committee recorded that an earlier internal review had found “an overabundance of reports [delivering] little in the way of policy change”, and that “horizon scanning is ignored when the strategic level is not open to challenge”. Then it recorded the Cabinet Secretary conceding, of work he had just called brilliant and path-breaking, that it had “not always translated into actual policy changes”. Good output, no uptake is the committee's verdict in five words.
Established An internal review reached the same place two years earlier, and its most damaging sentence is about the reviews themselves. A review of cross-government horizon scanning commissioned by the Cabinet Secretary found strong analyst networks producing informative products, “weak integration between Departments and with the policy agenda, particularly at the senior level”, products “often lengthy, and poorly presented, making them harder to digest and easier to ignore”, and at least twenty organisations scanning without cross-government oversight. Its own summary sentence is that horizon scanning “is wasted resource if it is not ultimately used to inform the policy agenda”. And the load-bearing admission, which is a statement about the reform record rather than about the practice: “previous reviews have attempted to embed cross-cutting horizon scanning into government structures without enduring success.” This is a government telling itself, in 2012, that it had already tried this repeatedly and it had not held.
Established The study commissioned to find the impact evidence reported that it could not. A foresight consultancy, paid by the science office, ran eight country case studies — Canada, Finland, Malaysia, the Netherlands, New Zealand, Singapore, the United Arab Emirates and the United States — and a thirty-expert workshop, and wrote: “During our research we looked for evidence of the impact of strategic foresight specifically on the policy process… it is much harder to show direct causal links between a piece of foresight work and a specific policy change.” Its blunter formulation is that “there is a limited evidence base on the impact of foresight work” and that the available case studies examine how projects or units used foresight rather than how governments as a whole did. Its substitute proxy for impact is “ongoing and increasing demand”. Apply the interested-party rule and the finding gets stronger, not weaker: the party with every reason to find impact, looking specifically for it, found none.
Established The peer-reviewed literature says the same thing in 2026 and is candid about not having tested it either. A Policy and Society paper on strategic foresight as a source of evidence for policymaking states that “scarce attention on the evaluation of SF exercises means that their quality remains untested”, and its authors then say plainly that they did not attempt it — “our objective was not to evaluate the substantive content… nor to evaluate their policy impact and uptake”, which “would have required further primary data collection and process-tracing”. That is the fourth of the four reviews, and it closes the pattern: each one looks, each one reports the absence, and each one identifies the missing study as somebody else's to run.
Established Practitioners on the receiving end report the same thing from outside government. A transport trade body giving written evidence to the parliamentary inquiry said that the science office's own infrastructure scenario work had “little lasting impact on subsequent policy”. That is an interested party of a different kind — a body that would benefit from more infrastructure foresight, not less — and its complaint is that the foresight it wanted was produced and then went nowhere.
Established There is exactly one place in government foresight with a numerator and a denominator, and it is a compliance count. A parliamentary research service examined sixty-three European Commission impact assessments from January 2020 to December 2023 and found nine supported by a dedicated foresight study. The Commission's own internal scrutiny board, grading the same corpus, reported that 73% of impact assessments “adequately considered” foresight elements, up nine points on the previous year. Both numbers count process. Neither asks whether any resulting policy was better, and the two are not in contradiction — they measure different things and both measure the wrong one.
Established The best available diagnosis of why this pattern recurs was written about something else entirely, and it fits foresight better than anything written about foresight. The standard account of why designed institutions are adopted without producing performance is isomorphic mimicry: “the tendency to introduce reforms that enhance an entity's external legitimacy and support, even when they do not demonstrably improve performance”, producing a capability trap in which organisations “constantly adopt ‘reforms’ to ensure ongoing flows of external financing and legitimacy yet never actually improve”. Read the foresight record against that definition. A unit whose output is praised as world-leading and used rarely; a duty to consider foresight satisfied in nine impact assessments out of sixty-three while a scrutiny board grades the same corpus at 73% compliance; a discipline that adopted a success criterion under which no negative result can be returned; and a reform history whose own reviewer records that previous attempts to embed it succeeded without enduring. This brief does not claim that foresight is mimicry. It claims that mimicry is a well-specified hypothesis that fits every observation in the record, that no evidence currently distinguishes it from the alternative, and that the field has never tested it.
Established And the failure to evaluate is not special to foresight, which changes what should be inferred from it. The UK's audit office found that in 2019 only 8% of government spend on major projects — £35 billion of £432 billion — had robust evaluation plans in place. Of the 108 most complex and strategically significant projects, just 9 were robustly evaluated, while 77, representing 64% of spend, had no evaluation arrangements at all. The Prime Minister's Implementation Unit concluded that “government has little information in most policy areas on what difference is made by the billions of pounds being spent”. If 8% of near-term, high-salience, well-funded projects are evaluated, the prior probability that a fifty-year anticipatory function has been evaluated is close to zero. The thinness of this literature needs no explanation specific to foresight, and attributing it to the field's philosophical objections to outcome measurement over-explains an observation that the base rate already covers.
3 · Frontier questions
The genuinely open questions in this subject divide into three, and only the second of them is a research question in the ordinary sense. The first is whether the products can be scored at all; the second is what accuracy is actually attainable and by whom; the third is whether accuracy is even the property that matters. Most writing on foresight collapses the three, which is how a field with a serious measurement problem comes to look as though it merely has a philosophical disagreement.
Frontier The one national instrument that has been scored is Japan's Delphi series, and this brief will not print the number yet. Japan has run large-scale national Delphi forecasts roughly every five years since the early 1970s; a 2026 study assessed the realisation status of transport topics forecast in the 1992, 1997 and 2001 rounds across 167 topics and 647 observations and reported a single headline accuracy rate. A corrigendum to that paper was subsequently published, and its content could not be obtained. Printing a figure whose correction is unread would be exactly the failure this map exists to avoid. The figure is well below half; it is stated here no more precisely than that until the corrigendum is read. Japan's own forecasting institute publishes no realisation-rate assessment of its surveys at all.
Established The adjacent field where accuracy is measured points two ways at once, and the two studies most often run together should not be. The expert-judgement programme reviewed in the forecasting literature accumulated 82,361 forecasts from 284 experts over roughly two decades. Its findings, in the reviewers' summary, are that “the experts barely if at all outperformed informed non-experts”, that “neither group of forecasters did well against simple rules and models”, and that when outcomes were reduced to three states the forecasters “frequently did worse than assuming that each state is equally likely”. Foxes beat hedgehogs; neither beat the algorithm. The forecasting-tournament work of 2011–15 found large accuracy differentials within that population: standardised Brier scores separating trained from untrained, teams from individuals, and a top-2% cohort furthest ahead of all, attempting more questions and updating far more often. The first is a finding about the ceiling on expert judgement; the second is about the spread beneath it. They have different designs and partly opposing implications, and a foresight unit that cites the second without the first is claiming a ceiling it has no evidence for.
Frontier And the tournament study's causal claims are contested in the journal that published them. A 2025 re-analysis using item response theory reports that the best-fitting models included extraneous variables that “substantially eliminated, reduced, and, in some cases, even reversed the effects of the experimental manipulations”. What is contested is not that accuracy differed between forecasters — it is whether the training and teaming interventions caused it. Anyone citing “training improves forecasting by X” must carry this.
Established The most useful calibration result for government foresight is not from a forecasting tournament at all. It is from an aid experiment, and its lesson is about which questions experts are good at. Before analysing an eleven-year follow-up of a randomised community institution-building programme in Sierra Leone, the research team elicited predictions from 126 experts — policymakers, academics and students. On physical infrastructure the mean forecast of 0.218 standard deviations almost exactly matched the realised 0.204. On institutional change experts predicted 0.095 against a realised 0.062, and the country's own policymakers predicted around 0.25 — roughly four times the truth. On one purely operational question, what share of communities would enter a competitive grants process, experts predicted 42% against an actual 98%. Expert judgement was well calibrated about things being built and badly calibrated about institutions changing, and proximity to the programme made the error larger rather than smaller. Government foresight is almost entirely in the second category, and the people staffing it are almost entirely the proximate kind.
Frontier The third frontier is the live design question and nobody has run it on a foresight product: how should an anticipatory document communicate its own uncertainty? This matters practically. The UK's committed post-pandemic reform is to move from a single reasonable-worst-case scenario to a wider range, which is a change in uncertainty communication before it is a change in analysis. There is now a randomised evidence base on exactly that manoeuvre — and it belongs to the advisory literature, where Scientific Advisory Institutions treats it as evidence about the advice relationship. Its transfer to foresight products is untested and is stated here as an open question rather than a result.
Established What the trials establish is that publishing uncertainty does not cost trust on average. Two pre-registered UK randomised experiments totalling 3,962 participants compared official persuasive text against balanced text disclosing the balance of risks and benefits, uncertainties and evidence quality. On vaccination (n = 2,928) there was no significant difference in rated trustworthiness of the information (5.28 against 5.39, p = 0.21, d = 0.08) or of its producer (5.16 against 5.28, p = 0.14, d = 0.09). On nuclear power (n = 1,034) the balanced version scored significantly higher on both information (4.70 against 4.92, p < 0.01, d = 0.17) and producer trustworthiness (4.35 against 4.63, p < 0.01, d = 0.20). A German experiment (N = 800) then showed that disclosing uncertainty in advance buffers the loss of trust when the statement later turns out to have been wrong, and that explaining the uncertainty does no additional harm.
Established But the effect is a crossover, not a main effect, and that is the correction that matters for a risk register. Two pre-registered US experiments (n = 600 and n = 1,001) found that whether uncertainty communication raises or lowers trust depends on whether the evidence agrees with the recipient's prior beliefs. Where beliefs and evidence were consistent, adding uncertainty reduced trust in the information and the source; where they were inconsistent, it increased both. The interaction terms were large and consistent, β between −0.26 and −0.41, against a strong main effect of belief–evidence consistency (β = 0.58 and 0.50) — people trust evidence that agrees with them. There was no evidence of polarisation. Transparency is therefore not a trust-maximising move but a trust-redistributing one, buying reach into the sceptical part of the audience and insurance against the day the assessment changes, at a small cost among those who already agreed. Those are precisely the two properties a pandemic risk assessment most needed in 2020 and did not have. That is a strong design case; it is not the case usually made; and an institution that adopts multiple scenarios expecting to be more believed will conclude within a year that the reform did not work.
Frontier What remains genuinely undecidable on current evidence is the counterfactual, and it is worth naming precisely rather than gesturing at. Nothing in the located literature distinguishes foresight does not work from foresight works and governments ignore it. The committee that called the output world-leading is the same committee that called it rarely used. Both hypotheses are consistent with every observation in this brief, and a third — that the practice is adopted for legitimacy and its non-use is the equilibrium rather than the failure — is consistent with them too. Three live hypotheses, no discriminating measurement.
4 · Technological bottlenecks
Established The first bottleneck is not analytical capacity. It is that nothing downstream is obliged to move. Every adverse finding in this brief is about wiring rather than method: strong analyst networks, weak senior integration; world-leading reports, rarely used; scenario work with “little lasting impact on subsequent policy” on its intended users' own account; a duty to consider foresight in impact assessments, satisfied nine times in sixty-three. A method problem would show up as disputes about technique. What actually shows up, in every review since 2012, is a transmission problem.
Established The second is that nobody publishes the cost. The last disaggregated figure for the UK programme is from 2009 — twenty-seven permanent staff and about £3 million a year, of which roughly £1 million for the horizon-scanning centre. A 2017 audit gives the whole science office £4.6 million and about sixty staff, down 27% in five years, with no separate foresight line. Singapore publishes neither headcount nor budget for either of its bodies. The cost side of any cost-benefit assessment of government foresight is simply unavailable, anywhere, currently.
Established The third is the most striking, because the data exists and the score does not. The UK has run an internal crowdsourced forecasting tournament since April 2020 — on the reported figures, 1,300 to 2,000 forecasters across forty-one government departments and allied countries, with more than ten thousand forecasts logged. Forecasting tournaments are scored by construction; a Brier score falls out of the platform whether anyone wants it or not. Six years in, no accuracy results have been published.
Frontier The fourth binds in the opposite direction from the first three, and it is the reason this brief no longer recommends measurement without qualification. Making an institutional target legible and scored reliably produces the score rather than the institution. The clearest demonstration is a governance index that ran for seventeen years: over 70 countries formed regulatory reform committees oriented specifically to its indicators; consultancies were paid to move rankings, with at least 9 of the twenty most rank-volatile countries in one region receiving over US$100 million in donor-funded business-environment projects between 2005 and 2019, one of them contracted with a top-forty position as a core objective and achieving it; and over fifteen years 111 countries swung more than 40 places out of 190 while 35 countries swung more than 75 — volatility no plausible account of underlying institutional change explains. An independent investigation then found scores altered under management pressure for four named countries, no codified written procedures for methodology change, and an environment its own interviewees described as toxic; the index was cancelled in September 2021. Any proposal to publish accuracy league tables for foresight units inherits this failure mode in full, and the proposal in this brief's experiments section is made with that caveat attached rather than without it.
Established The fifth is that measurement in public bureaucracies has a documented association with worse delivery, not better. In the only large-scale hand-coded study of project completion across a whole federal bureaucracy — 4,700 projects across 63 Nigerian federal organisations, with a companion study of 3,628 projects and tasks across 31 Ghanaian civil service organisations — an index of bureaucratic autonomy is robustly positively correlated with project initiation, full completion and average completion rate, while an index of incentives and monitoring is robustly negatively correlated with all three. The authors note that this runs counter to the private-sector management literature. The available summary reports directions and not coefficients, so no magnitude is attached here. What it establishes is that “score the foresight units” is not a costless prescription; it is an intervention with a prior against it.
Frontier The sixth is that the index you would score them against may not be a stable object. A comparison of seven established cross-national governance and state-capacity measures found pairwise correlations of 0.70 to 0.94 and a principal component absorbing 86.91% of common variance against 4.74% for the second — convergent validity, apparently measuring one thing. They nonetheless do not behave as one thing: 45 countries diverge by more than 0.40 standardised units between two of them and seven by more than 0.60, with divergence systematically largest at intermediate capacity levels, which is where most policy-relevant countries sit. Three published findings about democracy and state capacity reverse or vanish depending on which index is substituted. The author's conclusion is that convergent validity is high but “the interchangeability of the measures is low” and that “no measure of state capacity seems to be clearly superior to others”. Read alongside the finding that a pandemic-preparedness index correlated negatively with pandemic outcomes, the lesson is not that indices are useless but that an anticipatory capability scored by index is scored on an instrument whose properties nobody has established.
5 · Research dependencies
Established Nothing in this brief waits on a research result. The methods are decades old, their limits are documented, and no discovery is pending that would change what a horizon-scanning team is capable of. That is an unusual sentence in a frontier-research corpus and it is the correct one here.
Established What it waits on is measurement that institutions have chosen not to perform. Four things, all institutional: an agreed success criterion, which the field's own quality literature says does not exist — not even agreement on the criteria by which quality could be assessed; retrospective scoring of forecasts and risk assessments that were published with dates and probability bands and have never been marked; disclosure of programme cost, unavailable for every unit named here; and a mechanism connecting output to decision, whose absence is the finding every independent reviewer has reached. None of the four requires new method and none requires money on the scale of a single foresight project.
Established It also waits on something it cannot supply for itself: an evaluation culture in the surrounding government. Where only 8% of major-project spend carries a robust evaluation plan and 64% of that spend carries none, a demand that the anticipatory function be evaluated to a standard the delivery function is not held to will not survive a spending review. Foresight's measurement problem is a special case of a general one, and it will most likely be solved, if at all, as a by-product of solving the general one.
Established What depends on it is the more consequential direction, and the honest answer is that nothing demonstrably does. Every brief in this corpus with a long-horizon risk would in principle benefit from earlier warning. That relationship is exactly the one that has never been demonstrated, which is why no typed enabling edge is claimed below. Civilizational Planning and Long-Term Institutions both inherit this brief's measurement problem rather than depending on its output, and Existential Risk Governance inherits it at the tail of the distribution, where the base rates are worst.
6 · Required experiments
Established The cheapest and most valuable experiment in this entire subject is retrospective scoring of documents that already exist. The UK national risk register has been published since 2008 with ranked, banded, dated assessments. Eighteen years of falsifiable material sits in the public domain and no scoring of it has ever been published — not by the government, not by the audit office, not by parliament, not by the engineering academy that reviewed the methodology, not by the statutory inquiry, not in the academic literature. Every independent review located examined method, transparency or governance. None asked what the register got right.
Established Second: publish the tournament scores. A government forecasting platform with ten thousand-plus resolved and scorable forecasts is the closest thing to a calibration laboratory any state possesses. Publishing aggregate Brier scores by department and horizon costs nothing and would settle several questions in this brief at a stroke. It would also expose the field to the gaming risk in the bottlenecks section, which is a reason to publish aggregates by department and horizon rather than league tables by unit, not a reason to publish nothing.
Frontier Third, and harder: a counterfactual design. Nothing in the located literature is controlled, quasi-experimental or counterfactual. Randomising which of a set of comparable policy teams receives a scenario input, and measuring the decisions rather than the satisfaction, has apparently never been attempted in government. The one existing quantitative outcome study in the wider field is corporate, survey-based, n = 42–70, with unaddressed reverse causality — profitable firms can afford foresight units — and it is nonetheless the study most often cited in government advocacy.
Frontier Fourth, and this is the design suggestion that is new in this version: measure the removal rather than the founding. There is a methodological asymmetry visible in the neighbouring literature on central bank independence, and it generalises. A doubly robust causal estimate of whether creating that institution lowered inflation, across 60 countries from 1998 to 2010 with 17 measured variables, returns an average effect of +0.01 percentage points with a confidence interval from −1.48 to +1.50 — the authors say there is “only a weak causal link… if at all”. A 2026 study of 132 central bank governor transitions across 28 central banks, classifying 50 as politically motivated, finds that destroying it is followed by inflation rising about 2 to 4 percentage points. That second study is a working paper from an institution with a standing commitment to the design it is evaluating, and is flagged accordingly. The transferable point is methodological: institutional designs may be far easier to measure in the breach than in the observance, because destruction has a date and construction does not. Foresight has never been evaluated this way, and the natural experiment is running now — anticipatory and advisory bodies were closed at scale in the United States during 2025, an episode Scientific Advisory Institutions documents in detail. Whether the jurisdictions that lost them anticipate worse is answerable, dated, and nobody is measuring it.
Established Fifth: the uncertainty-communication trials are a ready-made design waiting to be transferred. The instrument already exists — randomised presentation of a balanced, uncertainty-disclosing version of a document against a confident version, with trust in the information, trust in the producer and stated intention measured separately. It has been run on vaccination advice, nuclear power and extreme-weather trends across roughly six thousand participants in three countries. Running it on two versions of a national risk register entry, or on a single reasonable-worst-case scenario against a scenario set, would cost a fraction of one foresight project and would produce the first experimental evidence about a foresight product that has ever existed.
Established A negative result already recorded, and worth stating as an experiment rather than as an embarrassment. The one cross-national quantitative test of whether anticipatory capacity produces better outcomes returned the wrong sign: an ex-ante national preparedness index scored against COVID outcomes across 36 OECD countries gave a negative moderate correlation. Carry its caveats — one index, one disease, n = 36, a mid-May 2020 snapshot preceding most of the mortality, and not a test of foresight units. It is nonetheless a completed experiment with a published result, and the field's response to it has been to not run another.
7 · Engineering requirements
Established The one genuinely methodological body of work here is on the Delphi technique, and its own reviewers dismantled it. A 1999 review counted comparative studies: Delphi better than statistical groups in twelve cases against two worse and two ties; better than face-to-face groups in five against one and two; and no consistent evidence of outperforming other structured procedures. The same authors then explained why the count is nearly worthless — laboratory Delphi uses simplified almanac questions and student subjects lacking genuine expertise, where the method was designed for disparate experts, and feeds back medians rather than reasoned justifications. Their conclusion is that the evidence cannot determine whether properly conducted Delphi performs substantially better.
Established What that literature does supply is a set of conditions associated with accuracy, which is more than the rest of the field offers: five to twenty experts with genuinely disparate knowledge, two to three rounds, feedback consisting of averages plus written justifications, equal weighting, balanced question wording, and responses elicited as frequencies. Frontier The canonical internal critique of the method, published in 1975, is cited here by existence only; the text was not obtained.
Established On the operational side, the risk-register methodology is documented and has changed under criticism. The 2025 UK register carries 89 risks across nine categories, likelihood banded 1–5 from under 0.2% to over 25%, impact scored across seven dimensions, and horizons of five years for non-malicious risks and two for malicious ones — a change from the uniform two-year outlook a Lords committee criticised in 2021. Its own disclaimer is explicit: the scenarios “are not a prediction of what is most likely to happen”.
Established Banded probability language is a technology with a measured failure rate, and the measurement was taken on the closest available analogue. A study of 556 respondents (from 841 invited, a 66% completion rate) randomised across three presentation formats tested how readers interpret the calibrated probability terms used by a major assessment body. Interpretations were regressive in every term and in both directions: very unlikely, defined as below 10%, was read at a mean of 41%; unlikely, below 33%, at 44%; likely, above 66%, at 54%; very likely, above 90%, at 62%. Presenting the verbal term alongside its numerical range raised consistency with the publishing body's own guidelines only from 20.76% to 30.12%, and the median rank-order correlation with the intended ordering from 0.33 to 0.67. Even the best format tested left about seven readers in ten outside the intended range. The instrument tested was a climate assessment rather than a national risk register, and the transfer is this brief's inference rather than a finding; but a register that publishes five likelihood bands and a matrix, to an audience of ministers and officials, is using the same technology, and a Lords committee has separately warned that its likelihood-impact matrix “may be misleading and lead to a false sense of confidence”.
Established The record of what has actually been scored is short enough to tabulate, and tabulating it is the clearest single statement this brief can make. Every row is an instrument that produced dated, falsifiable claims. The middle column is who marked them, where anyone did.
| Instrument | Scored by | Result |
|---|---|---|
| UK national risk register, 2008–2025, ranked and banded assessments | Nobody | Eighteen years of falsifiable material, never marked by government, audit office, parliament, inquiry or the academic literature |
| UK internal crowdsourced forecasting tournament, since April 2020 | Scored internally by construction | 1,300–2,000 forecasters, 41 departments and allied countries, 10,000-plus forecasts; no accuracy results published in six years |
| Japanese national Delphi rounds of 1992, 1997 and 2001, transport topics | Independent academic study, 2026 | 167 topics, 647 observations, one headline realisation rate; well below half; figure withheld here because the corrigendum could not be read |
| Energy outlooks published annually 1993–2022, solar deployment | Academic ex-post analysis (an interested party) | Systematic underestimation across the whole thirty-year run; direction and persistence are the defensible claims, not the magnitude |
| Megaproject cost and demand forecasts, ~70 years of comparable data | Megaproject research programme | Nine in ten over budget; rail averaging 44.7% cost overrun with 51.4% demand shortfall; overruns “high and constant” across seventy years |
| Expert political judgement, ~20 years | Peer-reviewed review of the programme | 82,361 forecasts by 284 experts; experts ≈ informed non-experts; both lost to simple rules |
| Geopolitical forecasting tournament, 2011–15 | Tournament researchers; re-analysed 2025 | Large within-population Brier differentials; causal attribution to training and teaming contested in the same journal |
| Expert predictions of a randomised institution-building trial | The trial's own authors, pre-registered | 126 forecasters: calibrated on infrastructure, out by half on institutions, local policymakers out by a factor of four |
| 2019 national pandemic-preparedness index against COVID outcomes | Independent academics, 2020 | 36 OECD countries; Spearman's rs = −0.41, p = 0.013 — higher scores, worse outcomes |
Established Read the table as a whole and one pattern is unmistakable. Every row that was scored was scored by somebody outside the producing institution, and every row that was not scored belongs to an institution that could have scored it at negligible cost. That is not a methodological limitation. It is a distribution of who is willing to be marked.
8 · Adjacent technologies
Within this map the nearest neighbour is Civilizational Planning, and the two briefs are best read as a pair. That one asks whether a society can adopt a plan whose horizon exceeds the careers of everyone who adopts it; this one asks whether the machinery that is supposed to see the future coming produces anything a decision-maker uses. They share four sources deliberately and split them by question: the same parliamentary and internal reviews supply the uptake finding here and the planning-capability finding there, and the same seventy-year megaproject series is a forecasting-accuracy result here and a bottleneck on delivery there. Read together they make an argument neither makes alone — that accurate long-range projection and effective long-range response are entirely separable capabilities, and that the evidence for the first is much better than the evidence for the second.
Scientific Advisory Institutions shares the transmission problem and now has better evidence on it than this brief does. The division is by whose product is at issue: an expert body writing for a minister is theirs, a standing unit inside the machine writing scenarios for its own colleagues is ours. The uncertainty-communication trials belong to both, and the split is that they own the effect on the advice relationship while this brief owns the untested transfer to anticipatory documents.
Long-Term Institutions supplies the case that makes this brief's epistemic position generalisable: an office with the strongest formal instrument any future-generations body holds anywhere publishes no count of what it has done. Existential Risk Governance inherits this brief's measurement problem at the tail, and Civilization Resilience Planning is where anticipation meets preparedness — the pandemic case belongs to both and the division is that identification is ours and readiness is theirs.
Outside the map: the forecasting and judgement literature, which is where the accuracy data actually lives and which government foresight writing barely cites; risk analysis; evaluation methodology, where the counterfactual designs this brief keeps asking for already exist and are routine; and public administration, where the uptake question properly belongs and where the most useful result in this brief — that monitoring is negatively associated with delivery in the one bureaucracy-wide study of it — was produced by people not thinking about foresight at all.
9 · Institutional requirements
Established The standing pattern in this subject is unusually lopsided and the reader should know it before reading any figure. Almost every positive claim originates with a foresight unit, a foresight consultancy, a commercial forecasting vendor, or a government describing its own machinery. Almost every adverse finding originates with a parliamentary committee, an audit office, a statutory inquiry or a peer-reviewed journal. That asymmetry is the single most important fact about this literature, and it is the reason the strongest sources in this brief are the ones with no stake in the answer.
Established Two of the interested-party sources are cited here because their interest runs against their finding. The consultancy report commissioned to demonstrate impact reported it could not establish causal links. The internal cross-government review, written by government about government, is more adverse than anything the field's own academics have published. Interest that cuts against the result raises the weight of the result.
Frontier The commercial forecasting claims are the sharpest trap. A vendor's own page asserts accuracy margins over competitors, a percentage advantage over intelligence analysts with classified access, and a headline accuracy rate across several hundred resolved questions. Most carry no peer-reviewed citation. If any of it appears anywhere, it must be attributed to the vendor and not stated as a finding.
Established One international body is frequently miscast. The OECD's government foresight community is a network and community of practice — an annual meeting, a virtual day, masterclasses — not a producing unit, and it publishes neither a founding year nor a membership count. No verified global census of government foresight capacity exists. The nearest available count is of a different object: a comparative study identifies roughly twenty-five institutional proxy representatives of future generations across seventeen countries, of which sixteen still exist, and it is Civilizational Planning that carries that figure. Even there, the study's author states that impact data has not been collected.
Established Who buys foresight is the question that explains most of the record, and the answer is that the buyer and the user are different people. Foresight is commissioned by centres of government — a prime minister's office, a cabinet office, a chief scientific adviser — and is supposed to be consumed by line departments making decisions. Nothing obliges the second group to read the first group's output, and every review since 2012 has found that they largely do not. A function whose purchaser is not its user has no feedback loop by construction, which is a more precise statement of the wiring problem than “senior buy-in” and points at a different remedy: not more seniority at the producing end, but a procedural hook at the consuming end. The one jurisdiction that built such a hook — a requirement to consider foresight in regulatory impact assessment — produced nine dedicated foresight studies across sixty-three assessments, and a compliance grade of 73% on the same corpus. That is what a procedural hook without an outcome test yields.
Frontier The institution that does not exist is a scorer. No audit office, parliamentary committee, statistical agency or academic centre has taken on the standing task of marking published anticipatory assessments against outturn. Every body that could has instead examined method, transparency or governance. The absence is not accidental: the producing institutions have no incentive to create it, the auditing institutions treat foresight as a process to be reviewed rather than a claim to be checked, and the academic field has spent thirty years arguing that accuracy is the wrong criterion. A scoring institution is the single missing piece of machinery in this subject, it would cost less than one foresight project, and nothing in the record suggests anyone intends to build it.
Frontier And the bodies that do exist are less secure than the literature assumes. Anticipatory and advisory committees were closed at scale in one large jurisdiction during 2025, including bodies whose remit was explicitly forward-looking; the design finding that emerged — that statutory bodies survive administrative abolition and non-statutory ones do not — belongs to Scientific Advisory Institutions and is recorded here because it applies to foresight units directly. Most government foresight capacity sits in non-statutory units inside executive departments, created by administrative decision and removable by one. A function that has never demonstrated its own value and sits on the removable side of that line is in a weaker institutional position than the volume of its output suggests.
10 · Ethical & societal considerations
Established An unfalsifiable success criterion is an ethical question, not merely a methodological one, because public money is spent against it. The field's most influential move — a 2006 argument that foresight evaluation should rest on behavioural additionality and a systems-failure rationale rather than outcome measurement — is intellectually defensible. It is also, in combination with the standing disclaimer that foresight is not prediction, a position from which no negative result can be returned. Both things are true simultaneously and neither cancels the other.
Established The field says as much about itself. A 2015 paper setting out quality criteria for futures work opens by conceding: “There is, however, no common understanding of these studies' quality — not even of the criteria according to which the quality of foresight can be assessed.” A 2024 review of 186 evaluation papers over thirty years concludes that foresight evaluation “has not yet been thoroughly studied” and “is still an emerging discipline”. Frontier After three decades, the replacement yardstick has not arrived. The result is not measured against a different standard; it is not measured.
Established The distribution of interest in this literature is lopsided in a way that is an evidence-quality question before it is a political one, and the institutional section below sets it out. The consequence for this brief's method is that sources with no stake in the answer carry the weight, and interested parties are cited chiefly where their interest runs against their finding.
Frontier There is a second-order cost to unearned assurance. Anticipatory machinery generates a public impression of preparedness. Where the machinery is not connected to delivery, that impression is unearned, and the pandemic record is what unearned preparedness looks like when it is tested: a risk ranked first for twelve years, two exercises warning of inadequacy, and 82% of plans unable to meet the demands of any actual incident when the incident arrived. The people who bear the cost of that assurance are not the people who produced it.
Established The opportunity-cost question is unanswerable in the direction that matters, and the reason is a disclosure choice. No unit named in this brief publishes a current disaggregated budget. Without a cost, no cost-benefit statement about government foresight can be made in either direction — neither “this is cheap insurance” nor “this is a waste”. Both are asserted regularly and neither is currently supportable. Publishing a headcount and a budget line is the single cheapest thing any foresight unit could do to make its own defence possible, and none has done it.
This brief's own unresolved questions, stated as an obligation rather than a caveat. Four remain open and are recorded so that a future version can be checked against them. First, the Japanese realisation rate: a corrigendum exists, was not read, and the headline figure is therefore withheld rather than estimated. Second, whether the calibrated-language reading result transfers from a climate assessment to a national risk register is this brief's inference and has not been tested. Third, whether the uncertainty-communication trials transfer from advice documents to scenario sets is likewise untested, and the brief recommends the transfer while flagging that it is a recommendation and not a finding. Fourth, and largest, the leading diagnosis this brief now adopts — that non-performance is the equilibrium rather than a shortfall — comes from authors who state explicitly that their proposed remedy has no outcome evidence, and none has been located since. Adopting a diagnosis whose cure is unevidenced is a defensible position and an uncomfortable one, and the discomfort should be visible in the text rather than smoothed out of it.
11 · Civilizational implications
Established The strongest evidence in this entire subject is not an absence. It is the pandemic, and it is a finding about foresight rather than a gap in one. Pandemic influenza was the top-ranked non-malicious risk on the UK register consistently from 2008. A 2007 exercise warned that business continuity plans lacked cross-organisational coordination. A 2016 exercise found preparedness “currently not sufficient to cope with the extreme demands of a severe pandemic”. In February–March 2020 a cross-government review found 82% of plans unable to meet the demands of any actual incident. The audit office recorded that no specific plan existed for a disease with COVID-19's characteristics, and no detailed plans for many non-health consequences. Estimated lifetime cost of the response, as at July 2021: £370 billion.
Established The statutory inquiry's verdict is one sentence and it is the sentence this brief is organised around: “The UK prepared for the wrong pandemic.” The inquiry concluded that the country “was ill prepared for dealing with a catastrophic emergency”, identified fatal strategic flaws in risk assessment, groupthink among advisers, “labyrinthine” institutional complexity, insufficient ministerial challenge, devolved administrations that “simply copied” UK assessments, and a 2016 exercise finding never adequately acted upon. It made ten recommendations, moving away from single reasonable-worst-case scenarios toward a wider range, mandating UK-wide exercises at least every three years with reports published within three months, and requiring regular external red-teaming.
Established The same failure was found by the foresight community's own commissioned study, before the inquiry reported. Examining eight countries, it recorded that “despite pandemics being identified as a key issue in many foresight and other planning exercises, there was a failure to integrate, act or sustain attention with their implications not fully understood or integrated into policy”. That is the profession's own consultants describing its sharpest single test and reporting that everything downstream of the identification failed.
Established And the one cross-national quantitative test available went the wrong way. A 2020 study tested the 2019 Global Health Security Index — an explicit ex-ante preparedness score — against COVID outcomes across 36 OECD countries and found a negative moderate correlation, Spearman's rs = −0.41, p = 0.013. Carry the caveats: one index, one disease, n = 36, a mid-May 2020 snapshot preceding most of the mortality, and not a test of foresight units. It is nonetheless the closest thing that exists to a quantitative test of whether anticipatory capacity produces better outcomes at national scale, and the sign is negative. Read alongside the finding that seven leading governance indices correlate highly and are not interchangeable, the safest reading is that an index of anticipatory capacity is a weak instrument rather than that anticipation is harmful.
Established The identification worked. The readiness did not follow. That is not “we do not know whether foresight works” — it is the highest-profile natural experiment available returning a negative result on the specific mechanism foresight claims. Frontier And yet nothing in the record distinguishes foresight does not work from foresight works and governments ignore it, or from foresight is adopted for legitimacy and its non-use is the design. All three hypotheses remain consistent with everything above, and that — not ignorance — is the honest terminal position of this brief.
Established The general principle the case illustrates is larger than foresight. Seventy years of comparable megaproject data show cost overruns “high and constant”, with nine projects in ten over budget. Seventy years of expensive, visible, extensively documented forecasting failure produced no measurable improvement. An institution that cannot improve its five-year forecasts across seventy years of feedback has no demonstrated mechanism by which it would improve its fifty-year ones, and the burden of proof on anyone claiming otherwise is heavier than this literature has ever accepted.
12 · Timelines
These horizons track institutional reform and publication decisions, not technology:
- 10 yr: Established The UK's committed reforms land or do not: multiple scenarios replacing single reasonable-worst-case in the national risk assessment, eight standing independent expert advisory groups, a chronic-risks process, annual tier-one exercises through 2028. Frontier They are commitments, not results, and no evaluation plan has been published alongside them — which means the reforms may be as unmeasured as what they replace. Frontier Whether any government publishes retrospective scoring of a risk register, or aggregate accuracy figures from an internal forecasting tournament, is the variable most worth watching and the cheapest to move. Frontier And the 2025 closures of anticipatory bodies in one large jurisdiction either produce a measurable degradation in anticipation or they do not, which is the first natural experiment this field has been handed in decades.
- 25 yr: Speculative The plausible split is that the forecasting-tournament strand, which is scored by construction, accumulates a real evidence base while the scenario-and-report strand does not, because only one of the two has a criterion. Frontier If that happens, expect institutional pressure to reclassify foresight as forecasting, and expect the field to resist it on the grounds it has held since 2006. Frontier Expect also, on the record of every scored governance indicator so far, that the first published accuracy league table is gamed within three cycles.
- 50 yr: Speculative Either an agreed evaluation criterion exists and this brief is rewritten as an assessment, or the field completes a second thirty years without one and the correct description becomes institutional rather than epistemic. Handwave Which of those occurs is not forecastable from anything in the current record, and this brief declines to supply a probability for it.
- 100 / 250+ yr: Handwave Beyond useful forecasting — a limit this subject asserts about itself more insistently than any other in this map, and the assertion is the thing under examination. The one adjacent series long enough to test it, seventy years of megaproject cost and demand forecasts, shows overruns “high and constant” throughout, which is to say seventy years of visible, expensive, documented feedback produced no measurable improvement in forecasting accuracy at all.
13 · Technology tree & dependencies
- Depends on Nothing on this map. Foresight practice waits on no research result produced by another brief; its methods are decades old and its constraints are institutional. The absence of a dependency here is a finding, not an omission.
- Requires (not on this map) An agreed criterion for what would count as success, which the field's own quality literature states does not exist — not even agreement on the criteria by which quality could be assessed. Retrospective scoring of assessments already published with dates and probability bands: eighteen years of national risk registers and six years of an internal forecasting tournament, none of it ever marked. Disclosure of programme cost, unavailable currently for every unit named in this brief. A route from output to decision, whose absence is the one thing every independent reviewer of this machinery has found. And, added in this version because the previous four are not sufficient, a scoring design that survives being gamed — the governance-index record says that a legible, scored institutional target reliably produces the score instead of the institution, and a naive accuracy league table for foresight units would reproduce that failure exactly. All five are institutional; none is a research result, and none requires new method.
- Enables In principle, earlier warning for every brief with a long-horizon risk. No typed enabling edge is claimed, and on this evidence none can be — the enabling relationship is exactly what has never been demonstrated, and a typed edge to an unbuilt claim would be worse than no edge at all.
- Adjacent The forecasting and judgement literature, where the accuracy data lives; risk analysis; public administration; and within this map Scientific Advisory Institutions, Civilizational Planning and Long-Term Institutions.
14 · Common misconceptions & speculative claims
“The Singapore foresight body is one organisation.” Established It is two. The Centre for Strategic Futures (2009, Prime Minister's Office, whole-of-government, long-range) and the Risk Assessment and Horizon Scanning programme (July 2005, national security coordination secretariat, two-to-five-year horizon) are separate organisations with separate remits. “CSF's RAHS system” describes nothing that exists. Nor is the UK unit still a standalone Foresight programme: it was renamed and its foresight function is now shared with the Cabinet Office.
“The national risk register is the operational face of the foresight machinery.” Established It is produced by a different part of government, on different horizons, by different methods, with its own explicit non-prediction disclaimer, and this brief's use of it is a framing choice rather than an institutional fact. Commentary conflates the two routinely. And the two-year-horizon criticism is out of date — since 2025 the horizon is five years for non-malicious risks.
“Expert forecasters beat intelligence analysts by about 30%.” Established This is not a research finding. It traces to a newspaper opinion column of 1 November 2013, published when fewer than a fifth of that year's forecast questions had resolved, describing data never released. The vendor's own verifiable figure is a different comparison entirely — an aggregation method against a prediction market, not forecasters against analysts. State it as a claim, name its origin, or leave it out.
“Scenario planning let one oil major foresee the 1973 price shock.” Established The founding success story does not survive the archive. Peer-reviewed archival research finds the exercises responded to threats already visible — the producers' cartel had formed, the price negotiations were under way — and functioned as narrative world-building rather than prediction. The corporate study most cited in government foresight advocacy is, separately, corporate: 83 firms surveyed, matched to performance for 70 and 42, self-reported maturity on 35 Likert items, no control for research intensity, and profitable firms can afford foresight units. There is no government equivalent.
“Foresight is unevaluated because governments will not fund evaluation.” Established The base rate says otherwise, and the correction is important because it changes what should be inferred. The UK audit office found 8% of major-project spend carrying robust evaluation plans in 2019, 9 of the 108 most significant projects robustly evaluated, and 77 projects representing 64% of spend with no evaluation arrangements at all. Foresight is not unevaluated because it is philosophically awkward. It is unevaluated because almost nothing in government is evaluated, and the field's stated objections to outcome measurement are a rationalisation of a condition it shares with everything around it.
“Publishing more foresight, more openly, would build public trust in it.” Established On the experimental evidence set out in the frontier section, that is the one thing transparency reliably does not do. The aggregate trust gain from balanced, uncertainty-disclosing text over confident persuasive text is null to small, and the effect is a crossover rather than a main effect — gained among those who disagree with the evidence, lost among those who already agree, with no net polarisation. Transparency redistributes trust rather than increasing it. What it buys is reach into the sceptical part of the audience and insurance against the day the assessment turns out to have been wrong. Those are the right reasons to publish uncertainty; popularity is not one of them, and an institution that adopts scenario ranges expecting to be better liked will read its own evaluation as a failure.
“Calibrated probability language means readers get the message.” Established Readers do not. Tested on the closest available analogue, a calibrated vocabulary was read regressively in every term: very unlikely (defined below 10%) at a mean of 41%, unlikely (below 33%) at 44%, likely (above 66%) at 54%, very likely (above 90%) at 62%. The best presentation format tested raised consistency with the publishing body's guidelines only from 20.76% to 30.12%. The instrument tested was a climate assessment, not a risk register, so the transfer is an inference; but a register that communicates in five likelihood bands and a likelihood-impact matrix is using the same technology, and a Lords committee has independently warned that the matrix “may be misleading and lead to a false sense of confidence”.
“Scoring foresight units on accuracy would fix the problem.” Frontier It might, and this brief still recommends scoring — but the unqualified version of the recommendation is refuted by the best-documented case of scoring an institutional design. A governance index made institutional quality legible and comparable, and over 70 countries reorganised regulation around its indicators, consultancies were paid to move rankings, 111 countries swung more than 40 places in fifteen years, scores were altered under internal pressure for four named countries, and the index was cancelled. Separately, in the only bureaucracy-wide hand-coded study of delivery, an incentives-and-monitoring index is robustly negatively correlated with project initiation and completion across 4,700 projects while an autonomy index is robustly positive. The honest position is that measurement is necessary and that the naive form of it has a documented failure mode, and that designing a scoring regime that survives contact with the incentives it creates is an unsolved problem this brief does not solve.
“The people closest to a policy area forecast it best.” Established Not for the kind of question foresight asks. When 126 experts forecast the eleven-year results of a randomised institution-building programme, the mean prediction for physical infrastructure (0.218 SD) almost exactly matched the outturn (0.204), while the prediction for institutional change (0.095) missed the outturn (0.062) by about half — and the country's own policymakers, the people with the most contextual knowledge, predicted around 0.25, four times the truth. On a simple operational question about participation they predicted 42% against an actual 98%. Domain proximity helped on delivery and hurt on institutions.
“No evidence that foresight institutions change decisions means they do not.” Established It does not, and the distinction is the most important epistemic point in this brief. The clearest illustration comes from an adjacent institution rather than a foresight unit: the ombudsman for future generations established under Article P) of a 2011 constitution holds the strongest formal instrument any such body has anywhere, including the power to propose that matters be referred to the constitutional court — and its own descriptive account of itself contains no staff size, no annual case or investigation statistics, no referral counts, and no instance of a changed decision or law. A peer-reviewed treatment of the same office, which calls its role outstanding, contains no case numbers either. That is a missing denominator, not a measured zero. “No recorded effect” here means the record was not kept, and a brief that reported it as evidence of no impact would be making the advocates' mistake in the opposite direction. The durability of such offices is Long-Term Institutions' subject; what is borrowed here is only the shape of the measurement failure, which is identical.
“Foresight units fail because they lack capacity or seniority.” Frontier That is the field's own diagnosis and there is a better-supported rival. The isomorphic-mimicry account holds that reforms are adopted to secure external legitimacy and financing, and that non-performance is the equilibrium rather than a shortfall — organisations “constantly adopt ‘reforms’ to ensure ongoing flows of external financing and legitimacy yet never actually improve”. Every observation in this brief is consistent with it, including the 2012 review's admission that previous embedding attempts succeeded without enduring. Handwave What must be carried with that diagnosis is that its own authors' proposed remedy — problem-driven iterative adaptation — is presented in the founding paper with the explicit statement that no outcome evidence for it existed, and none has been located since. A well-evidenced diagnosis attached to an unevidenced cure is exactly the position this brief is in, and it should not pretend otherwise.