1 · Concept overview

Megaproject governance is the set of arrangements that decide which very large projects get approved, on whose forecasts, under what contracts, and with what consequence for anyone when the forecasts prove wrong. The conventional threshold is a project costing upward of a billion dollars — a convention this brief could not source to a primary definition and does not treat as more than a convention. The more useful definition is behavioural: a project large enough that the people who approve it will not hold office when it is judged.

The category is worth interrogating before the numbers are. A cost threshold is a unit of account, not a causal kind, and pooling dams, Olympic Games, IT programmes and tunnels produces a statistic about a heterogeneous bag unless the pooling is earned. The field's own most recent work cuts against the pooling: it now compares twenty-three project types and finds one of them qualitatively different from the rest. Whether type explains more outturn variance than size does is answerable with data that exists and appears not to have been asked.

This is the keystone brief of its category because nine other briefs on this map inherit its assumptions. Every corridor, barrier, transit system and continental scheme assessed elsewhere here is appraised by a sponsor, approved on a forecast, and delivered across a decade. If the iron law holds, the most relevant published finding about all of them is not in their own literatures but in this one.

What the brief covers. The outturn record and what the distributions actually look like; the fat-tail result and its real scope; optimism bias against strategic misrepresentation as competing explanations with opposite remedies; reference-class forecasting, its provenance and the one national test that measured it; which reforms have been tried and what happened to them; and the exotic end of the question — whether overrun is an error at all, or the designed output of a selection mechanism working exactly as built. Section 14 carries a finding about the literature itself that belongs in any honest treatment of it.

2 · Current scientific position

Established The field's central claim is that large projects overrun systematically rather than randomly, and the claim has been given the name of a law. The formulation appears in Flyvbjerg's What You Should Know about Megaprojects and Why: An Overview, Project Management Journal 45(2):6–19 (2014), whose deposited abstract sets out the whole programme: measurement of project scale, global megaproject spending, four “sublimes” that explain why megaprojects keep growing, the iron law itself as systematic budget and time overruns, a “break-fix model” offered as the explanation, a critique of Hirschman's Hiding Hand, and the “survival of the unfittest” phenomenon by which inferior projects get built. Frontier The slogan usually attached to the law — over budget, over time, under benefits, over and over again — does not appear in the abstract obtained for this brief. The substance is sourced; the exact wording is not, and it is quoted here as a slogan rather than as a citation.

Established The same paper puts global megaproject spending at $6 to $9 trillion annually, or 8 per cent of total global GDP. Established That is a spending-scale claim and nothing more. It is not an overrun, not a loss and not a waste figure, and the paper performs no multiplication. Frontier The arithmetic that yields “trillions wasted every year” is done downstream by citers, using an overrun percentage drawn from a different sample with a different composition, which compounds two unrelated bases into a number nobody has measured.

Established The evidentiary base under the law moved by roughly two orders of magnitude in two decades, and it recomposed as it grew. The origin paper is Flyvbjerg, Skamris Holm and Buhl, Underestimating Costs in Public Works Projects: Error or Lie?, Journal of the American Planning Association 68(3):279–295 (2002), covering transport infrastructure only. Frontier Its sample size and headline proportions — commonly quoted as 258 projects across twenty nations with costs underestimated in roughly nine of ten — could not be verified for this brief. No abstract is deposited for the paper and the text was not obtained, so those figures appear here as what is quoted, not as what is established. Established By 2021 the behavioural synthesis reports data from 2,062 projects, and by then the population is all project types rather than transport alone. Frontier A figure of roughly 16,000 projects is widely attributed to the 2023 trade book How Big Things Get Done; no bibliographic record for it was obtained here and this brief does not state it as fact. Established The consequence is mechanical and it governs everything that follows: a percentage quoted without its vintage is a percentage quoted without its population.

Established The Olympics are the best-documented single project class in the literature, and the numbers are worth having exactly. The Oxford Olympics Study 2016, by Flyvbjerg, Stewart and Budzier, reports that Games held over the previous decade each cost USD 8.9 billion on average; that average actual outturn cost across 1960–2016 is USD 5.2 billion for Summer Games and USD 3.1 billion for Winter Games at 2015 price levels; that the most costly Summer Games to that date was London 2012 at USD 15 billion and the most costly Winter Games Sochi 2014 at USD 21.9 billion; that at 156 per cent in real terms the Olympics have the highest average cost overrun of any type of megaproject; that cost overrun is found in all Games without exception, which the study says is true of no other megaproject type; that 47 per cent of Games have overruns above 100 per cent; and that the largest overruns are Montreal 1976 at 720 per cent and Barcelona 1992 at 266 per cent for the Summer Games, Lake Placid 1980 at 324 per cent and Sochi 2014 at 289 per cent for the Winter. Established Rio 2016 came in at USD 4.6 billion with a 51 per cent real overrun of USD 1.6 billion, which the study notes is the same as the median overrun for Games since 1999.

Established Every one of those figures rests on sports-related costs only. The study states plainly that wider capital costs for general infrastructure — which are often larger than sports-related costs — have been excluded. Established The same abstract also carries two different averages on two different bases within a few sentences of each other: USD 8.9 billion for the past decade, and USD 5.2 billion Summer against USD 3.1 billion Winter for 1960–2016. Established Both are correct; quoting either without its window is not. Frontier The base-rate problem that afflicts this whole literature is therefore visible inside a single abstract, before any citer has touched it.

Established The 2024 update reverses the 2016 study's headline reform finding, and the reversal is by the same authors. In 2016 the Olympic Games Knowledge Management Programme “appears to be successful in reducing cost risk for the Games”, with the difference in cost overrun before (166 per cent) and after (51 per cent) reported as statistically significant. Established In 2024 Budzier and Flyvbjerg report that Olympic costs are statistically significantly increasing, that prior analyses did not show this trend, and that cost overruns were decreasing until 2008 but have increased since. Established Paris 2024 is given at USD 8.7 billion at 2022 price levels with a 115 per cent real overrun — in the authors' words, not the cheaper Games that were promised. Established The Games remain, on their account, the only project type that never delivered on budget, ever. Established Venue reuse, the IOC's specific cost lever, did not have the desired effect for Tokyo 2020 and also looks ineffective for Paris 2024. Frontier Note carefully what this is: the 2016 result was correct on the 2016 window, and 2024 is a longer-window reversal rather than an error. It is a research programme self-correcting in public, which is a credit to it. It also means that anyone citing the 2016 reform finding today is citing a conclusion its own authors withdrew inside an eight-year window.

Established Cost outcomes in at least one project class are power-law distributed rather than normal. Flyvbjerg, Budzier, Lee, Keil, Lunn and Bester established this for IT projects in the Journal of Management Information Systems (2022). Established The cross-type follow-up, The Uniqueness of IT Cost Risk: A Cross-Group Comparison of 23 Project Types, states that IT cost risk is uniquely more risky than risk for other project types, and that IT is the only project type with a fitted tail exponent (alpha) of 1 or below, indicating infinite mean and variance. Established The scope of that result is the paper's entire point: it is asserted of IT against twenty-two other types, and the structure of the finding implies the other twenty-two are not in that regime. Frontier “Infinite mean and variance” is moreover a property of a fitted distribution rather than of realised expenditure, and real budgets are truncated by cancellation and insolvency long before any theoretical divergence could be realised. Frontier Discriminating a true power law from a heavy-bodied lognormal at achievable sample sizes is a known hard problem, and the confound is named in the statistical literature rather than being an outsider's objection.

Frontier The two candidate explanations imply opposite remedies, and the field's leading proponent does not hold the purely cognitive one. Flyvbjerg's 2021 inventory of the top ten behavioural biases in project management, selected from a field of more than 200 biases, lists them in order: strategic misrepresentation, optimism bias, uniqueness bias, the planning fallacy, overconfidence bias, hindsight bias, availability bias, base rate fallacy, anchoring, and escalation of commitment. Established Strategic misrepresentation is ranked first and is explicitly distinguished as political rather than cognitive. Established The same paper names base rate neglect as a primary reason projects underperform. Frontier The selection argument is separate and sharper. In Survival of the unfittest, cost–benefit analysis is held to be unreliable for these projects because of perverse incentives encouraging underestimation of costs: if error is biased rather than random, a competitive appraisal process selects the project whose promoter shaded hardest, not the project that was best. Frontier Cantarelli, Chorus and Cunningham formalise the same intuition as a signalling game, which converts overrun from a mistake into an equilibrium object with comparative statics.

Established Reference-class forecasting is the field's flagship remedy, and its own framing is revealing. Flyvbjerg's From Nobel Prize to Project Management: Getting Risks Right (2006) describes RCF as using historical performance data from comparable projects to improve accuracy, bypassing cognitive distortions rather than correcting them. Established The provenance is the inside-view / outside-view distinction of Kahneman and Tversky, and the Nobel of the title is Kahneman's 2002 economics prize; the specific 1979 primary citation could not be verified for this brief and is named rather than quoted. Frontier Where RCF has been mandated — the UK Treasury's optimism-bias uplifts, the Danish transport ministry, Swiss and Dutch applications — the actual uplift percentages could not be obtained here, because the government sources were unreachable. This brief prints no uplift figure. Frontier That gap is not decorative. The uplift is the method, and Zani and Adey's 2025 cross-validation of weighted against non-weighted RCF on real project data emphasises the importance of practitioners selecting appropriate percentiles. Frontier The percentile is where the number actually comes from, and the percentile is not itself reference-classed.

Established The single most useful finding in the topic is a dissociation, and most citers collapse it. Odeck, Welde and Volden, examining external quality assurance of cost estimates in the Norwegian road sector, report that quality assurance has led to a reduction in cost overruns, and that quality assurance has not, however, led to improved accuracy of the estimates made by authorities. Established The number got better; we did not get better at knowing the number. Frontier Two readings follow and the paper separates them better than its citers do: either the assurance regime worked as an incentive device against shading, or the approved baselines were inflated so that an unchanged outturn cleared a higher bar. Frontier Both readings are live, both are decidable with data that already exists in Norwegian project files, and neither has been settled. Everything in the governance debate turns on which one is right.

Frontier On the magnitude of the effect the evidence does not converge, and the honest terminal position is to declare the tie. The largest verified sample located for this brief is Dembélé's study of optimism bias in economic internal rates of return across 3,674 World Bank projects, which reports that the bias is real but moderate — materially weaker language than the field's headline framing. Frontier Nilsson's study of seven road and railway construction projects, honest about its own denominator, finds early cost estimates of limited utility for assessing project returns and concludes that institutional frameworks prevent meaningful learning from previous projects' implementation experiences. Frontier Park's before-and-after test of reference-class forecasting on 107 major projects exists and is exactly the right instrument; its direction could not be established for this brief and is therefore not characterised here. Frontier Odeck's econometric meta-regression across the reported literature exists precisely because reported magnitudes vary systematically between studies. Established The directional finding — costs come in high, more often than not, across countries, decades, methods and independent research groups — is robust. Frontier A single headline number is not available. Any treatment that hands you one without a vintage, a baseline convention and a scope boundary is handing you an artefact of those three choices.

3 · Frontier questions

Frontier The single iron law is being replaced from within by a taxonomy. The 23-project-type cross-group comparison moves the field from “megaprojects overrun” to “project types have distinguishable tail signatures and the differences between them are large.” That is a more useful object and a more fragile one: it invites the question of whether the category “megaproject” is a causal kind at all, or merely a cost threshold, which is a unit of account. Speculative If type explains more outturn variance than size does, pooling dams, Olympics, IT systems and tunnels produces a statistic about a heterogeneous bag. The 23-type dataset can answer this. As far as could be established here, it has not been asked.

Frontier Whether any governance reform has durably moved outturns is genuinely open, and the two best pieces of evidence point opposite ways. On one side, the Olympic reform verdict reversed within eight years, and venue reuse is a specific, falsifiable policy that was tested and produced a negative result for two consecutive Games. On the other, the Norwegian quality-assurance regime is a real national intervention, evaluated by an independent group, with a measured reduction in overruns. Frontier The reconciliation nobody has published is a baseline-inflation control: if approved budgets rose as fast as assurance tightened, the Norwegian result is a measurement artefact, and if they did not, it is the field's one durable win.

Frontier The live methodological seam in reference-class forecasting has moved from whether to which. Weighted RCF is now being cross-validated on real project data with explicit attention to percentile selection. Love's constructive alternative, decomposed contingency, treats RCF as a coarse instrument rather than a wrong one and proposes building contingency from decomposed risk elements instead of lifting it off a class. Frontier The two approaches make different bets about where the information is: RCF says it is in the outturn distribution and not in the project; decomposition says the reverse.

Frontier A revisionist position has entered mainstream project journals. Odeck — the field's most prolific independent replicator, working on Norwegian roads for two decades — publishes under the title Beyond the narrative of failure: Road project cost overruns are manageable. Love and Ahiaga-Dagbui argue that the cost-underestimation literature rests on what they call plausible untruths, definitional and base-selection artefacts rather than deception. Frontier Neither paper's text could be obtained for this brief, and both are named here as positions in print rather than summarised. That is itself worth stating: the strongest published counter-case to the field's headline claim sits behind a paywall that deposits no abstract, which is a small illustration of how the accessible version of a literature diverges from the literature.

Frontier Front-end escalation and post-baseline overrun are becoming two measurands rather than one. Welde and Odeck measure cost escalation in the front end of Norwegian road projects — the movement between an early sketch estimate and a formal baseline — separately from overrun measured after the baseline is set. Frontier Which of the two a study measures largely determines the size of the number it reports, and this is the mechanical source of much of the drift between papers. Speculative If most of what is called overrun turns out to be scope maturation between a sketch and a design, the entire governance agenda is aimed at the wrong stage of the process.

Frontier Two methodological turns are running in opposite directions at once. Flyvbjerg's 2024 work on heuristics elicits practitioners' tacit fast-and-frugal rules of thumb, framed through Aristotelian phronesis and drawn from leadership training and from Pixar. Frontier That sits awkwardly beside the outside view, which exists precisely because practitioner judgement was held to be unreliable. Whether the two are compatible is unresolved and worth watching. Frontier Meanwhile machine-learning overrun prediction is being claimed — one 2025 paper asserts a 15 to 25 per cent reduction in forecasting error against traditional methods, in a low-prestige venue and cited here as a claim rather than a result — against Flyvbjerg's own counterblast, published under the title AI as artificial ignorance. Speculative The unresolved question is whether a large model can add anything to a reference class, or whether it merely relearns the reference class with extra steps and less transparency.

Speculative Modularity is the most interesting live hypothesis about cause. If overrun is a function of how much of a project is one-of-a-kind, the cure is not better forecasting but a different unit of construction. Ansar and Flyvbjerg's platform analysis of space launch is the best available existence proof in the adjacent direction: costs fell far enough to increase demand for services in space. The 2024 Olympics paper points the same way, recommending the Games look to other megaproject types for how to generate positive learning curves. Frontier The warning attached is that venue reuse — modularity-adjacent, though not modularity — was tested and failed twice.

Frontier The Hiding Hand dispute remains open in print. Flyvbjerg and Sunstein's Malevolent Hiding Hand makes a frequency claim, not an existence refutation: on a larger sample the opposite of Hirschman's pattern is more common. Room's rejoinder and Flyvbjerg's reply appear in the same volume of World Development. Frontier The pivot of the argument is Hirschman's own sample size, which could not be verified for this brief — so the load-bearing number in the field's central intellectual dispute is, in the accessible record, unconfirmed.

4 · Technological bottlenecks

Frontier The binding bottleneck is not scientific. Take the target state seriously — large projects that come in on budget as a matter of course, meaning a median outturn within a stated tolerance of the approval estimate, a symmetric distribution, holding across jurisdictions and types for a decade — and ask what stands between here and there. Almost none of it is an unknown about the world.

Established The measurement layer is missing first. There is no adopted specification fixing baseline date, scope boundary, deflator and treatment of scope change, such that two independent teams scoring the same project agree within a stated tolerance. Frontier This is not hard science; it is hard politics. Every jurisdiction's definition is entangled with its own accountability regime, and a common definition would make cross-jurisdiction comparison — and therefore blame — possible for the first time. Established Nilsson's finding that institutional frameworks prevent learning from previous projects points at exactly this, as does Winch's observation that megaproject stakeholder instruments are largely instrumental and fail to address the stakes of the natural environment and of future generations, which are precisely the constituencies that cannot allocate blame.

Frontier The second bottleneck is data access rather than method: resolving how much measured overrun is scope maturation between sketch and design, and how much is true delivery escalation, requires paired estimates at both stages for a large sample, and pre-baseline internal estimates are rarely archived. Speculative If the answer is mostly the former, most of the governance literature is aimed at the wrong stage.

Frontier The third is escalation of commitment as a structural feature. Drummond's account is that once megaprojects are underway, decision-makers face mounting pressures to persist regardless, and the pressure to reinvest in economically poor megaprojects is the specific pathology. Established A partly built megaproject is close to uncancellable, and every promoter bidding for approval knows it, which makes the first tranche of spending an option purchase on the rest.

Frontier The genuinely scientific unknown is narrower than the field's rhetoric implies. Establishing tail exponents with confidence intervals, principled minimum-cutoff selection and a lognormal likelihood-ratio comparison, per project type, is real statistics with a real difficulty. Frontier But it affects tail pricing, insurance and portfolio construction. It does not touch the central governance question, which is whether forecasting is the problem at all.

5 · Research dependencies

Frontier Progress depends on outturn databases that are public and complete rather than curated by sponsors, because reference-class forecasting is only as good as its reference class and the class is currently assembled by interested parties; on ex-post evaluation funded and published as a matter of course rather than commissioned when something has gone visibly wrong; on paired front-end and post-baseline estimates being archived rather than discarded, which is a records-management decision that forecloses a research question for decades when it goes the wrong way; and on some mechanism carrying accountability across the decade between approval and outturn. Established That last one is a constitutional design problem rather than a project-management one, and it is the reason this brief sits in an infrastructure category rather than a methods one.

Frontier On this map the dependency runs outward rather than inward. Energy corridors, intercontinental rail systems, water infrastructure megaprojects, high speed transit networks and underground cities are all delivery-limited rather than capability-limited on this map's own assessment, and each inherits whatever this brief concludes. Frontier Infrastructure resilience and civilization resilience planning inherit something sharper: if appraisal systematically selects the worst-forecast project, then a resilience portfolio assembled through ordinary appraisal is selected on the same criterion.

Frontier Two dependencies run the other way and are worth naming because they are unusual. The first is on research-integrity infrastructure: a citation graph that propagates retractions and supersessions would have removed one canonical paper from circulation in 2012 and flagged a reform finding as withdrawn in 2024, and neither happened. Speculative The second is on the archival practice of individual project offices, which is nobody's research agenda and quietly decides whether the front-end question is answerable at all — a pre-baseline estimate discarded at project close is a datum that cannot be recovered at any price later.

6 · Required experiments

Frontier The workback chain to on-budget delivery is seven links, and the striking thing about it is how few require a discovery.

Established One: a standardised outturn measurement protocol. Publish and adopt a specification fixing baseline date, scope boundary, deflator and scope-change treatment. Measure by inter-rater agreement: two independent teams score the same project set and agree within a stated tolerance. Binding constraint institutional; nearest existing instrument is Odeck's meta-regression. Frontier Two: decompose front-end escalation from post-baseline overrun. Paired estimates at sketch and baseline stages for a large sample; report the split. Binding constraint is archival access, not method. This link is early because if it comes back mostly front-end, links four to six change shape.

Speculative Three, and the highest-leverage unrun study in the field: the discriminating test between error and misrepresentation. A two-by-two: competitive against non-competitive approval, crossed with externally set against promoter-set uplift, with outturn as the dependent variable and difference-in-differences as the estimator. Frontier The variation already exists in the world — single-promoter utility investment against competitive transport appraisal, external assurance regimes against self-assessed ones. Nobody has assembled it. There is no scientific barrier whatsoever, only the fact that assembling it requires several governments to hand over comparable data about their own decisions.

Frontier Four: distributional characterisation per project type — tail exponent with confidence interval, principled minimum cutoff, lognormal likelihood-ratio test, out-of-sample replication on a disjoint set, for each of the 23 types. This is the one genuine scientific unknown in the chain. Speculative Five: an incentive-compatible approval mechanism, under which a promoter's optimal disclosed forecast equals its private expectation. The measurement is revealed truthfulness — divergence between internal and disclosed estimates falling toward zero. The mechanism-design content is standard; the blocker is that public bodies cannot easily impose personal downside on promoters, and the promoter is frequently the same body that appraises.

Speculative Six: a repetition index, measurable across types, with a demonstration that overrun falls monotonically in it controlling for scale. Frontier Seven: a decade-long, multi-jurisdiction outturn panel with a baseline-inflation control — the study without which the claim that no reform has ever worked cannot be falsified. Established No such study was located. It is arguably the missing study in the field, and its binding constraint is temporal and institutional: it requires standing data infrastructure that survives electoral cycles.

7 · Engineering requirements

Established The engineering requirements here are informational, not physical. Nothing in this brief waits on a machine. The enabling infrastructure is outturn data collection at project close, in a schema comparable across projects and jurisdictions, and it is mostly absent — not because it is difficult to build but because nobody is obliged to populate it.

Frontier A workable specification has four fields that the literature shows are load-bearing and that current practice leaves to the analyst: the baseline date and which decision it corresponds to; the scope boundary, including whether general infrastructure attributable to the project is inside or outside it, which is precisely the boundary the Olympic figures draw around sports-related costs; the deflator and its price-level year; and the treatment of authorised scope change, which is the mechanism by which an overrun can be reclassified out of existence.

Established Beyond the schema: probabilistic rather than point estimation, so that an approval carries a distribution and not a number; earned-value measurement that cannot be gamed by scope reclassification; and archival retention of pre-baseline internal estimates, which costs almost nothing and is the only way link two of the workback chain ever gets run. Frontier Modularity belongs here as a design objective rather than a governance one: a project assembled from many repeated units has a reference class of its own and a learning curve, which is the proposed mechanism behind the observation that repetitive construction behaves better than singular construction. Speculative Whether that mechanism survives measurement against a repetition index is untested, and the one modularity-adjacent intervention that has been tested — Olympic venue reuse — failed twice in a row.

Frontier One requirement is easy to overlook because it is a negative: the schema must be resistant to the two moves that make an overrun disappear on paper. Rebaselining converts an overrun into a new plan, and authorised scope change converts it into a different project. Established Both are legitimate operations. Both are also the exact operations a sponsor performs when the alternative is reporting a failure, which is why the treatment of each has to be fixed in the specification rather than settled case by case by the party being measured.

8 · Adjacent technologies

Public procurement law and contract theory, which supply the instruments any reform has to be written into; mechanism design, which is where an incentive-compatible approval rule would come from and which the field has barely drawn on; behavioural economics, particularly the anchoring, planning-fallacy and outside-view literature the whole programme descends from; extreme-value statistics and the power-law-versus-lognormal discrimination problem, which is the technical core of the tail dispute; and cost engineering and probabilistic estimation as practised.

The subject also has an unusual adjacency to bibliometrics and research integrity, for reasons section 14 sets out: a literature about institutional accountability turns out to be a good place to observe how citation practice does and does not self-correct. A retraction that has not propagated in fourteen years and a headline finding withdrawn by its own authors are both, in this field, objects of study rather than embarrassments.

Two further adjacencies are underused. Auditing and public accounts scrutiny is the one institution that already produces outturn data at scale, in a format nobody has harmonised; and insurance and surety underwriting is the one commercial actor with a direct financial interest in the tail of the cost distribution, and therefore the one party that prices these projects without a stake in their approval. Neither literature talks much to this one.

On this map the governance problem recurs wherever scale does. See automated construction systems, whose productivity case rests on organisational structure rather than machinery; historical megaprojects for the outturn record before the databases existed; technology forecasting for the same inside-view problem in a different domain; arcologies for the case where the whole capital cost precedes the first revenue; northern development corridors and smart cities, whose documented failures are largely of this kind; and future public administration for the machinery that would have to carry any of it.

9 · Institutional requirements

Established This brief is institutional requirements throughout, so the section states the uncomfortable part plainly. The reforms with the best evidence — independent forecasting, published outturn ranges, reference classes, real cancellation options at defined gates — all have the same effect: they make a project look more expensive and less certain at the moment approval is sought. Frontier That is not a side effect to be engineered away. It is the mechanism. A governance reform in this field works by reducing the number of projects that clear the bar, which means its beneficiaries are diffuse and future while its costs are concentrated and immediate.

Speculative The minority reading is that this is why the instruments survive. On that account, stage gates, assurance reviews and uplift schedules persist despite weak outturn effects because their real function is not to reduce cost but to distribute accountability for a bad outcome that everyone already expects. Nilsson's finding that institutional frameworks prevent learning, Drummond's account of persistence regardless, and Winch's finding that the instruments systematically omit the environment and future generations all fit it. Frontier The Norwegian quality-assurance result is the counterexample and it is a real one. Speculative The test nobody has run is whether reform adoption tracks litigation and audit exposure rather than measured cost performance. National reform chronologies exist; the comparison does not.

Established There is a sharper illustration available, and it is inside the literature rather than in its subject matter. The canonical paper on transport demand-forecast inaccuracy — Flyvbjerg, Skamris Holm and Buhl, Inaccuracy in Traffic Forecasts, Transport Reviews 26(1):1–24 (2006) — carries a formal retraction in the Crossref registry, dated 21 February 2012. Established It continues to be cited as canonical. Frontier A body of scholarship whose central subject is what happens when nobody is accountable for a stale forecast contains, at its own foundation, a paper whose withdrawal fourteen years ago has not propagated through its citation graph. Section 14 sets out exactly what the registry records and what it does not.

10 · Ethical & societal considerations

Frontier Systematic underestimation is a transfer, not merely an inaccuracy. The difference between forecast and outturn is paid by taxpayers or users who were never shown the real number, and it is paid to whoever the project employs, finances and enriches. Established Whether to call that error or misrepresentation is a substantive empirical claim rather than a courtesy, and the field's own leading taxonomy ranks strategic misrepresentation first and classifies it as political. Frontier The charitable reading is available and defensible; it is not the default.

Established Distribution compounds it. Megaprojects concentrate benefits on the corridor or city served and costs on the general fund, and the people displaced are rarely among the beneficiaries. Established Winch's finding that megaproject stakeholder management is largely instrumental — and in particular fails to address the stakes of the natural environment and of future generations — identifies the two constituencies with the largest exposure and the least standing, and they are systematically absent from the instruments rather than accidentally so.

Speculative Against all of which sits Hirschman's defence, which deserves better than the dismissal it usually gets. The principle of the Hiding Hand holds that planners' ignorance of difficulty is providential: it induces them to start projects they would otherwise refuse, and human ingenuity then rises to the unforeseen problems. Some projects that would never have survived honest forecasting are now regarded as obviously worth having. Frontier Flyvbjerg and Sunstein's answer is a frequency claim — on a larger sample, the opposite pattern is more common — not a refutation, and the dispute is live in print. Frontier A mechanism that fires rarely is a bad thing to design a system around; a mechanism that fires sometimes is a real cost of designing the system to prevent it. Both halves of that sentence are true and the field has not priced the second.

11 · Civilizational implications

Frontier If the selection account is right, the aggregate consequence is not wasted money but a systematically wrong portfolio. Humanity would be building the projects that were best at surviving approval rather than the projects that were best, and it would have been doing so for as long as competitive appraisal has existed. Speculative That is a claim about civilisational capital allocation at the scale of several per cent of global output, and it is testable in principle by the link-three experiment nobody has run.

Speculative The upside case is worth stating because it is unusually cheap. If reference classes were genuinely mandatory, uplifts genuinely external, and outturns genuinely public, the same capital would plausibly deliver materially more — and no new technology would be required to achieve it. That is rare on this map: a large civilisational gain available from an institutional change rather than a scientific advance. Frontier It is also the reason the pessimistic reading matters. If overrun is an equilibrium rather than an error, the gain is not available at all, because realising it requires the people who benefit from the current arrangement to end it.

Handwave The deepest version of the pessimism is that the shading is load-bearing for civilisation-scale ambition — that a species which forecast its largest undertakings accurately would attempt fewer of them, and that the record of what got built under optimistic forecasts is the argument for the practice rather than against it. Frontier It is unfalsifiable in the form usually stated, and it becomes tractable only if someone can price the counterfactual value of the projects that were never proposed because their sponsors forecast honestly. Nobody has tried.

12 · Timelines

These horizons track institutions and data rather than capability. Nothing in this brief awaits an invention, which makes the timelines unusually pessimistic rather than unusually optimistic: where a forecast rests on a discovery, one can at least imagine the discovery arriving early, and where it rests on several governments agreeing to be measured against each other, one cannot. The dates below are therefore best read as the earliest points at which the relevant question could be answered by someone who wanted to, not as predictions that anyone will.

  • 10 yr: Frontier reference-class methods nominally standard in most large public procurement, and nominal is the operative word; outturn databases still fragmentary and largely sponsor-curated; the aggregate overrun rate substantially unchanged; the 23-type tail taxonomy consolidated or abandoned. The cheapest available advance in the whole field — a published measurement protocol with demonstrated inter-rater agreement — is achievable inside this window by one well-funded group.
  • 25 yr: Frontier enough jurisdiction-level natural-experiment data to settle the error-versus-incentive question, if anyone assembles it; the baseline-inflation control run against the Norwegian result, deciding whether the field has one durable win or none. Speculative Expect the answer to be unwelcome to whichever side loses and to be absorbed slowly.
  • 50 yr: Speculative plausible that modularity displaces singularity as the default construction unit for large infrastructure — on cost grounds rather than governance ones, with the learning curve doing the work reform could not. Speculative If that happens, the iron law will look less like a law and more like a property of a construction paradigm that ended.
  • 100 / 250+ yr: Handwave no basis for forecasting institutional arrangements on this horizon. The only relevant datum is that the pattern has now survived every reform aimed at it within the documented record, which is either a strong prior or a small sample, depending on whether you think the last seventy years were a fair test. Handwave The interesting far-horizon possibility is not that the problem gets solved but that it gets dissolved — that the category of the bespoke civil megaproject stops existing, the way bespoke shipbuilding largely did, and the question becomes a historical one about a particular way of organising construction.

13 · Technology tree & dependencies

  • Depends on Public and complete project outturn databases with an adopted measurement protocol; archived paired front-end and post-baseline estimates; mandatory independent ex-ante forecasting; funded and published ex-post evaluation; contract forms that price scope uncertainty rather than assuming it away; and the statistical work to characterise tail behaviour per project type rather than in aggregate.
  • Requires (not on this map) Reference-class forecasting is only as good as its reference class, and the class is at present assembled by the same sponsors whose projects populate it; ex-post evaluation is commissioned when something has gone visibly wrong rather than funded as a matter of course, which is a sampling rule that guarantees the record is unrepresentative. The second requirement is the one the Norwegian dissociation encodes — external quality assurance reduced overruns without improving the accuracy of authorities' own estimates, so what moved the outcome was who set the number and not what anyone knew — which makes independence of the forecaster, not the quality of the forecast, the governing variable. The deepest requirement is the one section 9 refuses to soften: every reform with real evidence behind it makes a project look dearer and less certain at the moment approval is sought, so it needs a mechanism carrying accountability across the decade between approval and outturn. That is a constitutional design problem rather than a project-management one, and no jurisdiction has solved it.
  • Enables Every delivery-limited slot in this category inherits this brief directly — energy corridors, intercontinental rail, water infrastructure, high-speed transit, underground construction and northern corridors are all appraised by sponsors on forecasts this literature is about. Beyond the category, it enables honest portfolio selection at civilisational scale: the ability to say which large undertakings are worth starting is downstream of the ability to know what the last ones actually cost. It also enables something narrower and more immediately useful: an insurable, priceable tail for large public projects, which does not exist today because nobody can characterise the distribution per type.
  • Adjacent Public procurement law and contract theory; mechanism design, which is where an incentive-compatible approval rule would have to come from and which this field has barely drawn on; behavioural economics and the outside-view literature the whole programme descends from; extreme-value statistics and power-law estimation, including the power-law-versus-lognormal discrimination problem that decides the tail claim; cost engineering and probabilistic estimation as practised; public audit, which already produces outturn data nobody has harmonised; and, unexpectedly, research integrity and bibliometrics, which section 14 shows to be a live adjacency rather than a decorative one.

14 · Common misconceptions & speculative claims

Established “Traffic forecasts are wrong by X per cent — see Flyvbjerg, Skamris Holm and Buhl 2006.” That paper is retracted. Inaccuracy in Traffic Forecasts, Transport Reviews 26(1):1–24 (2006), DOI 10.1080/01441640500124779, appears in the Crossref registry under the title RETRACTED: Inaccuracy in Traffic Forecasts, carrying an updated-by relation of type “retraction”, labelled Retraction, with an update date of 21 February 2012, pointing to DOI 10.1080/01441647.2012.664813, which resolves to a retraction notice by David Banister in Transport Reviews 32(2):265. The registry's retraction assertion is sourced from Retraction Watch. Frontier A retraction is not a finding of falsity. The reason for this one could not be established from anything reachable for this brief; the notice itself was not read. Retractions for duplicate publication or procedural defect are common and impeach nothing substantive, and nothing here asserts the paper's findings are wrong. Established What can be said is narrow and sufficient: the paper is formally withdrawn, it should not be cited without the retraction, and it is still cited without it. Frontier This matters because the benefit-shortfall limb of the iron law — the “under benefits” half — leans substantially on demand-forecast inaccuracy, and this is the canonical demand-forecast-inaccuracy paper. Established The substitute citation is the unretracted companion from the same year: Næss, Flyvbjerg and Buhl, Do Road Planners Produce More ‘Honest Numbers’ than Rail Planners?, Transport Reviews 26(5):537–555 (2006), which stands and covers overlapping ground. Speculative The open question, straightforward to answer and not answered here, is a citation audit: how much of the field's benefit-shortfall apparatus propagated into secondary and tertiary sources before 2012 and was never re-sourced.

Established “Reform fixed the Olympics.” The 2016 finding — overrun of 166 per cent before the IOC knowledge programme and 51 per cent after, reported as statistically significant — is real and was correctly reported on its window. It has since been superseded by the same authors, who in 2024 report costs statistically significantly increasing and overruns increasing since 2008, explicitly noting that prior analyses did not show this trend. Established Citing the 2016 reform result today is citing a withdrawn conclusion.

Established “The Olympics run 156 per cent over budget.” The 156 per cent is an average, in real terms, across 1960–2016, computed on sports-related costs only. The study states that wider capital costs for general infrastructure are often larger than sports-related costs and have been excluded. Frontier So the claim simultaneously understates total public exposure and overstates the coverage of the statistic, which is an unusual way for a number to be wrong in both directions at once.

Established “Megaprojects have infinite-variance cost risk.” The tail-exponent result is asserted of IT specifically, as the property distinguishing IT from twenty-two other project types in a cross-group comparison. Generalising it to megaprojects inverts the paper's central finding. Established And infinite mean and variance is a property of a fitted power law, not of money that anyone spent.

Established “Six to nine trillion dollars a year is wasted on megaprojects.” That figure is total megaproject spending, roughly 8 per cent of global GDP, from the 2014 overview. It is not an overrun, a loss or a waste figure, and multiplying it by an overrun percentage drawn from a different sample compounds two unrelated bases.

Frontier “Nine out of ten projects go over budget, and here is the number.” The origin of that proportion is the 2002 Journal of the American Planning Association paper, for which no abstract is deposited and whose sample size and headline proportions could not be verified for this brief. Frontier Meanwhile the largest verified sample located here — 3,674 World Bank projects — describes the bias as real but moderate. The strong version and the large-sample version are not obviously the same claim, and reconciling them is real work that has not been published.

Frontier “The database is one database.” It is a growing, recomposing collection: transport projects only in 2002, 2,062 projects of all types by 2021, and a figure of roughly 16,000 commonly attributed to the 2023 trade book but not verifiable here. Established Composition changes with size. A percentage from one vintage does not carry to another, and most citation practice in this field ignores that completely.

Frontier “Reference-class forecasting is a validated corrective.” Its own literature reports that the uplift depends on an analyst-chosen percentile; that the one verified national quality-assurance regime improved outturns without improving accuracy; and that a 107-project before-and-after test was thought necessary at all, which is not what a settled method looks like. Frontier Treat RCF as a promising and partially evidenced device rather than a validated corrective. Speculative The most interesting hypothesis about it is that its causal power is political rather than statistical — an externally supplied uplift is hard for a promoter to argue down, so it works as a commitment device against shading and its numerical accuracy is incidental. That predicts RCF's benefit vanishes where the promoter sets the uplift and persists where an external party does, independent of accuracy. It is a clean two-by-two and nobody has run it.

Established “Cost overrun is a well-defined quantity.” It is not, until you fix a baseline. Front-end escalation and post-baseline overrun are separately measurable and separately reported, and an econometric meta-regression across the literature exists precisely because reported magnitudes vary systematically across studies.

Speculative “Hirschman was refuted.” Flyvbjerg and Sunstein make a frequency claim on a larger sample, not an existence refutation, and Room's rejoinder and Flyvbjerg's reply sit in the same volume of World Development. Frontier The pivot — Hirschman's own sample size — could not be verified for this brief, which is a peculiar position for the field's central dispute to be in.

Speculative And the claim this brief will not resolve. The exotic hypothesis is that overrun is not a forecasting failure at all but a stable political equilibrium: promoters compete for scarce approval, approval is granted on forecast merit, forecasts are cheap to shade, and therefore the winning forecast is the most-shaded one and overrun is the designed output of the mechanism working as built. Frontier The strongest case for it is that the field's own taxonomy ranks strategic misrepresentation first and calls it political; that cost–benefit analysis is held unreliable because of perverse incentives encouraging underestimation; that the phenomenon has been formalised as a signalling game; that external assurance moved outturns without moving accuracy, which is what you expect if the binding constraint was incentive rather than knowledge; and that Olympic cost overrun is found in all Games without exception, a distribution with no left tail being hard to generate from unbiased error. Frontier The strongest case against is that the most-cited direct challenge holds the underestimation literature to rest on definitional and base-selection artefacts rather than deception; that the largest sample calls the bias real but moderate, where strategic capture should be large where it operates; that front-end scope maturation offers a mundane alternative explanation; and that overruns respond to ordinary management, which an equilibrium account predicts they should not. Speculative The falsifier is specific: find a regime with competitive approval, high stakes and symmetric forecast error, and the equilibrium account dies. Handwave Until someone runs the two-by-two, the honest verdict is a declared tie between an error theory and a design theory that predict identical data at the point of decision and opposite responses to every reform anyone has proposed.