1 · Concept overview

This brief is about measuring individual and population wellbeing — life satisfaction, affect, mental health, flourishing — and about whether those measurements are good enough to be targeted by policy rather than merely described. It sits with two others on this map where the same structure appears: a number is standing in for a thing, and the number has acquired institutions the thing never had. Here the thing is wellbeing; in Human Development Metrics it is development; in Reputation Economies it is trustworthiness. In all three the measurement outran the validation, and in all three the interesting evidence is about what the measurement did rather than what it captured.

The framing under test — that wellbeing can be measured well enough to be a policy target — contains two claims that are routinely run together and that fail in different ways. (a) The numbers mean something, are comparable across people, and respond to real changes. (b) Those numbers can be an allocation instrument, embedded in appraisal, capable of changing a decision. Frontier Claim (a) is contested at the foundations by a live technical dispute most practitioners route around rather than resolve. Frontier Claim (b) has an infrastructure — a monetised unit, official Treasury guidance, distributional weights, a wellbeing budget, a national screening tool — and an almost empty record of decisions demonstrably changed.

The unit of analysis fixes the boundary with the neighbouring brief and is worth stating at the top. Everything here is person-level: a scale a person answers, a domain a person is sufficient in, a satisfaction point a person gains. Bhutan's Gross National Happiness belongs here, despite its national headline, because its arithmetic runs over persons. Country-level composites — the HDI, the Multidimensional Poverty Index, the SDG indicator framework, the Doing Business record — belong next door.

2 · Current scientific position

Established Start with the objection, because it is the spine and it has not been withdrawn. Bond and Lang's Journal of Political Economy paper makes a mathematical claim that nobody disputes: reported happiness is an ordinal categorical response, comparing group means requires cardinalising the scale, and there are infinitely many arbitrary cardinalisations, each producing a different set of means. A group-mean ranking is invariant to cardinalisation only if one group's distribution first-order stochastically dominates the other's, which cannot generally be established from discrete survey categories. Established What makes it bite is the empirical half: applying simple monotonic lognormal transformations reverses headline findings in the literature, and they name the parameter at which each flips. The US income-happiness time series reverses at c = −0.68. The unemployment-inflation misery tradeoff, usually weighted 1.73:1, becomes 1:1 at c = 0.375. The finding that parents are less happy reverses at about c = −0.74 for men and −0.64 for women. The U-shape of happiness in age disappears, and some countries become monotone.

Established Their test of the conditions is the number this brief keeps returning to. Non-parametrically, valid rank-order comparison requires that one group never reports the lowest category and the other never the highest — a condition the authors say will never be satisfied in practice. Parametrically, ordered probit under normality requires equal happiness variance across groups. Across nine prominent literatures — Easterlin, age profiles, the unemployment-inflation tradeoff, country rankings, Moving to Opportunity, marriage, children, the female happiness decline, disability — the non-parametric conditions are satisfied in not a single case, and in the eight cases where equality of variances is testable, equality is rejected in every one. Their verdict is that happiness rankings from survey data are “essentially uninformative” absent assumptions they consider unjustifiable. Zero out of nine, and nought for eight.

Frontier The rebuttal is real, and it does different work than it appears to. Plant's Oxford Wellbeing Research Centre working paper argues, first, that Bond and Lang show only that if people use a strongly non-linear reporting function the results reverse, and that they “do not provide evidence to support their claim that individuals do use a (strongly) non-linear reporting function”; he cites Kaiser and Vendrik as finding the required non-linearity implausible for almost all variables of interest. His supporting evidence is indirect — correlations with height, where verbal labels sit on the numeric scale, homoscedasticity tests — and his vignette result is that the largest adjustments produce an average cardinal change of about 0.2 on a 5-point scale, with a worst case in which women are reclassified as less satisfied than men. Speculative His second argument is theoretical: a Grice-Schelling claim that people cooperatively interpreting a subjective scale converge on a linear, comparable one because that is the focal point that makes communication work. That is a plausible story about linguistic coordination with no direct measurement behind it. Frontier Read the exchange precisely: the mathematics is undisputed, and what is disputed is whether real reporting functions are non-linear enough for it to matter. The field's answer is “probably not, and here is indirect evidence.” That is a reasonable answer and it is not a resolution. Note the standing on both sides — Plant is at a wellbeing research centre and founded an organisation advocating that resources be allocated by wellbeing measures; Bond and Lang are labour economists with no organisational stake in the answer.

Frontier And the strongest evidence that the reporting function is not stable comes from Plant's own institution. Harrison, at the same centre, proposes that the scale itself stretches over time — the ruler changes. His evidence: the estimated effects of life events on reported satisfaction declined by roughly 35 to 40 per cent between 1991 and 2022 in German panel data. If underlying happiness rose while the scale expanded, measured effects would shrink in exactly that way, and his conclusion is that latent happiness could be up to 50 per cent higher than reported happiness from scale expansion alone. Established This creates a tension the field has not resolved and should not smooth over: you cannot use Plant to dismiss Bond and Lang and simultaneously use Harrison to dissolve the Easterlin paradox. They point in opposite directions on the stability of the reporting function.

Frontier The Easterlin paradox at fifty is still contested, and the contest has moved. Oparina, Clark and Layard, on Gallup World Poll data across more than 150 countries for 2009–2019, report a cross-sectional individual coefficient of about 0.4 on log income — doubling income raises Cantril-ladder wellbeing by roughly 0.3 points out of 10. Frontier The crude cross-country GDP coefficient of 0.6 falls to insignificance once health, social support, freedom, corruption and generosity enter, except in low-income countries; and in rich countries income growth shows no time-series correlation with happiness growth. They locate the disagreement with Stevenson and Wolfers in the omission of the social mediators. Frontier Standing matters here and runs with the finding: Layard is the leading academic advocate of wellbeing-based policy and a founder of the World Happiness Report, and a result saying income does not buy happiness in rich countries is congenial to that programme.

Established Against all that, the institutional half of the record is unusually concrete. HM Treasury's supplementary Green Book guidance defines a WELLBY as a one-point change in life satisfaction on a 0–10 scale, per person, per year, and prices it at £13,000 in 2019 prices, range £10,000 to £16,000, to be uprated to the appraisal base year, with an illustrative distributional weight of 2.4 for bottom-quintile recipients. Established Its status is load-bearing and usually dropped: it is supplementary guidance, used within the Green Book five case model “alongside existing welfare estimation methodologies”. It is not a replacement for cost-benefit analysis and it is not mandatory as the primary appraisal metric. Established The Treasury's own caveats are candid to the point of being against interest: monetise only where there is “high confidence” in causal estimates; not all wellbeing impacts are proportionate to monetise; adaptation means some improvements do not persist; and utility misprediction means people's experience differs from their forecasts.

Established The usage record is where the framing struggles, and the detail is the finding. The What Works Centre for Wellbeing — the body established precisely to build this evidence base — presents as its “first practice example” of WELLBY use a charity's cost-benefit analysis of its own Mentalization-Based Therapy programme, run 2019–2021, described on the platform as “an initial approximation”. The same page states that it is “not always possible to build credible counterfactuals”, that children's outcomes could not be easily quantified and were excluded from monetisation, and that there is “a lack of consensus on the monetary value of the life satisfaction of children”. Established The What Works Centre for Wellbeing ceased operations on 30 April 2024. Five years after a government created a monetised wellbeing unit with an official price, the flagship worked example on the sector's own knowledge platform is a charity evaluating itself, and the centre that curated it has closed.

Established New Zealand is the largest process experiment, and the independent account is more equivocal than the official one. Budget 2019 organised spending around five wellbeing priorities — mental health with emphasis on the under-24s, child wellbeing, Maori and Pasifika aspirations, building a productive nation, and transforming the economy to low emissions — selected through the Treasury's Living Standards Framework across four capitals: financial and physical, human, natural, social. Ministers had to show how initiatives served the priorities, and Cabinet committees assembled cross-agency packages using wellbeing analysis. The largest named allocation is NZ$455.1 million for primary mental health services. That is the Treasury describing its own budget, which is maximal interest. Frontier A Journal of Public Policy case study built on semi-structured interviews with 22 key informants across Treasury, ministries, politicians and academics — 45 minutes to two hours each, conducted in 2021 — records that adoption was driven by politics, internal direction and the international policy environment; that Treasury guardians disagreed among themselves on implementation and measurement; that the Treasury initially viewed the framework as having “not gone anywhere,” with significant staff alienation; and that internal debate over how to apply it ran roughly eight years. Speculative The retrievable portion of that study did not contain its conclusions on whether allocations changed substantively, so it is cited here for the process account and not for a verdict.

Established Bhutan's Gross National Happiness is a real instrument and deserves to be reported as one. It is built on the Alkire-Foster method with the Oxford Poverty and Human Development Initiative: nine domains — psychological wellbeing, health, time use and balance, education, cultural diversity and resilience, good governance, community vitality, ecological diversity and resilience, living standards — over 33 weighted indicators. A person counts as happy at sufficiency in at least 66 per cent of weighted indicators, and the population splits into deeply happy (77–100 per cent), extensively happy (66–77), narrowly happy (50–66) and unhappy (below 50). Thresholds are concrete: six years of schooling; Nu. 32,951.27 of annual household per capita income. Established Results: the index ran 0.743 in 2010 to 0.756 in 2015 to 0.781 in 2022; the 2022 survey covered 11,440 people aged 15 and over with a 96.6 per cent response rate and fieldwork from April to July 2022; the 2022 shares are 9.5 per cent deeply happy, 38.6 extensively, 45.5 narrowly and 6.4 unhappy, so 93.6 per cent are “happy” by the index's own definition. Established The point in its favour as a measurement instrument is the domains that fell while the headline rose: healthy days, cultural participation, political participation, mental health, and Driglam Namzha. A propaganda instrument does not report a decline in mental health.

Frontier What is not documented is the bite. The policy screening tool was developed in 2008 by the Centre for Bhutan Studies and the GNH Commission, and all new policies and projects are supposed to pass it, scored against the nine domains; the 2022 survey report states the index is “the foundation for National Key Result Areas” in five-year planning. Frontier The most detailed independent assessment retrievable for this brief provides no concrete example of a policy the tool blocked, rejected or substantially modified. It reports the 2010–2015 programme as associated with a 1.8 per cent increase in GNH, driven by living standards, service delivery, health and cultural participation, and quotes a senior Bhutanese official's concern that GNH “has been overused by some people and how they have been distracted from the real business at hand”. Its own verdict on whether the tool altered policy is mixed.

Frontier So the honest answer to the question the whole subject turns on — has a wellbeing measure changed a real allocation decision? — is split. Established Yes at the level of framing and process: New Zealand named priorities and attached NZ$455.1m to the first one, Bhutan's domains anchor its National Key Result Areas, and the UK created a priced unit and put it in Treasury guidance. Frontier No at the level of a documented counterfactual: no source retrieved for this brief identifies a specific decision where the wellbeing measure produced an allocation different from what conventional appraisal would have produced, with the alternative on record. That is an absence in the retrievable record, not a proof of absence — and it is a conspicuous absence, because this is the one thing the movement most wants to be able to show.

3 · Frontier questions

Frontier Question one, and everything else is downstream of it: are reported life-satisfaction scales cardinal enough for group means to mean anything? The live positions are three. Bond and Lang say no, with a mathematical proof and nine failed test cases. Plant and Kaiser and Vendrik say the objection is technically right and empirically inert because real reporting functions are near-linear, on indirect evidence. Harrison says the function is not even stable across decades. Speculative All three cannot be right, and no fetched work adjudicates between them by measuring a reporting function directly.

Frontier Question two: is the Easterlin paradox a fact about happiness or a fact about rulers? Oparina, Clark and Layard support it for rich countries; Harrison, from the same institution, argues it is an artefact of scale drift; Stevenson and Wolfers rejected it on a specification that omits social mediators. Frontier A fourth position — that wellbeing plateaus above an income threshold, against the claim that it continues rising — was the subject of a 2023 adversarial collaboration that could not be retrieved for this brief. It is named here so its absence is visible, and no claim above rests on it.

Frontier Question three: is evaluative life satisfaction even the right target? The experienced-wellbeing minority holds that momentary affect is the object policy should care about, and the two measures behave differently on income, parenthood and commuting. Nobody has built parallel appraisals on both and compared the rankings, which is the test. Frontier Question four, the objective-list position: the capability tradition holds that sufficiency across a list of domains beats any satisfaction scale, and Bhutan's 33 indicators are that instrument in the field. Against it: a 66 per cent sufficiency cutoff is itself an unvoted weighting, so the objective list may only relocate the aggregation problem it was built to escape.

Speculative Question five, almost entirely unstudied: is wellbeing targeting Goodhart-vulnerable? Make life satisfaction a target and it may stop measuring wellbeing. There is essentially no empirical work on this for wellbeing measures specifically — the analogous evidence sits in Human Development Metrics, where a high-salience index produced reform aimed at the indicator rather than the outcome. Frontier Question six: does adaptation make targeting self-defeating? HM Treasury names it as a caveat against its own instrument. Hedonic adaptation is real; its size for policy-relevant changes is unresolved, and an appraisal that ignores it overstates benefits by an unknown factor.

Frontier Question seven, the minority position that deserves more space than it gets: is GNH presentational? The evidence for is suggestive rather than strong — no documented blocked policy in the retrievable record, and a senior official's own complaint that the concept has been overused and distracts from the real business at hand. The evidence against is that the index reports declines in the domains a national brand would suppress. Speculative The test is simple and nobody has run it: publish the screening tool's decisions, including its rejections.

4 · Technological bottlenecks

Established The first bottleneck is that no wellbeing scale has been validated against a criterion that is not itself a self-report. Predictive validity for behaviour and stability of correlates are real and they are consistency evidence, not cardinality evidence. The whole apparatus — means, differences, WELLBY counts, monetised benefits — assumes an interval scale that has never been demonstrated to be one. Frontier Plant's best direct evidence is vignette adjustments averaging 0.2 on a 5-point scale; that is an argument that the error is small, not that the scale is interval.

Established The second is the counterfactual, and it is an infrastructure problem rather than a statistical one. The sector's own knowledge platform states that it is “not always possible to build credible counterfactuals”. Appraisal happens before a decision, so a wellbeing appraisal that changed nothing and one that changed everything leave the same paper trail unless the alternative ranking is published alongside. Nobody publishes the alternative ranking. Established The third follows: the body that was supposed to accumulate these examples closed on 30 April 2024, so the evidence base has no curator.

Established The fourth is children, and it is a hard stop rather than a gap. The flagship WELLBY example excluded children's outcomes from monetisation entirely, on the platform's stated ground that there is no consensus on the monetary value of a child's life satisfaction. Frontier An appraisal instrument that cannot price the beneficiaries of education, child mental health and family policy is missing the applications where wellbeing evidence would matter most.

Frontier The fifth is adaptation and misprediction, named by the Treasury against its own instrument. If a wellbeing gain decays, its present value depends on a decay parameter nobody has estimated for policy-scale interventions; if people mispredict their own future experience, stated-preference cross-checks inherit the error. Speculative The sixth is scale drift. If life-event effects really attenuated 35 to 40 per cent over three decades in one long panel, every time-series wellbeing target is measured on a ruler that is changing length, and no appraisal guidance in the record adjusts for it.

5 · Research dependencies

Established Nothing on this map produces a result this brief waits on. The binding constraints are one missing scientific result and two institutional facts, all three recorded as typed requirements below. The missing result is a cardinal wellbeing scale validated against something that is not a self-report — a physiological, behavioural or revealed-choice criterion against which a claimed one-point difference can be checked. The institutional facts are a published appraisal in which the wellbeing analysis reversed the conventional ranking, with the alternative on record, and a wellbeing-evidence body that outlives a funding cycle.

Frontier What it waits on from research is narrower than the field's own agenda suggests. Not more surveys: the survey infrastructure is good, large and regularly fielded. What is missing is direct measurement of individual reporting functions — how a given person maps an internal state onto the integers 0 to 10, and whether two people map alike. Everything in section 2 turns on that single unmeasured object. Speculative A second, cheaper dependency is replication of the scale-stretch result outside German panel data.

Speculative Two areas are empty rather than thin. There is no study of Goodhart effects on a targeted wellbeing measure, because no jurisdiction has targeted one long enough for gaming to appear. And there is no published comparison of an appraisal built on evaluative life satisfaction with the same appraisal built on experienced affect, although both instruments exist and both are fielded at scale.

6 · Required experiments

Established The highest-value experiment is also the cheapest, and it is not a survey. Measure the reporting function directly: present the same described life to respondents across countries, languages and income levels, anchor with vignettes at scale, and estimate the mapping from described state to reported integer per person. Frontier Bond and Lang's conditions are testable in any existing dataset — check first-order stochastic dominance between the groups whose means are being compared, and report the result next to the mean. Doing that routinely would either rescue the literature or retire specific findings, and it requires no new data collection.

Frontier Second: replicate the scale-stretch finding in other long panels. The attenuation of life-event effects by 35 to 40 per cent between 1991 and 2022 is a single-panel result with large consequences, and the British, Australian and American long panels can all test it on their own life-event series. Speculative If the attenuation replicates, every published wellbeing trend is measured on a moving ruler; if it does not, one of the two strongest current objections to the Easterlin literature disappears.

Established Third, and the one a government could do alone: publish a WELLBY appraisal with the conventional ranking beside it. Take any programme appraised under the supplementary guidance, publish both orderings, and state which option each would have chosen. Frontier The counterfactual column is empty because nobody has been asked to fill it, not because the arithmetic is hard. One published pair would move claim (b) further than another decade of guidance.

Frontier Fourth: publish the Bhutanese screening tool's decisions, including rejections. Eighteen years of screening exists; the number of policies modified and the number blocked would settle a question currently answered by anecdote in both directions. Speculative Fifth: run the parallel-instrument comparison. Appraise the same intervention on evaluative life satisfaction and on experienced affect and report whether the rankings differ. Speculative Sixth: pre-register a Goodhart test. A jurisdiction adopting a wellbeing target should register in advance what gaming would look like — question-order effects, survey-timing effects, selective sampling, provider coaching — before the incentive exists.

7 · Engineering requirements

Established The machinery is more built than the evidence base underneath it, and it is worth describing precisely because the precision is what creates the false confidence. A WELLBY is a one-point change in life satisfaction on a 0–10 scale, per person, per year. It is priced at £13,000 in 2019 prices with a stated range of £10,000 to £16,000, uprated to the appraisal base year, with an illustrative distributional weight of 2.4 applied to bottom-quintile recipients. It enters the Green Book five case model alongside existing welfare estimation, not instead of it. Frontier A range of £10,000 to £16,000 is a band of roughly 23 per cent either side of the central value before any uncertainty in the wellbeing estimate itself is counted, and the guidance is explicit that monetisation belongs only where causal confidence is high.

Established New Zealand's apparatus is a framework rather than a price. The Living Standards Framework arrays outcomes across four capitals — financial and physical, human, natural, social — and the 2019 budget process required ministers to demonstrate how initiatives served five named wellbeing priorities, with Cabinet committees assembling cross-agency packages. Frontier That is a procedural gate, not a valuation rule, and the two are constantly conflated in commentary: a framework can reorder attention without ever pricing an outcome, and there is no exchange rate in the New Zealand system for anything.

Established Bhutan's is the most fully specified instrument of the three, and it is an objective list rather than a satisfaction scale. Thirty-three weighted indicators across nine domains; sufficiency thresholds set concretely at six years of schooling and Nu. 32,951.27 of annual household per capita income; a person counted happy at sufficiency in 66 per cent of weighted indicators; a four-band population split; a national survey of 11,440 respondents at 96.6 per cent response. Frontier Every one of those numbers is a choice, and the 66 per cent cutoff does the same work in Bhutan that a cardinalisation assumption does in a life-satisfaction mean. The objective-list approach relocates the aggregation judgement; it does not remove it.

Established The engineering gap is on the output side, not the input side. Survey design, sampling and fielding are mature; response rates in the Bhutanese survey exceed what most rich-country statistical agencies achieve. Frontier What has never been engineered is the appraisal artefact that would make the instrument auditable: a published document showing the wellbeing ranking, the conventional ranking, and the decision. Until that exists, the machinery is measuring instruments all the way down and never producing a decision record.

8 · Adjacent technologies

The boundary that matters most on this map is with the neighbouring measurement brief, and it is resolved by unit of analysis rather than by subject. Human Development Metrics owns country-level composite indices and what they do to policy — the HDI's construction, the Multidimensional Poverty Index, the SDG indicator framework, Goodhart effects and the Doing Business record. This brief owns person-level wellbeing and its domains: life-satisfaction scale validity, the Easterlin dispute, the WELLBY, and Bhutan's GNH, which is assigned here because its arithmetic runs over persons and its content is wellbeing domains. The Alkire-Foster method appears in both because both objects use it; each brief cites it for its own object. Established Bond and Lang's ordinality problem belongs here and does not transfer next door — HDI and MPI components are objective quantities, not reported scales, and their aggregation difficulty is a weighting problem, not a cardinality one.

Within this map, also: Intelligence Measurement, the same question about scale validity asked of a different latent construct; Reputation Economies, where the measured quantity is trustworthiness and the pathology is inflation rather than cardinality; Future Public Administration, whose finding that evaluation is discretionary explains why appraisal counts exist and outcome studies do not; Public Policy Foresight and Institutional Design, for the machinery a target has to run through; Scientific Advisory Institutions, for who gets to certify a measure; Long-Term Institutions, given that the sector's evidence body lasted less than a decade; and Civilizational Planning, which inherits the question of what a civilisation should be optimising.

Outside it: psychometrics and measurement theory, which owns the ordinal-to-cardinal problem in general; welfare economics and cost-benefit analysis, which supplies the appraisal frame; the capability approach, which supplies the objective-list alternative; and clinical psychology, which supplies the only wellbeing measures routinely validated against non-self-report outcomes.

9 · Institutional requirements

Established Almost every source in this subject is an interested party, and saying which way the interest runs is load-bearing. The defence of scale validity comes from a researcher at a wellbeing research centre who founded an organisation advocating wellbeing-based resource allocation. The Easterlin result comes from the leading academic advocate of wellbeing policy and a World Happiness Report founder, and the interest runs with the finding. The New Zealand budget account comes from the Treasury describing its own budget. The GNH results come from the index's publisher and from its method's co-developer. Established Three sources run against interest and are weighted up accordingly: HM Treasury documenting adaptation and misprediction as weaknesses of its own new instrument; the wellbeing-evidence platform stating that credible counterfactuals are often impossible and that children cannot be priced; and the Bhutanese index reporting declines in mental health and political participation while its headline rose.

Established The institutional failure that matters most is a closure. A national centre existed to accumulate exactly the evidence claim (b) needs, and it ceased operations on 30 April 2024. The instrument outlived its evidence body. Frontier That is a fact about how this kind of measurement gets institutionalised: the appraisal guidance sits inside a finance ministry with permanent funding, and the evaluation function sits in a grant-funded centre with a term. When the two diverge, the priced unit survives and the scrutiny does not.

Frontier The constraint that actually binds is not scientific. A government could publish, tomorrow, one appraisal showing both rankings and the decision taken; a screening authority could publish its rejections. Neither requires a research advance. Established Both are recorded as typed institutional requirements below, because they are things an institution could choose to supply and has so far chosen not to.

10 · Ethical & societal considerations

Frontier The sharpest ethical problem in the subject is adaptation, and it cuts against the people the instrument is most often invoked to help. If reported satisfaction returns toward baseline after a permanent change in circumstances, a WELLBY-based appraisal systematically undervalues remedies for conditions people have adapted to — long-term disability, chronic illness, entrenched deprivation — while overvaluing novel gains. Speculative Bond and Lang's disability case is one of the nine literatures where the comparison conditions fail, which means the population where adaptation matters most is also where the measurement is least defensible.

Established The distributional weight is a value judgement wearing a coefficient. An illustrative factor of 2.4 on bottom-quintile recipients is a defensible egalitarian choice and it is not a measurement; changing it changes which programmes pass. Frontier The same is true of Bhutan's 66 per cent sufficiency cutoff, which determines who counts as happy, and of the £13,000 central value, which determines how much wellbeing a pound is worth against every other benefit in the same appraisal. None of the three was voted on, and all three are more consequential than most things that are.

Established The exclusion of children is the most concrete inequity in the current apparatus. Where a child's life satisfaction has no agreed monetary value, the honest appraisal excludes it — which means it is counted at zero in the arithmetic that decides. Frontier An instrument that prices adults and drops children will systematically favour programmes for adults, and the platform that documented the exclusion said so plainly before it closed.

Frontier And there is a legitimacy question the field rarely states. A headline that 93.6 per cent of a population is happy, on a definition set by the body publishing it, is a claim with political consequences whoever makes it. Speculative The defence is that the same index reported falls in mental health and political participation. That is a real defence and it is the correct standard: the test of a state wellbeing measure is whether it publishes the numbers that embarrass the state, and on the record fetched here, this one does.

11 · Civilizational implications

Established The terminal position is a tie, and both halves have to be said or the brief is propaganda in one direction or the other. Wellbeing measures are technically vulnerable in ways their advocates under-state: across nine literatures the conditions for cardinalisation-invariant group comparison hold in zero cases, variance equality is rejected in all eight testable ones, headline findings reverse under simple monotonic transformations, and the strongest evidence that the reporting scale is unstable comes from the same research centre that supplies the strongest defence of it. And they have been built into real budgetary machinery in at least three countries and have reported findings that cut against their sponsors. Picking a side requires ignoring half the record.

Frontier What follows is a change of question rather than a compromise. The useful question is not whether wellbeing can be measured but what a wellbeing number is being asked to carry. As a description — this population reports lower satisfaction than that one, and here are the domains — the instruments are serviceable and their correlates are stable. As an exchange rate, converting a satisfaction point into £13,000 and trading it against every other benefit in an appraisal, they are carrying a cardinality assumption that has never been validated. Established The same numbers are adequate for the first job and unproven for the second, and the institutional apparatus was built for the second.

Frontier The long-run risk is not that wellbeing measurement fails but that it succeeds administratively without ever being validated scientifically. A priced unit inside permanent Treasury guidance, an evaluation centre that closed after less than a decade, and an unresolved ordinality dispute that practitioners route around is a stable configuration — it can persist indefinitely, because nothing in it generates the evidence that would disturb it. Speculative Instruments that become infrastructure stop being tested; that is what infrastructure means.

Speculative And the largest open question is what a civilisation should be optimising if not this. The critique of GDP that motivated the whole programme is correct: national income is a poor summary of how life is going. The wellbeing alternative is at least measuring the right object. Handwave But “measure the right thing badly” beats “measure the wrong thing well” only if the badness is bounded, and the size of the error is precisely what nobody has established. That is an argument this brief cannot settle, and it declines to pretend the direction of the inequality is obvious.

12 · Timelines

These horizons track survey waves, guidance revisions and the closure and possible re-founding of evidence bodies rather than technology:

  • 10 yr: Frontier Bhutan's GNH survey cycle continues on roughly a five-to-seven-year rhythm, so one or two further index values arrive within this window and the 2022 declines in mental health and political participation are either confirmed or reversed. Frontier The UK's supplementary guidance is uprated rather than revised; expect the £13,000 central value to move with prices and the range to stay proportionally as wide. Speculative Expect at least one replication attempt on the scale-stretch result in a non-German panel, because it is cheap and the payoff is large in either direction. Speculative Whether the counterfactual column gets filled is a decision, not a forecast, and nothing in the record suggests anyone has been asked to fill it.
  • 25 yr: Speculative The plausible split is that measurement improves where a non-self-report criterion is available — clinical mental health, sleep, physiological stress markers, revealed choices with real stakes — and stays unresolved for general life satisfaction, because the object has no external referent to validate against. Speculative If wellbeing targeting spreads far enough for gaming to appear, this is the window in which the first Goodhart evidence for wellbeing measures specifically arrives. Handwave Which way it comes out is an assertion about administrative behaviour, not an extrapolation from anything measured.
  • 50 yr: Speculative If a cardinal scale is ever validated against an external criterion, the entire post-1970 wellbeing literature has to be re-analysed under the validated mapping, and some fraction of its findings will not survive — Bond and Lang say the fraction is large and nobody has estimated it. Speculative If no such validation arrives, the likeliest equilibrium is the one already visible: wellbeing measures as a permanent supplementary layer in appraisal that never becomes the primary metric. Handwave Both are extrapolations from a single unresolved technical dispute.
  • 100 / 250+ yr: Handwave Beyond useful forecasting. The only datum at that horizon is that societies have been generating official statistics about their populations for roughly two centuries, and that the categories they chose look arbitrary in retrospect within decades. Handwave One trend, on a class of measurement, is a story rather than a base rate.

13 · Technology tree & dependencies

  • Depends on Nothing on this map. No brief here produces the result this one waits on, and the wait is not on another brief's discovery but on a validation nobody is attempting. No typed depends-on edge is claimed.
  • Requires (not on this map) Two of these are missing scientific results and two are institutional choices. The scientific ones are specific. A cardinal wellbeing scale validated against a criterion that is not itself a self-report: at present a one-point difference on a 0–10 life-satisfaction scale is checked only against other reported quantities, and the whole monetisation chain — means, differences, WELLBY counts, £13,000 per point — assumes an interval scale nobody has demonstrated to be one. And a direct measurement of the individual reporting function, meaning how a given person maps an internal state onto the integers 0 to 10 and whether two people map alike: the entire Bond and Lang dispute turns on that single object, the defence of near-linearity rests on indirect evidence, and a paper from the defending institution reports the mapping drifting 35 to 40 per cent over three decades. The institutional ones require no discovery at all. A published appraisal showing the wellbeing ranking and the conventional ranking side by side with the decision taken would fill the empty counterfactual column that this brief's second claim rests on; the flagship worked example available instead is a charity's evaluation of its own therapy programme, described by the platform hosting it as an initial approximation. And a wellbeing-evidence body that outlives a funding cycle: the What Works Centre for Wellbeing was established to accumulate exactly this evidence and ceased operations on 30 April 2024, leaving a priced unit inside permanent Treasury guidance with no standing curator for the evidence that would test it.
  • Enables In principle, any policy programme on this map that proposes to justify itself by improved wellbeing — income support, mental health provision, urban design, working-time policy — would rest on this brief's measurement holding. No typed enabling edge is claimed, and the reason is a finding rather than an omission: no appraisal in the retrievable record demonstrates a decision changed by a wellbeing measure, so the enabling relationship has never been shown to operate.
  • Adjacent Psychometrics and measurement theory, which owns the ordinal-to-cardinal problem; welfare economics and cost-benefit analysis, which supplies the appraisal frame; the capability approach, which supplies the objective-list alternative; clinical psychology, which supplies the few wellbeing measures validated against non-self-report outcomes; and within this map Human Development Metrics, Intelligence Measurement and Future Public Administration.

14 · Common misconceptions & speculative claims

Frontier “Bhutan's GNH screening blocked WTO accession, plastic bags, or a specific mining project.” These claims circulate widely and could not be verified from any source fetched for this brief. They are flagged here rather than repeated, and the flag is the honest state: the screening tool exists, has operated since 2008, is scored against nine domains, and the most detailed independent assessment retrievable contains no concrete example of a policy it blocked, rejected or substantially modified. Speculative The stories may well be true; nothing fetched here establishes them, and a site that repeats an unverified anecdote about a real government is doing something worse than leaving a gap.

Established “The UK prices happiness at £13,000 and appraises policy that way.” Half right and misleading. The price is real, official and specific — £13,000 in 2019 prices, range £10,000 to £16,000, distributional weight 2.4. Established But the guidance is supplementary, to be used alongside existing welfare estimation within the five case model, and is not the primary appraisal metric. The Treasury itself restricts monetisation to cases of high causal confidence and names adaptation and utility misprediction as limits on its own instrument.

Frontier “Money doesn't buy happiness — that's settled.” It is not settled in either direction, and the disagreement is specification rather than data. Individual cross-sections show about 0.4 on log income, so doubling income buys roughly 0.3 points out of 10. Cross-country GDP effects vanish once social variables enter except in low-income countries. Rich-country time series show no correlation. Frontier And a paper from the institution most associated with the paradox argues it may be an artefact of a stretching scale, with latent happiness possibly 50 per cent higher than reported. Three live positions, and anyone quoting one as consensus is quoting a third of a literature.

Established “Bond and Lang proved wellbeing research is meaningless.” They did not. They proved that group-mean comparisons of an ordinal scale are not invariant to cardinalisation, showed the conditions for invariance failing in nine literatures, and demonstrated reversals under plausible transformations. Frontier That is an argument about comparisons of means, not about whether the underlying construct exists or whether the measures carry information. The measures have stable correlates, predict behaviour, and in Bhutan's case recorded declines in mental health and political participation while the headline rose — which is what a measurement instrument does.

Frontier “New Zealand's Wellbeing Budget reallocated spending to wellbeing.” The process changed and money was named — five priorities, mental health first, NZ$455.1m for primary mental health services. Frontier The independent interview study of 22 budget actors records roughly eight years of internal disagreement about how to apply the framework and Treasury staff who felt it had “not gone anywhere”, and the retrievable portion of that study does not establish whether allocations changed substantively. Speculative Naming a priority and attaching a figure is not the same as demonstrating a different allocation, and this brief declines to treat it as one in either direction.

Frontier “Objective measures avoid the subjectivity problem.” They relocate it. Bhutan's index is an objective list — 33 indicators, concrete thresholds, no satisfaction scale anywhere — and its central judgement is a 66 per cent sufficiency cutoff that determines who counts as happy. Established A cardinalisation assumption and a weighting cutoff are the same kind of unvoted choice at different points in the pipeline.

Established One prominent source is deliberately not used, and naming it stops its absence reading as an oversight. The 2023 adversarial collaboration on income and emotional wellbeing — the standard reference for whether wellbeing plateaus above a threshold — could not be retrieved for this brief, so no claim here rests on it and none of its numbers are quoted. Speculative It is the most conspicuous hole in this pass, it bears directly on the curvature question in section 3, and it should be closed before anyone builds an argument about income thresholds on what is written here.

Established And the framing itself: “if you can measure it you can manage it” is doing the work of an argument it has not made. The measurement infrastructure for wellbeing is now substantial — a defined unit, an official price, distributional weights, a four-capitals framework, a national screening tool, surveys at 96.6 per cent response. Frontier The managing has not been demonstrated once with a documented counterfactual, and the closure of the body whose job was to demonstrate it is the most informative single event in the recent record.