1 · Concept overview

Ask whether a machine can be creative and you have asked three questions at once, which is why the literature answers in three incompatible voices. Margaret Boden’s 1990 taxonomy separates combinatorial and exploratory creativity — new points inside a conceptual space — from transformational creativity, which changes the space. It separately distinguishes P-creativity, novel to the agent, from H-creativity, novel to the world. Mihaly Csikszentmihalyi’s framework makes creativity a social judgment that no isolated generator can possess. On Boden’s P-creativity current systems are plainly creative and the question is boring; on Csikszentmihalyi’s account the question is malformed; on transformational creativity nobody has an instrument that could return an answer. Most public argument about machine creativity is two people using different definitions and reporting the disagreement as a fact about machines.

What can be measured has been measured, at increasing scale, and the result is stable: models beat the average human on divergent-thinking tasks and lose to the best humans, with the largest human-scored study reversing even the average. Meanwhile the most convincing machine creativity anywhere — FunSearch improving a published bound in additive combinatorics, AlphaEvolve improving on Strassen’s algorithm in the complex setting — happens in formal domains where an exact evaluator, not a language model, decides what is true, and would not register on any psychometric creativity test. Those two literatures barely cite each other. This brief also carries an unusual source condition, stated here rather than buried: eleven of the thirty entries in its research pack are bibliographic records whose text was never obtained, because creativity psychology sits in journals that refused to serve. Where that bites — most sharply on the validation history of the Torrance tests, the field’s own standard instrument — this brief names the absence and prints no number from it.

2 · Current scientific position

Established The operative definition is Boden’s, and its two halves come apart under pressure. The Creative Mind: Myths and Mechanisms (Weidenfeld & Nicolson, 1990) defines creativity as the ability to produce ideas or artifacts that are new, surprising and valuable, then draws two orthogonal distinctions: P-creativity against H-creativity, and exploration within an established conceptual space against deliberate transformation or transcendence of that space, the latter described as far more radical, challenging and rarer. Frontier A provenance note this brief will repeat rather than hide: Boden’s text was not obtained in the research pass behind this page, and these formulations reached it through an encyclopaedic secondary source. They are standard in the literature and should nonetheless be checked against the book before being leaned on. Established Newell, Shaw and Simon’s 1963 criteria — that an answer be novel and useful, that it demand rejection of previously accepted ideas, that it follow intense motivation, and that it emerge from clarifying a vague problem — remain embedded in modern operationalizations, and are stricter than anything the current benchmarks test. Established Csikszentmihalyi’s DIFI framework locates the judgment socially: the field, meaning other people, determines whether work qualifies. Frontier These are not shades of one definition. Under DIFI, machine creativity is not a property a machine can have on its own; under P-creativity it obviously can. The answer to the headline question is fixed by the definition chosen, before any evidence is consulted, and almost no popular treatment says which one it is using.

Established The measured record is a distributional statement, and quoting any single moment of the distribution is the standard error in this field. Four studies, four teams, four instruments. Hubert, Awa and Zabelina (Scientific Reports 14:3440, 2024) compared 151 humans against GPT-4 on the Alternative Uses Task, the Consequences Task and the Divergent Associations Task, and reported that AI was robustly more creative on each measurement, more original and more elaborate when fluency of responses was controlled. Established Koivisto and Grassini (Scientific Reports 13:13601, 2023), published five months earlier with 256 humans and three chatbots on the Alternate Uses Task, found that the chatbots beat the human average but that the best human ideas still matched or exceeded them. Established Bellemare-Pepin, Lespinasse, Thölke, Harel, Mathewson, Olson, Bengio and Jerbi compared state-of-the-art models against a dataset of 100,000 humans on the Divergent Association Task and on creative writing under identical objective scoring: some models surpass average human performance and remain below the mean of the more creative participants, and top humans outperform every model tested. Established Wang, Huang, Shen and Uzzi (Nature Human Behaviour, 2025) ran the largest comparison — 9,198 humans against 215,542 model observations — and found that human creativity on average is slightly higher than that of LLMs, that differences are pronounced at the extremes with humans showing greater variability and a higher right-hand tail, and that persona prompting lifts performance only to a threshold beyond which output becomes opposite to real-life patterns. Established Note what that last study does to the first: at roughly sixty times the sample size, it reverses the headline on the mean. Frontier The honest summary is the shape rather than the winner — models beat the human average, humans own the right tail — and it is the most robust empirical generalization in this subject.

Established Creative performance has not improved with model generation, which is the single most surprising number here. Haase, Hanel and Pokutta (arXiv:2504.12320, 2025) ran fourteen models on the Divergent Association Task and the Alternative Uses Task and found no evidence of increased creative performance over the preceding eighteen to twenty-four months. Established On the Alternative Uses Task all models performed on average better than the average human, with GPT-4o and o3-mini leading — and only 0.28 per cent of model-generated responses reached the top ten per cent of human creativity benchmarks. Established The same study documents enormous intra-model variability: one model, one prompt, produced outputs ranging from below-average to original. Frontier That variability is not a footnote. It means a single sample from a model is a draw from a wide distribution, and that every comparison in the paragraph above is comparing a human population maximum against a model sample maximum at an unmatched and usually unreported sampling budget.

Established Every automated creativity claim sits on top of a judge that agrees with humans three times in four. Li, Zhu, Xu, Wang and Mao (arXiv:2504.15784, EMNLP 2025) built a reference-based Likert scorer grounded in the Torrance Test of Creative Writing and reached a pairwise accuracy of 0.75, a fifteen per cent improvement on prior methods. Established That is the current published ceiling for automated creativity judgment. Frontier It is therefore a noise floor under every large-scale creativity comparison that used automated scoring, and it is a plausible mechanism for why the largest human-scored study is the one that reverses the smaller results. Frontier A related record this brief could not read: Transactions of the ACL (2024) carries a study reporting that model-based story evaluation beats automatic measures but cannot explain its own ratings; the record was verified and the text was not obtained, so no figure from it appears here. Speculative If the explanation-gap result holds, the field’s judges are better than its metrics and worse than its experts, and nobody has characterized the region in between.

Established The interpolation argument has been made measurable, and it favours the skeptics on text. Lu, Sclar, Hallinan, Mireshghallah, Liu, Han, Ettinger, Jiang, Chandu, Dziri and Choi (arXiv:2410.04265) define a CREATIVITY INDEX — how much of a text can be reconstructed from existing web snippets — computed by a dynamic-programming algorithm they call DJ SEARCH. Established Professional human authors score 66.2 per cent higher than language models on it; Hemingway scores measurably higher than other human writers; and the index outperforms DetectGPT by 30.2 per cent as a zero-shot machine-text detector, which is independent evidence that the quantity it measures is real. Established The result with the most consequence for practice: alignment reduces a model’s creativity index by 30.1 per cent. Established Zhang, Diddee, Holm, Liu, Liu, Samuel, Wang and Ippolito’s NOVELTYBENCH (arXiv:2504.05228) evaluated twenty frontier models on producing multiple distinct and high-quality outputs, found that state-of-the-art systems generate significantly less diversity than human writers, and found — the counterintuitive one — that larger models within a family often exhibit less diversity than their smaller counterparts. Frontier In-context regeneration narrows the gap without closing it. Frontier Read together, the training pipeline that raises benchmark scores appears to lower creative diversity; read strictly, nobody has run the controlled experiment that would establish it causally.

Established The strongest machine creativity in existence is in formal domains, and the psychometric literature cannot see it. FunSearch (Romera-Paredes et al., Nature 625:468–475, 2024) paired a pretrained language model with a systematic evaluator and produced new cap set constructions: a cap set of size 512 in n=8, a cap set capacity lower bound improved from 2.2180 to 2.2202, a full-size admissible set giving capacity at least 2.219486, and a partial admissible set in A(24,17) of size 237,984 — plus online bin-packing heuristics beating first-fit and best-fit and reaching 0.03 per cent from the optimality lower bound on 100,000-item Weibull instances. Established AlphaEvolve (Novikov et al., arXiv:2506.13131, 2025) found a scheme multiplying 4×4 complex-valued matrices in 48 scalar multiplications, described by its authors as the first improvement over Strassen’s algorithm in that setting after 56 years. Established A correction that matters and is routinely botched: AlphaTensor’s 47 multiplications (Nature 610:47–53, 2022) were for 4×4 matrices in finite fields, a different quantity, so a reader who sees 47 and then 48 and concludes the newer system did worse has been misled by the arithmetic rather than by the papers, both of which state their settings in their own abstracts. Frontier These results are H-creative in Boden’s sense — new to the world, valuable, surprising to specialists — and neither would score a single point on a Divergent Association Task. Established The structural feature both share is that the language model proposes and a cheap exact verifier disposes; the creativity is a property of the loop, not of the generator. Frontier The literature that measures LLM creativity and the literature that produces machine creativity are close to disjoint, and that disjunction is the most important unremarked fact in this subject.

Established The canonical exhibit is mis-cited more often than not. Silver et al., “Mastering the game of Go with deep neural networks and tree search” (Nature 529:484–489, 2016), reports a 99.8 per cent winning rate against other Go programs and a 5–0 defeat of the human European champion. Established Move 37 is from the March 2016 Lee Sedol match and does not appear in that paper; any citation of Move 37 to that DOI is a mis-citation. Frontier This brief did not obtain a primary source for the Lee Sedol match itself, so it prints no move-level description, no commentary quotation and no probability figure for the move. Speculative The strongest version of the argument that survives that omission is structural rather than anecdotal: a policy trained to predict human play assigned low probability to a move that specialists later judged strong, which is a claim about a distribution over human moves and not, on its own, evidence of a new conceptual space. Frontier And the case is architecturally identical to FunSearch — a search inside a fully specified space with an exact evaluator — which makes it good evidence for exploratory creativity of superhuman depth and no evidence at all for the transformational kind.

3 · Frontier questions

Established The right-tail result is now replicated across four independent teams and four instruments, which is more replication than almost anything else in this category enjoys. Koivisto and Grassini at N=256, Bellemare-Pepin and colleagues against a 100,000-human dataset, Haase and colleagues across fourteen models, and Wang, Huang, Shen and Uzzi at 9,198 humans against 215,542 model observations all find the same shape. Frontier The open question is whether the shape is a fact about creativity or a fact about sampling.

Frontier The sampling objection is the sharpest live argument and nobody has run the experiment that settles it. The human right tail is drawn from thousands of independent minds; the model right tail is drawn from repeated samples of one distribution, usually at an unreported temperature and an unmatched N. Established Brown et al.’s “Large Language Monkeys” (arXiv:2407.21787) showed that coverage — the fraction of problems solved by any sample — scales log-linearly with sample count over four orders of magnitude, taking DeepSeek-Coder-V2-Instruct from 15.9 per cent at one sample to 56 per cent at 250 on SWE-bench Lite. Frontier The creativity analogue is obvious, cheap and undone: human maximum from N people against model maximum from N independent samples, N matched, judged blind. Speculative If matched-N sampling reaches the human maximum, the entire right-tail literature is a budget artifact and the field’s central empirical result dissolves. Speculative If it does not, the tail is a genuine capability gap and the deflationary reading of machine creativity strengthens considerably.

Frontier The diversity–capability inversion suggests the training pipeline is the problem, and it has never been tested directly. NoveltyBench finds larger models within a family less diverse; the Creativity Index finds alignment costing 30.1 per cent of linguistic novelty; Haase and colleagues find a flat eighteen-to-twenty-four-month trend. Speculative Together these imply that post-training buys benchmark score with creative range, and that every lab is currently choosing a point on a creativity–helpfulness Pareto frontier without reporting the trade. Frontier One base model, five RLHF strengths, both instruments measured at each checkpoint would settle it, and the compute cost is small by frontier standards.

Frontier Novelty and value come apart, and machines are good at exactly one of them. Si, Yang and Hashimoto (arXiv:2409.04109), in a blinded study with more than a hundred NLP researchers, found LLM-generated research ideas judged more novel than human expert ideas at p<0.05 while being judged slightly weaker on feasibility. Established The authors’ own caveats travel with the result and are usually stripped from it: failures of model self-evaluation, lack of diversity in generation, and the difficulty of novelty judgments even for experts. Frontier NoveltyBench makes the same split structural by scoring distinctness and quality separately and finding systems trade one for the other. Frontier Most of the creativity literature reports a scalar, which is precisely the wrong shape for a two-part construct.

Frontier Measurement is moving to the population level, where the interesting effects probably are. Brinkmann and colleagues’ “Machine culture” (Nature Human Behaviour 7:1855–1868, 2023) argues that intelligent machines simultaneously transform the cultural evolutionary processes of variation, transmission and selection: recommenders alter social learning, chatbots become cultural models, and machines contribute traits from game strategies and visual art to scientific results. Frontier Sourati, Ziabari and Dehghani (arXiv:2508.01491, Trends in Cognitive Sciences) argue models reflect and reinforce dominant styles while marginalizing alternative voices and reasoning strategies, and that unaddressed homogenization could flatten the cognitive landscapes that drive collective intelligence. Frontier Doshi and Hauser’s Science Advances paper provides the first controlled individual-versus-collective measurement; its record was verified for this brief and its text was gated, so its effect sizes are not printed here. Speculative Nothing yet measures the longitudinal diversity of a real creative corpus under AI adoption, which is where a civilizational-scale effect would first become visible.

Frontier Task-specific evaluations are proliferating faster than construct definitions. A causality-aware multimodal paradigm built on the Oogiri game (IEEE TPAMI, 2025), a semantic-diversity study in the Journal of Creative Behavior (2024), a figurative-language comparison at NLP4DH 2025, and a human-based novelty-plus-appropriateness evaluation at EACL 2026 each define creativity differently. Frontier All four were surfaced as verified bibliographic records for this brief and none of their texts were obtained, so their numbers are absent from this page by policy rather than by oversight. Speculative A field with four new instruments a year and no criterion validity for any of them is generating incomparable numbers at an accelerating rate.

4 · Technological bottlenecks

Established The binding constraint is that there is no validated instrument distinguishing a new point in a known space from a new space. Everything interesting in Boden’s taxonomy lives on that distinction, and no published measure operationalizes it. Established The Creativity Index measures reconstructibility from web text, which is a good proxy for combinatorial novelty and explicitly not a measure of conceptual transformation. Frontier Divergent-thinking scores measure fluency, originality and elaboration within a task frame supplied by the experimenter, so by construction they cannot register a change of frame. Frontier The consequence is symmetrical and uncomfortable for both camps: until such an instrument exists, the claim that AI cannot be transformationally creative is unfalsifiable, and the claim that AI is creative is unmeasurable.

Established The second constraint is the judge. The best published automated creativity evaluator agrees with human pairwise judgment 75 per cent of the time, so any comparison scored automatically inherits that noise floor, and the field routinely reports point estimates without intervals. Frontier The third is the scalar: novelty and value are two quantities and are reported as one, which makes systems that are highly novel and low value indistinguishable from systems that are neither.

Frontier The fourth constraint is a validity debt this brief could not even audit. The field’s standard instruments — the Torrance Tests of Creative Thinking, the Alternative Uses Task, the Divergent Association Task — have contested criterion validity for predicting real creative achievement in humans, and importing them wholesale to machines imports that problem intact. Handwave This brief states the contest and prints no effect size from it: the Torrance validation literature sits in psychology journals that were gated to the research pass behind this page, and a brief that guessed at those numbers would be committing the exact error it is warning about. Frontier What can be said with confidence is narrower and still damaging: a machine score on an instrument normed on a human population is not obviously a measurement of anything, because norming establishes what a score means for the population sampled and models were not in that population.

Speculative The fifth is that H-creativity is a social fact and nobody tracks the social process. Boden’s H-creativity and Csikszentmihalyi’s DIFI both make adoption constitutive, and no registry tracks whether other creators build on machine-originated artifacts. Frontier This is the same missing infrastructure that blocks the field-uptake criterion in Artificial Scientists, and it is missing for FunSearch’s cap sets as much as for generated images.

5 · Research dependencies

Established This brief waits on a result that Intelligence Measurement is producing, and the dependency is not decorative. Creativity research is measurement research with a harder construct, and every pathology documented for benchmarks recurs here in a literature with a smaller methodological toolkit. Established Raji, Bender, Paullada, Denton and Hanna (arXiv:2111.15366) set out the construct-validity requirement that a benchmark claiming generality must state its construct, argue that its items sample it, and show that scores predict criterion behaviour outside the benchmark; the creativity instruments satisfy the first, gesture at the second, and have contested evidence for the third.

Established Miller’s “Adding Error Bars to Evals” (arXiv:2411.00640) makes the general point that evaluations are experiments and should be analysed as such; human-versus-model creativity comparisons are the clearest case on the site of a literature reporting differences of means without reporting whether they are separable. Established Schaeffer, Miranda and Koyejo’s demonstration that discontinuous metrics manufacture apparent emergence transfers directly: a creativity score is a metric choice, and the choice between a mean, a maximum and a top-decile hit rate is exactly what makes Hubert and Haase appear to disagree. Frontier Chollet’s framing in “On the Measure of Intelligence” — that skill is modulated by priors and experience, so unlimited training data lets an experimenter buy skill in a way that masks generalization — is the argument the creativity literature needs and has not adopted; the Creativity Index is the nearest thing to it, and it measures priors directly by measuring reconstructibility from the training distribution’s ancestor.

Frontier The dependency that bites hardest is the one about human comparison groups. Benchmark research has learned, expensively, that a human baseline is not a number but a protocol: which population was sampled, under what time budget, with what incentive, and with the distribution reported rather than the mean. Established The creativity studies span undergraduate convenience samples in the hundreds, a repurposed archive of 100,000 respondents, and a purpose-recruited panel of 9,198, and they are routinely quoted against each other as though those were the same reference class. Frontier Until the measurement brief’s standard is imported wholesale — a named population, a reported spread, an interval on every difference — the four studies at the centre of this subject will keep appearing to disagree for reasons that have nothing to do with machines.

6 · Required experiments

The workback plan to a machine that is transformationally creative in Boden’s sense, in dependency order. The target is an artifact that specialists in a domain judge to have required a new conceptual space, produced by a system that was not pointed at that space.

Frontier L1 — report novelty and value as two numbers. Re-score every existing model-creativity study on the NoveltyBench distinct-versus-quality split and the Si et al. novelty-versus-feasibility split. Trivial, and undone. Frontier L2 — match the sampling budget. Human maximum from N independent people against model maximum from N independent samples, N matched, decoding parameters reported, judged blind by raters whose agreement is reported. This is the single most informative cheap experiment in the topic and no one has run it. Speculative L3 — the binding link: a validated instrument for transformational creativity. Construct a gold-standard set of historical H-creative artifacts labelled by domain experts as exploratory or transformational; demonstrate that a candidate metric separates them; only then apply it to machine output. Frontier The ordering matters and is usually inverted: applying an unvalidated metric to machines first is how the field acquired its current stock of incomparable numbers. Frontier L4 — decouple creativity from alignment. One base model, five RLHF strengths, Creativity Index and NoveltyBench at each checkpoint, publish the Pareto frontier. Speculative L5 — population-level longitudinal measurement. Corpus diversity of a real creative field — short-fiction submissions, advertising copy, chemistry paper titles — tracked before and after adoption, with model-market concentration as a covariate. Doshi and Hauser is the lab-scale version; the field-scale version is unbuilt. Speculative L6 — adoption as ground truth. Register machine-originated artifacts and track derivative human work at twenty-four months, which is the only measurement that would satisfy Boden’s H-creativity and Csikszentmihalyi’s DIFI at the same time. Frontier L1, L2 and L4 are each achievable inside one year on an academic budget. L3 is the one nobody can shortcut, and it is a humanities-and-mathematics labelling exercise before it is a machine-learning problem. Frontier Every one of these experiments must carry the protocol that the field currently omits: who administered the instrument, which human population supplied the comparison, what the sampling budget was on each side, and whether the instrument has ever been validated for the population it is being applied to. Established A machine score on a test normed on undergraduates is not a measurement until that last question has an answer, and at present it has none.

7 · Engineering requirements

Frontier The instrumentation this subject needs is unglamorous and mostly absent. A creativity score without its decoding parameters is not a fact about a model: temperature, top-p, repetition penalty and the number of samples drawn all move the measured quantity, and the majority of published human-versus-model comparisons do not report them. Frontier Serving infrastructure that logs sampling configuration alongside every scored generation is a small engineering task and would retroactively make the existing literature comparable.

Speculative Matched-N sampling requires an evaluation harness that treats a model as a population, not as a respondent. Drawing 250 independent samples per prompt across twenty models is affordable; scoring them is not, because human scoring is the only judge that clears the 0.75 automated ceiling. Frontier The realistic design is a two-stage funnel — automated pre-screening to a shortlist, blind human scoring of the shortlist — with the screening step’s false-negative rate measured rather than assumed, which no published creativity study does.

Speculative The RLHF ablation needs a checkpoint series no lab currently ships. Releasing a base model together with graded post-training checkpoints would let anybody measure the creativity–alignment frontier; the labs release the endpoints and nothing in between, so the trade-off can be inferred across model families and never measured within one. Handwave An open reproduction at small scale would be scientifically adequate and is the obvious move for an academic group with a few thousand GPU-hours.

Speculative The adoption registry is a database problem wearing a philosophy problem’s clothes. To make H-creativity measurable you need persistent identifiers for machine-originated artifacts, a record of what generated each one and under what configuration, and a crawler that detects derivative human work — citation for a mathematical construction, reuse for a heuristic, stylistic descent for an image. Frontier The first two are solved technology and nobody deploys them; the third is the hard part and is the same unsolved provenance problem the detection literature is already working on, which is a reason to build the two together rather than separately.

8 · Adjacent technologies

Established Program search with an exact verifier is the adjacent technology that actually produces the goods. FunSearch and AlphaEvolve are evolutionary loops in which a language model supplies the proposal distribution and a systematic evaluator supplies the truth, and that architecture is the only one on this map with an undisputed H-creative output. Frontier Its limitation is equally clear and is the same limitation described in Artificial Scientists: it requires a cheap, exact, automatic verifier, which exists in combinatorics and compiler optimization and does not exist in poetry.

Established Machine-text detection and creativity measurement turn out to be the same instrument. The Creativity Index outperforms DetectGPT by 30.2 per cent as a zero-shot detector, which means the quantity that separates human from machine text is the same quantity the field calls linguistic novelty. Frontier That equivalence has consequences for provenance policy that nobody has worked through: a good creativity metric is a good detector, and a model optimized to score well on the former becomes harder to detect by the latter.

Frontier Recommender systems belong here as the selection operator in the machine-culture frame rather than as a generative technology; open-endedness research in evolutionary computation supplies the only formal literature on generating novelty without a fixed objective; and the generative-art perspective in Science (2023) by Epstein, Hertzmann, Akten, Farid, Fjeld, Frank and colleagues is the field-defining statement on the aesthetics side, cited here from a verified record whose text this brief did not obtain. Frontier Collective Intelligence and Multi-Agent Intelligence Systems carry the population-level machinery that the homogenization results depend on. Speculative The one genuinely underexplored adjacency is quality-diversity search, an evolutionary-computation tradition that optimizes for an archive of behaviourally distinct solutions rather than for a single best one; it is the only mature body of work that treats diversity as the objective rather than as a diagnostic, and the creativity-evaluation literature does not cite it. Handwave A model post-trained against a quality-diversity objective instead of a preference model is the cheapest available test of whether the alignment–creativity trade-off is a law or a choice.

9 · Institutional requirements

Frontier The binding experiment needs a panel nobody funds. Labelling historical artifacts as exploratory or transformational requires art historians, musicologists, mathematicians and historians of science working to a shared protocol, and the output is an instrument rather than a paper. Speculative No standing body commissions that kind of work: creativity psychology funds studies, computer science funds benchmarks, and an interdisciplinary labelling corpus falls between them. Frontier It is the same institutional shape as the missing artifact registry — useful to everyone, owned by no one, and unpublishable in the venues that would have to staff it.

Established The field’s own literature is not open, and that is a finding rather than an excuse. The research pass behind this brief could reach preprint servers and Nature-family article pages and could not reach the psychology and ACM venues where the validation history of the creativity instruments lives; eleven of thirty catalogued sources are bibliographic records whose text was never obtained. Frontier A subject whose foundational measurement literature is inaccessible to a well-resourced literature search is a subject in which secondary claims propagate uncorrected, which is exactly the pattern the corrections register behind this category documents elsewhere: a peer-reviewed correction two years old, still absent from the encyclopaedia entry where the wrong number does argumentative work.

Frontier Copyright and authorship regimes are being asked to make originality determinations that no instrument supports. Registration systems, plagiarism policies and prize eligibility rules all now require a judgment about machine contribution, and the best available measure of that judgment agrees with human raters three times in four. Speculative A disclosure norm requiring labs to publish creativity-diversity metrics alongside capability benchmarks would cost little and would make the alignment trade-off visible; no lab does it, no regulator asks, and the metric that would be reported is the one whose validation is incomplete.

Frontier There is no prize institution for creativity and the absence is instructive. Capability measurement has a funded adversarial ecosystem — private evaluation sets, standing prizes, refresh budgets — because a benchmark that becomes important gets optimized against. Speculative Creativity has none, which means its instruments are safe from Goodhart only because nobody is trying hard enough to game them; the moment a creativity score carries money, the field will discover its measures were never adversarially robust and will have no private-set tradition to fall back on. Frontier Building the reserved evaluation corpus before the incentive arrives is cheap now and impossible later.

10 · Ethical & societal considerations

Frontier The population-level effect is the ethically serious one and it runs opposite to the individual-level effect. Doshi and Hauser’s title carries both halves — generative AI enhances individual creativity and reduces the collective diversity of novel content — and the second half is the one that compounds. Frontier A writer who uses a model produces better work than they would alone; a thousand writers using the same three models produce a narrower corpus than a thousand writers working alone would have. Speculative If that is right, the harm is not attributable to any individual choice, which puts it outside the reach of consent-based ethics and inside the reach of competition and procurement policy.

Frontier Homogenization is a systems property, not a model property. Even a maximally creative model, adopted universally, narrows the variation operator in the cultural evolutionary loop; the prediction that follows is that diversity loss should scale with model-market concentration rather than with model capability, and nobody has measured it. Speculative That prediction is testable with a market-concentration index and a corpus diversity measure, and it would reframe an aesthetic complaint as an antitrust question.

Frontier Using an unvalidated creativity instrument to make decisions about people is the near-term concrete harm. Divergent-thinking scores already appear in selection and admissions contexts; a machine-scored version with a 0.75-accurate judge would extend a contested instrument with an additional noise source and an unexamined bias profile. Established The minimum defensible standard is the one the measurement literature already states: name the construct, show the items sample it, and show the score predicts something outside the test.

Frontier Attribution is the unresolved distributive question. Every machine-creative artifact is downstream of a training corpus of human work whose authors are not identified in the output and cannot be, and the Creativity Index result gives that intuition a number for text by measuring how much of a generation is reconstructible from existing writing. Speculative A credit mechanism that paid out in proportion to measured reconstructibility is technically imaginable and politically unattempted; the honest observation is that the measurement exists, the mechanism does not, and the argument is currently conducted without either.

11 · Civilizational implications

Frontier If creativity is a population process, the machine intervenes at three points and only one of them is the one people argue about. The machine-culture frame identifies variation, transmission and selection as separately transformed: models generate candidate traits, recommenders decide which propagate, and chatbots become cultural models that others imitate. Frontier Public debate is almost entirely about the first, and the measured effects so far are strongest on the second and third.

Speculative The pessimistic scenario is a civilization with more artifacts and less variance. A cultural corpus whose novelty is high per item and low across items would look productive by every current metric and would explore less of the space of possible ideas per decade, and the flattening would be invisible to any instrument the field currently deploys because they all score items rather than corpora. Handwave On the long view this is the one machine-intelligence risk that does not require any capability advance to arrive: it follows from adoption plus concentration alone.

Speculative The optimistic scenario is verified search in every domain that admits a verifier. FunSearch and AlphaEvolve suggest that wherever a cheap exact evaluator can be written, machine proposal distributions plus that evaluator will out-search human intuition, and the civilizational payoff is a steady stream of H-creative results in combinatorics, algorithm design, materials screening and circuit synthesis. Handwave The limit of that programme is the verifier, and the domains people care most about — art, ethics, institutions — are precisely the ones where writing the verifier is the whole problem. Speculative A civilization that automates creativity everywhere a verifier exists and nowhere else will find its formal domains accelerating and its interpretive ones static, which is a strange shape for a culture and not one anybody has planned for.

Handwave The workback plan to the impossible version is short and none of its steps are technological. To get a machine that changes a conceptual space rather than searching one, you first need a way to tell the two apart, which is a labelling project; then a way to tell whether a change was valuable, which is an adoption-tracking project; then a system whose representation of the space is explicit enough to be altered, which is the only step that is a research problem. Handwave That ordering is worth stating because the field spends almost all of its effort on the third step and almost none on the first two, and the first two are what make the third legible as progress.

12 · Timelines

These horizons track measurement rather than capability, because in this subject the measurement is the blocker.

  • 10 yr: Frontier The cheap experiments land. Two-number reporting of novelty and value becomes standard; a matched-N sampling comparison settles whether the human right tail survives a fair budget; at least one open RLHF checkpoint series makes the creativity–alignment frontier measurable within a single model family. Frontier Expect the right-tail result to survive in weakened form and the alignment trade-off to be confirmed. Speculative The one plausible surprise on this horizon is a matched-N result showing that a single model at 250 samples reaches the human population maximum, which would retire the field’s central empirical claim in a single paper.
  • 25 yr: Speculative A validated instrument separating exploratory from transformational creativity exists, built on an expert-labelled corpus of historical artifacts, and is applied to machine output for the first time. Speculative Until this lands every claim in this brief’s headline question remains formally undecidable, and there is no technical reason it could not have landed twenty years earlier — the obstacle is that the instrument is an interdisciplinary labelling corpus, which is nobody’s publication and therefore nobody’s project.
  • 50 yr: Speculative Adoption becomes the criterion. Machine-originated artifacts are registered and their derivative human work tracked, and the first defensible H-creativity claim for a machine is made on social evidence rather than on a rating scale. Handwave Whether the count is large or negligible is genuinely open, and this brief declines to guess. Handwave The interesting failure mode on this horizon is a large count that nobody accepts, because the adoption was of the artifact and the credit went to the operator.
  • 100 / 250+ yr: Handwave Transformational creativity as a demonstrated capability: a system that changes the conceptual space rather than searching it, verified against the instrument built at the twenty-five-year mark. Handwave The strongest argument that it is impossible — that transcending a space requires representing the space, which a text model does not do — is an argument about current architectures and not a theorem, and the strongest argument that it is achievable is that human transformational creativity is itself a physical process with no known exemption. Handwave A reader who wants the honest bottom line should notice that this row is dated to a century not because the engineering is a century away but because the measurement is not scheduled at all.

13 · Technology tree & dependencies

  • Depends on Intelligence Measurement, directly and unusually heavily. This brief’s binding constraint is an instrument that does not exist, and the methodology that would build it — a stated construct, evidence that the items sample it, and criterion validity against behaviour outside the test — is the same methodology that brief is assembling for capability benchmarks. Construct validity, item-level parameters, error bars on differences of means, and specified human comparison populations are all prerequisites here, and the creativity literature is roughly a decade behind the benchmark literature in adopting any of them. The dependency runs one way: no improvement in generative models resolves anything in this brief, and an adequate measurement theory would resolve most of it. The specific import needed is the human-baseline protocol — a named comparison population, a reported distribution rather than a mean, a time budget, and an interval on every difference — without which the four large studies at the centre of this subject will continue to appear to contradict each other for reasons that are entirely methodological.
  • Enables A validated creativity instrument would put a floor under several claims that currently float. Automated scientific discovery needs a way to say whether a machine-proposed hypothesis is a new point in a known space or a new space, which is exactly the exploratory-versus-transformational distinction; open-ended search and curriculum design need a novelty signal that is not reconstructibility from the training corpus; and any regulatory or copyright regime that turns on originality needs a measure that survives adversarial use. None of these are typed edges here, because the instrument does not exist and an edge from a brief to a result nobody is producing is a promise rather than a dependency. The same reasoning governs the absence of a requires token: the constraint here is not a factory, a licence or a supply chain but a labelling corpus and a validation study, and naming those as external constraints would overstate how far outside the research system they sit. They are cheap, they are unglamorous, and they are simply not being done.
  • Adjacent Artificial Scientists shares this brief’s architecture and half its evidence: the verified-search loop that produces machine H-creativity in formal domains is the same loop, and the missing artifact-adoption registry is the same gap. Collective Intelligence and Multi-Agent Intelligence Systems supply the population-level machinery that the homogenization and machine-culture results depend on, and the diversity-of-generators question is structurally identical to the error-correlation question in ensembles. Artificial General Intelligence is adjacent rather than upstream: nothing in this brief depends on general capability, and the flat eighteen-to-twenty-four-month creativity trend across fourteen models is direct evidence that capability gains and creative gains have come apart. Machine Consciousness is adjacent for a reason worth stating rather than assuming: both subjects are blocked by a construct nobody can operationalize, and both have literatures in which the philosophical work and the empirical work cite each other without constraining each other. The difference is that creativity has instruments, they are simply the wrong ones, which is a more tractable position than having none at all.

14 · Common misconceptions & speculative claims

Established “AI is more creative than humans.” The four largest studies agree that models beat the human average and lose to the human right tail, and the largest human-scored study, at 9,198 participants, finds humans slightly ahead even on the mean. Established The enthusiast error and the skeptic error are the same error: quoting one moment of a distribution. Frontier The correct sentence names the moment, the instrument and the sample size, and it does not have a winner in it.

Established “GPT-4 scored higher than humans on the Torrance test, so it is creative.” Three things are wrong at once. Established Only 0.28 per cent of model-generated responses reach the top decile of human responses, so the claim is true of the mean and false of the range. Frontier Divergent-thinking instruments have contested criterion validity for predicting real creative achievement in humans, which means a machine score inherits a validity problem rather than escaping it. Handwave And this brief cannot tell you how contested, because the Torrance validation literature was not obtainable in the research pass behind this page — a gap named here rather than filled with a plausible-looking number. Established Note also that the automated evaluator most often invoked in this context is grounded in the Torrance Test of Creative Writing; it is not the Torrance test, and its 0.75 pairwise agreement with human judgment is a property of the evaluator, not of any model it scores.

Established “Models are getting more creative.” Across fourteen models on two instruments, Haase and colleagues found no evidence of improvement over eighteen to twenty-four months. Frontier That is a striking dissociation from the capability benchmarks over the same window and is the single best argument that creativity is not a by-product of scale.

Established “Bigger models are more creative.” NoveltyBench finds that larger models within a family often exhibit less diversity than their smaller counterparts. Frontier The most likely mechanism is post-training rather than parameter count, which is testable and untested.

Established “Move 37 proves machine creativity” — cited to the 2016 Nature AlphaGo paper. That paper reports a 5–0 win over the European champion and a 99.8 per cent win rate against other Go programs. Move 37 is from the later Lee Sedol match and is not in it. Frontier The citation error is worth more than pedantry: it means a large share of the popular argument for machine creativity rests on an anecdote whose primary source most of its citers have never opened. Speculative The defensible version of the claim is about a search finding a strong move that a human-play model rated unlikely, which is exploratory creativity of superhuman depth inside a fixed and fully specified space.

Frontier “AI art is derivative because models only interpolate.” The Creativity Index turns this from rhetoric into a measurement — a 66.2 per cent gap to professional authors on reconstructibility from web text — and the measurement is for text. Handwave Extending it to images or music without running the equivalent measurement is unsupported, and the equivalent measurement is not obvious, because there is no snippet-level web index for images that plays the role DJ SEARCH’s corpus plays for text.

Frontier “Generative AI makes people more creative.” Doshi and Hauser’s title carries both halves: individual creativity up, collective diversity of novel content down. Established Quoting the first clause alone inverts the paper’s point. Frontier This brief reports the direction and not the magnitude, because the text was gated and only the record was verified.

Established “Persona prompting reliably boosts creativity.” Wang and colleagues found that genius personas and demographic role prompts lift performance up to a threshold beyond which output becomes opposite to real-life patterns, and that broader prompt-engineering efforts yielded mixed to negative results. Frontier The standard practitioner technique has a measured failure mode and it is not widely known.

Established The skeptic-side error: “no machine has produced a genuinely new idea.” FunSearch improved a published lower bound on cap set capacity from 2.2180 to 2.2202 and produced a size-512 cap set at n=8; AlphaEvolve found a 48-multiplication scheme for 4×4 complex matrices, the first improvement over Strassen in that setting in 56 years. Established Both are in the literature with released artifacts and both are H-creative by Boden’s own criteria. Frontier What they are not is transformational: each was a search in a space whose objective a human specified completely.

Speculative The exotic claim, steelmanned: transformational creativity requires a world model rather than a text model. The argument, made in a 2023 Critical Humanities paper this brief reached as a verified record and could not read, is that transcending a conceptual space requires representing the space, and that a next-token predictor represents trajectories through a space it never explicitly holds. Speculative The prediction that follows is sharp and testable: systems with explicit structured world models should show qualitatively different creative behaviour, and the autonomous-discovery systems now shipping structured world models are the nearest available testbed and have never been evaluated for creativity. Handwave The mainstream reply is equally clean — that a sufficiently rich implicit representation is a world model by any behavioural test, and that demanding an explicit one smuggles in an architectural preference as a criterion. Handwave Neither side can currently be wrong, because the instrument that would adjudicate is the binding link in section 4, and that is the honest terminal position of this brief: the interesting question about machine creativity is not contested, it is unmeasured.