1 · Concept overview

Established A research front is a measurable object, and it was defined sixty years ago. Derek de Solla Price, writing in Science in 1965, observed that the citation network of scientific papers is not uniform: a small, recent, tightly interlinked slice of the literature does most of the work of the moment, and he called it the research front. Everything in this brief is an attempt to find that slice while it is forming rather than after it has been named.

Established Three families of method compete, and they use different raw material. Citation-structure methods cluster papers by who cites whom and look for clusters that are new, growing and coherent. Text methods look for terms, topics or concepts appearing at anomalous rates. Expert methods ask people. Each has a patent-side twin, and in practice most operational systems blend all three and then have a human write the label.

Frontier The claim this brief lands is that the methods are real and their validation is backwards. Almost every published evaluation selects a field already known to have emerged — graphene, CRISPR, deep learning — and shows the method would have flagged it. That is selection on the outcome. What does not exist, anywhere at scale, is a dated public register of emergence predictions with resolution criteria fixed in advance, scored afterwards by someone with nothing to sell. Without it, no sensitivity and no false-alarm rate can be quoted for any method in this brief, and none is.

Established Two structural limits bound the achievable performance regardless of method. Citation-based signals cannot fire until citations exist, which puts a floor of roughly three years between a result and its detectability; and the base rate of consequential emergence among nascent topics is low enough that a highly sensitive detector still returns mostly false alarms. Both are arithmetic, not criticism, and both are developed in sections 4 and 14.

Established The boundaries with two neighbouring briefs are drawn deliberately. Technology Forecasting owns rate and date forecasting — experience curves, S-curves, the tracked accuracy of dated predictions, the hype cycle — and none of that is re-derived here. Public Policy Foresight owns government foresight units, scenario practice and horizon scanning as an institutional activity. This brief owns the detection instruments themselves and the honest record of what they miss.

Established A note on sourcing. This brief was commissioned in September 2026 from the Institute’s research base. Reading-list entries without links are cited from the bibliographic record rather than re-fetched, and claims are dated no later than early 2026 unless carried by a linked source.

2 · Current scientific position

Established The field had no agreed definition of its own object until 2015. Rotolo, Hicks and Martin, reviewing the literature in Research Policy, found the term emerging technology applied to almost anything recent, and proposed five attributes that have since become the working standard: radical novelty, relatively fast growth, coherence, prominent impact, and uncertainty and ambiguity. The last attribute is the awkward one. A thing is emergent partly because nobody yet knows what it is, which makes the label unavailable at the moment it would be most useful.

Established Clustering the citation network works, and the choice of link type measurably matters. Klavans and Boyack compared direct citation, co-citation and bibliographic coupling on the same corpus and found direct citation produced the most accurate taxonomy of the literature. That is a rare thing in scientometrics: a methodological question settled by a head-to-head on a common dataset rather than by preference.

Frontier Applied to emergence, the same machinery returns a small and contested harvest. Small, Boyack and Klavans clustered recent literature by direct citation and applied explicit emergence criteria, reporting that only a small minority of clusters qualified. The counts are carried here from the bibliographic record rather than re-derived, and the deeper problem is that the criteria were the authors’ own: change the growth threshold and the harvest changes with it.

Established Burst detection is the one component with a clean mathematical foundation. Kleinberg’s 2002 algorithm models a document stream as a two-state process and identifies intervals where a term’s arrival rate jumps, with a cost for switching states so that noise does not trigger it. Chen’s CiteSpace made it a standard research tool by combining bursts with co-citation clustering and betweenness centrality. Both are honest instruments. Neither claims predictive validity, and the software’s own documentation is careful about that.

Established Novelty can be measured as unusual recombination, and it correlates with impact. Uzzi and colleagues analysed roughly 17.9 million papers in Science and found that the highest-impact work is overwhelmingly conventional in its combinations of prior knowledge while carrying an unusual tail: mostly familiar pairings, with a few genuinely atypical ones. Youn and colleagues found the same shape in patents, where the great majority of filings recombine existing technology codes and wholly new codes are rare.

Frontier The field’s most famous recent indicator turned out to be partly an artefact, and the sequence is instructive. Park, Leahey and Funk reported in Nature in 2023 that papers and patents have become less disruptive over time, using the CD index across roughly 45 million papers and 3.9 million patents. Petersen, Arroyave and Pammolli then showed in Quantitative Science Studies that the index is biased by citation inflation: reference lists have grown, and that growth alone drives much of the measured decline. Both papers are careful and both are peer-reviewed. The lesson for detection is that an indicator computed on a growing network is measuring the network’s growth until proven otherwise.

Established Patent-based detection rests on an indicator whose limits were catalogued in 1990 and have not changed. Griliches’s survey of patent statistics as economic indicators is still the reference: patents measure a legal act, not an invention; the value distribution is extremely skewed; propensity to patent varies by industry and by country. Moser’s work on nineteenth-century exhibition catalogues supplies the number that should end any argument — of the innovations displayed at the great exhibitions, only on the order of one in ten was patented at all, with the share varying by country and by industry. A detector reading patents was, in that period, blind to most of its subject.

Established The haystack grows faster than the needles. Bornmann and Mutz measure modern science growing at a few percent a year in publications and cited references, a rate that compounds to a doubling time comfortably inside a single career. Any detector with a fixed false-positive rate therefore returns a growing absolute number of false alarms every year while the number of genuinely consequential emergences does not obviously grow at the same rate.

Frontier Where a detection-adjacent correlation has been published, it has not always survived correction. Benson and Magee regressed patent metrics on measured improvement rates across technological domains and reported a best single-metric correlation of 0.76; a 2016 correction to the same paper replaced a figure and a table and rewrote the result to a Pearson correlation of 0.38 with a p-value of 0.043. Technology Forecasting owns the rate-prediction question this paper belongs to. It is carried here for a narrower point: the published correlations underpinning patent-signal detection are fragile enough that one recomputation halved one.

Established The only large scored corpus of dated technology predictions returns a third. A 2012 contractor study for a US defence sponsor extracted 1,055 forecasts, found roughly 70 percent in scope for scoring and about a third correct within a generous window. It has never been replicated under a different sponsor or coding rule. That corpus is the evidential floor of the whole enterprise, it is owned and analysed in the forecasting brief, and its existence is why the absence of an equivalent for emergence detection is conspicuous rather than excusable.

Frontier Asking experts does not stabilise the answer. Grace and colleagues surveyed 2,778 authors published at major machine-learning venues; between the 2022 and 2023 rounds the aggregate fifty-percent date for high-level machine intelligence moved from 2060 to 2047, and for full automation of labour from 2164 to 2116. The date question belongs to the forecasting brief. What belongs here is the implication for elicitation-based scanning: if a population of specialists moves its own central estimate by decades in a year, its judgement about which topics are emerging deserves the same error bars, and is almost never given them.

3 · Frontier questions

Frontier Can anything fire before the citations exist? Preprint servers, grant award records, conference programmes, job advertisements, code repositories, instrument orders and clinical-trial registrations all carry signal earlier than the citation network does. Several groups have built detectors on each. No published comparison scores them against a citation baseline on the same held-out future, so their claimed earliness is unaudited.

Frontier Do language-model scanners beat keyword and citation baselines, or only read faster? Since 2023 the obvious move has been to hand a model the literature and ask what is emerging, and many organisations have done it. The bottleneck is not reading. It is the label: there is no agreed target variable for the model to be right about, so the outputs are evaluated by plausibility. Speculative A model that has ingested the retrospective literature on emergence may also be reproducing its hindsight rather than detecting anything.

Frontier Is emergence one phenomenon? A new instrument, a new method, a new problem, a new community and a new funding line produce different citation signatures and probably need different detectors. The literature mostly treats them as one class, which is a plausible reason detectors trained on one type generalise poorly to another, and it is testable on existing data.

Frontier Does delayed recognition set a ceiling on recall? Ke, Ferrara, Radicchi and Flammini showed that delayed-recognition papers do not form a distinct class of sleeping beauties but sit on a continuous spectrum. Any detector tuned to early citation growth misses that tail by construction, and nobody has published what fraction of consequential results lies in it.

Frontier What happens to a detector once money follows it? An indicator used to allocate funding becomes a target. Keyword adoption, citation clubs and strategic framing of proposals are all cheap. No emergence detector has been tested for robustness under adversarial use, because none has yet been used at a scale worth gaming.

4 · Technological bottlenecks

Established There is an arithmetic floor on how early citation-based detection can be. A result must be written, reviewed and published; it must then accumulate enough citations for a burst or a cluster to be statistically distinguishable from noise. Publication alone commonly takes six to twelve months and citation accrual two or more years, so the earliest a citation detector can honestly fire is about three years after the work was done. Patent signals are worse: filings are not published for eighteen months by statute, before any citation accrues.

Frontier The base rate defeats precision, and the calculation is worth doing explicitly. Suppose one nascent topic in a thousand turns out to matter. A detector with ninety percent sensitivity and ninety-five percent specificity flags 0.9 true cases and about fifty false ones per thousand: roughly fifty-five false alarms for each hit, and under two percent precision. The numbers are illustrative because nobody has measured the base rate, and that is the point — without it, no detector’s output can be interpreted.

Established There is no ground truth, so sensitivity cannot be computed even in principle. Scoring a detector requires an agreed operational definition of a topic having emerged: a threshold on what, measured when, by whom. The field uses citation growth, which is the same signal the detector uses, so the test is partly circular. Independent outcome measures — products shipped, clinical approvals, standards published, jobs created — exist but have never been assembled into a scoring key.

Frontier Cluster identity is unstable between runs, which is a practical defect people underrate. Re-run a clustering pipeline on a corpus a year later and clusters split, merge and change membership; tracking a front across snapshots requires matching rules that are themselves parameters. Published emergence claims rarely report how sensitive the result is to those rules.

Established Data access shapes what can be detected. The comprehensive citation indexes are commercial products with licence restrictions on redistribution, so the most-cited detection results cannot be reproduced by anyone without a subscription. Open alternatives have improved coverage substantially in the 2020s but differ in disambiguation quality, and the two families do not always agree on whether a cluster exists.

5 · Research dependencies

Established Detection depends on identifier infrastructure more than on algorithms. Author, institution and funder disambiguation determines whether a cluster is a research community or an artefact of name collisions. Persistent identifiers for people, organisations and outputs are the unglamorous precondition, and coverage is uneven outside a handful of well-resourced countries and disciplines.

Established It depends on open corpora with stable snapshots. A scored detector needs a frozen corpus as of a date, so that a later evaluation can establish what was knowable then. Commercial indexes are updated in place and retroactively corrected, which quietly destroys the counterfactual.

Frontier It depends on the statistics of rare-event detection, which the field borrows loosely. Precision-recall analysis, calibration and proper scoring rules are standard in other prediction settings and largely absent here. Adopting them is not a research problem; it requires only a labelled outcome set, which is exactly what is missing.

Speculative It may come to depend on automated research systems as both subject and instrument. If systems of the kind surveyed in Artificial Scientists begin generating a material share of hypotheses and papers, the citation record changes character: volume rises, provenance blurs, and the assumption that a growing cluster reflects human attention weakens.

6 · Required experiments

Frontier The decisive experiment is a prospective register of dated emergence forecasts with resolution criteria fixed in advance. Several teams publish, on the same frozen corpus and the same date, a ranked list of research fronts they predict will be consequential, together with the outcome measures and thresholds by which they agree to be judged. An independent body holds the lists and scores them at five and ten years. That single design would produce the first honest sensitivity and false-alarm rates the field has ever had, would let citation, text, patent and expert methods be compared on identical terms, and would settle whether any of them beats a simple baseline of extrapolating current publication growth. Nothing about it requires new science; it requires a custodian and a decade of patience, and no funder has supplied either.

Established The cheapest version can be run today on lists that already exist. The Government of Canada publishes dated emerging technology trend cards; other governments and vendors publish comparable inventories. Scoring last decade’s published lists against outcomes now — retrospectively, but against documents whose dates are fixed and public — costs a research assistant a summer. It would be evidence about the output, not validation of the method, and the distinction matters: an inventory is not a forecast unless it names what would count as being wrong.

Frontier A third design is a held-out-future bake-off. Freeze a corpus at a past date, hide everything after it, and have a language-model scanner and a citation-burst baseline each produce a ranked list. Score both against what actually happened. This is cheap, repeatable and uses only archived data, and it is the only way to find out whether model-based scanning adds anything beyond faster reading.

Speculative A fourth is an adversarial test nobody has run. Offer a modest prize for moving a named topic up a published detector’s ranking within a year, by any legitimate means. How cheaply it can be done is a direct measure of how much weight the indicator can bear in a funding decision.

7 · Engineering requirements

Established A working detector is a data pipeline with four hard stages and one easy one. Ingest and deduplicate records; disambiguate authors, institutions and funders; build the citation graph and cluster it; track cluster identity across snapshots; and finally rank. The ranking, which is where the published papers concentrate, is the easy stage. The disambiguation and the cross-snapshot tracking determine whether the output means anything.

Established Scale is a real but solved constraint. Clustering a global citation graph of hundreds of millions of nodes by direct citation is routine on commodity hardware with modern graph libraries, and the incremental update problem is tractable. Computation is not what limits this field.

Frontier Versioning is the requirement most systems skip. To be scoreable, a detector must publish immutable, dated snapshots of both its corpus and its output, with parameters recorded. Almost no operational system does this, which is why almost no operational system can be audited after the fact.

Frontier The human label remains load-bearing and is rarely costed. Every system this brief examined ends with an analyst writing a name and a paragraph for each candidate front. That step encodes judgement, absorbs most of the marginal cost at scale, and is where the vendor’s prior enters the product.

8 · Adjacent technologies

Established The two boundaries that define this brief are with forecasting and with foresight. Technology Forecasting owns extrapolation: how fast an existing technology improves, when a capability arrives, and the scored record of dated predictions. Public Policy Foresight owns the institutional practice: government units, scenarios, national risk registers, and the finding that the field stopped measuring its own outputs. This brief owns the instruments that decide what goes on the list in the first place.

Established Within the map it also touches four others. Innovation History owns invention and its measurement, including the patent-series audits this brief leans on; Scientific Revolutions owns the account of change that a detected front is supposed to be an early sign of; Scientific Funding Models owns what happens when a detector is wired to an allocation rule; and Technology Trees of Civilization owns dependency structure, which is the alternative to detection — inferring what must come next from what a field requires rather than from what it is citing.

Frontier Outside the map: scientometrics and information science, which supply every method here; the economics of innovation, which supplies the patent-indicator critique; epidemiology and signal detection theory, which supply the precision-recall discipline the field lacks; and research-security practice, which is the largest present customer.

9 · Institutional requirements

Established The largest operational customer for emergence detection is security, not opportunity, and that changes the instrument. Canada’s emerging technology trend cards are published under research-security guidance, as an aid to assessing collaboration risk. A detector built for screening optimises against false negatives, because a missed dual-use technology is the costly error; a detector built for funding optimises against false positives, because backing a dead end wastes money. These are different machines, and the same list cannot serve both without saying which error it is minimising.

Established The suppliers are interested parties and should be labelled as such. The comprehensive citation indexes, the emerging-topic reports and the annual technology lists are commercial products of a small number of analytics firms, sold to the institutions whose performance the same firms’ indicators are used to assess. (Vendors on their own instruments.) That is not a charge of bad faith; it is the reason an independent scoring body is the load-bearing missing institution.

Frontier No institution has an incentive to publish its own miss rate. A vendor loses a selling point, a ministry loses deniability, a research council invites an audit of its portfolio. The result is a field in which every participant can point to a hit and none can produce a denominator, and this is a governance problem rather than a methodological one.

Established This Institute’s own longlist is an instance of the object under study. A curated map of research frontiers is a dated emergence claim, and treating it as one imposes obligations: record when each entry was added and on what evidence, keep the entries that went nowhere, and publish the resolution criteria in advance rather than reconstructing them afterwards. The cost is embarrassment. The benefit is that the list becomes evidence about detection rather than another undated inventory.

10 · Ethical & societal considerations

Frontier Emergence lists have become instruments of control as well as of curiosity. When a technology appears on a national research-security list, collaborations are reviewed, visas are scrutinised and funding conditions change. The people affected are individual researchers who had no part in the classification and no route to appeal it, and the classification rests on methods whose error rates are, as this brief argues throughout, unmeasured.

Frontier False positives are not evenly distributed. Screening attaches to fields, institutions and, in practice, national origins. A method with an unquantified false-positive rate applied to people rather than to topics transfers the whole cost of its imprecision to the least powerful party in the system.

Established Detection is reflexive: publishing a front can help create it. A named emerging field attracts grant applications, special issues and job postings, all of which generate the citation growth the detector was looking for. This is not hypothetical; it is the ordinary mechanism by which research agendas form, and it means a detector in wide use is partly measuring its own output.

Speculative Concentrating attention has an opportunity cost that is never counted. Every front promoted is a set of unfashionable programmes not funded, and the delayed-recognition literature says some of those are where the returns were. No allocation system has been evaluated for what it forgoes, only for what it backs.

11 · Civilizational implications

Frontier The strongest argument against the importance of this whole enterprise is that discovery is robust to being missed. Independent and near-simultaneous discovery is common enough that the standard historical question is not whether something would have been found but when and by whom. If the multiples literature is right, a detector that misses a front costs its owner priority rather than costing the world the result.

Established The counter-argument is about positioning, not about knowledge. Being early confers standards influence, supply-chain placement, trained people and regulatory agenda-setting. Those are the returns a state or an institute is actually buying when it funds scanning, and they are distributional rather than epistemic. Stating it plainly clears away the pretence that horizon scanning is a search for truth.

Frontier Detection capability is unequally held, and it compounds. The comprehensive indexes, the analysts and the compute sit in a small number of countries. A state without them is not merely slower to spot a front; it is dependent on lists produced elsewhere, under loss functions set elsewhere.

Speculative If automated research systems raise output by an order of magnitude, detection becomes the binding constraint on attention rather than a convenience. A literature no human community can read at all is one where whatever filter is in use is effectively the research agenda. The filter’s unmeasured error rate then stops being a methodological footnote.

12 · Timelines

These horizons track the evidence base for early detection, not the pace of the technologies it is pointed at.

  • 10 yr: Frontier Open bibliographic corpora with dated immutable snapshots become the default substrate, making held-out-future evaluation cheap; at least one government or funder publishes a scored retrospective of its own past emergence lists; language-model scanning is ubiquitous and still, on present trends, unvalidated against a citation baseline.
  • 25 yr: Speculative Either a prospective register exists and the field has real sensitivity and false-alarm numbers, or the practice has been absorbed entirely into research-security screening, where nobody publishes error rates by design. These two futures are distinguishable now by whether any funder commissions the register.
  • 50 yr: Speculative If machine-generated research is a large share of output, detection shifts from finding fronts in a human literature to allocating limited human attention across a machine one, and the methods in this brief are replaced rather than improved.
  • 100 / 250+ yr: Handwave Any claim at this range about what a research front is presupposes that science remains an activity organised around citing publications. That format is about 360 years old and has already changed twice in the last thirty.

13 · Technology tree & dependencies

  • Depends on nothing on this map in the strict sense: no brief here produces a result that early detection is waiting on. The methods are published and the mathematics is settled. What it does consume is the scored-forecast record assembled in Technology Forecasting, which supplies the only large corpus of dated predictions anyone has graded, and the institutional diagnosis in Public Policy Foresight, which explains why the scoring stopped.
  • Requires (not on this map) a public register of dated emergence forecasts whose resolution criteria are fixed before the fact, because without one no method in this brief has a sensitivity; an agreed operational definition of a topic having emerged, measured by something other than the citation growth the detectors already use, or the evaluation stays circular; open citation and patent corpora published as dated immutable snapshots, since indexes corrected in place destroy the counterfactual a later scoring needs; author and institution disambiguation at global coverage, without which clusters are partly name collisions; a scoring body with no detection product of its own to sell; and a buyer willing to pay for a measured miss rate rather than for another annual list of twelve technologies.
  • Enables defensible portfolio allocation by research funders; research-security screening that can state its error rate; early standards and regulatory engagement; and the maintenance of any curated frontier map, including this one, as something scoreable rather than merely published.
  • Adjacent to Innovation History for the patent-indicator record, Scientific Revolutions for what a front is an early sign of, Scientific Funding Models for what happens when a detector allocates money, and Technology Trees of Civilization for the dependency-based alternative to detection.

14 · Common misconceptions & speculative claims

Handwave “Bibliometrics predicted CRISPR, graphene and deep learning.” Those are the cases the methods were demonstrated on, after the fact. A demonstration that a method flags a known success is a fit, not a forecast, and it carries no information about how many other clusters the same settings would have flagged that went nowhere. Until someone publishes the full ranked list as of a past date, this claim has no content.

Established “An emerging technology list is a forecast.” Almost none of them are. A forecast names what would count as being wrong; an inventory names what looks interesting now. Most published lists — governmental, commercial and academic — carry no thresholds, no dates by which the named technology should have arrived, and no prior list against which the current one can be compared. They are useful documents and they are not predictions.

Established “Citation growth measures importance.” It measures attention, which correlates with importance and is not it. Fields grow for reasons including fashion, funding calls, low entry cost and controversy. The delayed-recognition literature shows the converse directly: work whose value is recognised late is common, and it sits on a continuum rather than forming a rare exceptional class.

Frontier “The disruption index shows science has stopped being disruptive.” The Nature result is real and carefully done, and a subsequent peer-reviewed analysis shows the index is biased by citation inflation, with growing reference lists accounting for much of the decline. The honest reading is that the measured decline is partly an artefact of how the indicator interacts with a growing network, and that the underlying question remains open.

Established “Patent landscapes show where the innovation is.” They show where the patenting is. Griliches catalogued the gap in 1990 and it has not closed: propensity to patent varies enormously by sector and country, patent value is extremely skewed, and the historical exhibition data indicate that only on the order of one in ten innovations was patented in the periods where an independent count is possible. Software, methods and process improvements remain systematically under-represented.

Frontier “Language models have solved horizon scanning because they can read everything.” Reading was never the constraint. The constraint is the absence of a labelled outcome, which means a model’s output can be assessed only for plausibility — precisely the failure mode it is best equipped to exploit. Speculative A model trained largely on retrospective accounts of how fields emerged may be very good at producing narratives of emergence and no better than a citation baseline at identifying one.

Frontier “Ask the experts in the field; they know what is coming.” The best-documented elicitation in an adjacent area moved its own aggregate central estimate by thirteen years for one question and forty-eight for another between consecutive annual rounds. Specialists are the right people to ask what a result means. Treating their aggregate as a calibrated instrument requires evidence of calibration, and in this domain nobody has collected any.

Speculative “Better detection would have prevented the big misses.” The historical misses that get cited — mRNA therapeutics funded thinly for years, neural networks written off as a dead end — were not primarily failures to see the work. The work was published and visible. They were failures of judgement about whether it would scale, which is a different problem and not one any of the methods in this brief addresses.