1 · Concept overview

This is the applied, near-term end of augmenting human cognition: the things people actually use. Notebooks and note systems. Search, and the habit of not knowing something because you can look it up. Spaced review and self-testing. Checklists and decision aids. Spreadsheets, version control, integrated development environments, and now language models sitting inside all of them. Nothing here is implanted, ingested or wired. The question is whether any of it makes a healthy adult measurably better at thinking, by how much, on whose test, and at what cost to the capacity it replaced.

The honest summary is that the established core of this subject is thin, mostly negative, and considerably more interesting than the positive claims that get repeated. The techniques with the strongest evidence are behavioural, unglamorous and free. The technologies with the loudest claims have the weakest instruments. And the field’s most reliable single finding is a measurement artefact: when an intervention is assessed on a test it helped shape it looks transformative, and when it is assessed on a test it did not shape it usually looks small.

Scope, because this cluster divides its labour deliberately. Intelligence Amplification carries the Engelbart and Licklider tradition and the theory of augmentation as a research programme. Human-AI Integration carries the complementarity evidence and the anchor meta-analysis on when a human-plus-machine pairing beats its better half. Cognitive Enhancement carries the pharmacological and electrical routes and their nulls. Distributed Cognition carries the philosophical dispute about where a mind ends. This brief owns the tools and the practice: what is on the desk, what the measured gain is, who it helps, and what it costs. It also owns the pivot question of the subject — whether offloading a capacity to a tool is a loss — on which the literature is genuinely mixed and both sides get a hearing below.

2 · Current scientific position

Established The most consequential number in this subject is not an effect size but a ratio between two ways of measuring the same intervention. Kulik and Fletcher’s meta-analytic review of intelligent tutoring systems gives a median effect of 0.66 across fifty controlled evaluations, then splits it by instrument: 0.73 on studies using local tests, 0.13 on studies using standardised tests. Established That is a 5.6-fold collapse produced by nothing but the choice of measuring stick, and the reviewers’ own conclusion names it: alignment of test and instructional objectives is a critical determinant of evaluation results. Established The same collapse appears independently in the mastery-learning literature, where an average effect near 0.59 was reported alongside the observation that these dramatic effect sizes essentially disappeared when standardised tests were used. Frontier No equivalent split has been published for note-taking software, search interfaces, retrieval tools or AI writing assistants, because almost none of those studies use an independent instrument at all. Speculative The prior to hold is that the same deflator applies here: a gain measured on the task a tool was built for should be expected to shrink several-fold when measured on an outcome it did not shape. That is an expectation, not a result, because nobody has run the comparison.

Established Cognitive offloading is the pivot of this subject, and the flagship evidence that it damages memory has failed to replicate twice. Sparrow, Liu and Wegner’s 2011 Science paper reported that people are primed to think of computers when asked general-knowledge questions even when they know the answer, that they do not tend to remember information they believe will remain available to look up, and that they remember where information sits far better than the information itself. Established The replication record: a 2018 study could not be successfully replicated, and a 2020 replication incorporating methodological recommendations found no conclusive evidence for it. Frontier Storm and Schooler showed the effect is replicable, but only when participants had prior experience that saved information would remain accessible — which converts a claim about human memory into a claim about rational trust in a filing system. Established A 2024 meta-analysis by Gong and Yang across thirty-five studies frames its outcome as changes in how people process and remember information. Established Changes are not damage. Frontier A conditional, twice-unreplicated priming effect has been upgraded into a civilisational claim, and the upgrade happened in the popular literature rather than the empirical one.

Established The case that offloading is normal rather than injurious is older and better grounded than the case against it. Clark and Chalmers’ 1998 argument — Otto with his notebook and Inga with her biological memory are functionally equivalent, so the notebook is part of Otto’s cognitive system — treats the offload target as constitutive of cognition rather than a threat, and the standing objection to it is philosophical rather than empirical. Speculative Wegner’s transactive memory is the pre-digital version; no primary citation for it was obtained and it is stated here as background. Speculative The strongest version of the benign reading is historical: writing, numerals, notation and indexes are all offloading, each was denounced on arrival, and the measured cost falls on the specific capacity offloaded — which is the point rather than the problem. Handwave That a worry is old does not make it wrong; the Phaedrus is too often cited as though recurrence were refutation. Frontier The adjudication offered here is that the historical analogy holds for tools that fail loudly and breaks for tools that fail silently. See Distributed Cognition.

Established Where offloading does measurably hurt, the harm has a repeatable structure: the tool fails without signalling, the human’s residual skill has decayed, and the human is nominally responsible for catching the failure. The cleanest instance is a computer-aided detection study in which cancers diagnosed in 46% of cases without automated aids were found in only 21% of cases when an aid was present and had missed the finding — a twenty-five-point absolute drop caused by a tool that said nothing. Established The same literature reports that decision support improved clinicians’ answers by 21 percentage points, from 29% to 50%, while 7% of previously correct answers were changed to incorrect ones. Established Automation bias has a settled taxonomy underneath those numbers — commission errors from redirected attention and from discounting information that contradicts the aid, omission errors from failing to notice what the automation missed — and training reduces the first but not the second. Frontier In endoscopy, a 2025 multicentre study reported adenoma detection in unassisted colonoscopy falling from 28.4% to 22.4% after routine exposure to an AI polyp detector; this brief read that figure in a secondary source rather than the paper, and four correspondence responses followed within months, so treat it as contested. Established Saturation is the same failure inverted: 566 catalogued deaths from ignored alarms between 2005 and 2008, and roughly 8,000 track-circuit alerts a week before the 2009 Washington Metro collision, which investigators concluded would have thoroughly desensitised the dispatchers. Frontier An augmentation that speaks constantly is functionally an augmentation that is silent.

Established The best-evidenced cognitive interventions available to an individual today are retrieval practice and spacing, and they are free. Testing knowledge produces better learning, transfer and retrieval than study built on recognition — re-reading, highlighting — repeated testing produces superior transfer relative to repeated studying, and classroom effects have been shown to last years in real educational contexts. Established Spacing increases the effect, with greater impact as the delay grows, at the cost of forgetting if it runs long. Frontier This brief could not obtain meta-analytic effect sizes for either. The source it could fetch gives conclusions and not numbers, so no d or g appears here. That matters, because the interesting claim — that a schedule beats every electrical, pharmacological and neural intervention on the market for durability and cost — needs a common-metric comparison nobody has run. Established Alongside it sits cognitive load theory: the worked-example effect, whereby novices learn more from studying solved problems than from solving them, and the expertise-reversal effect, whereby the same support becomes a hindrance once the learner has the schema it was scaffolding. Frontier Expertise reversal is the most under-used idea in the field, because it predicts that every augmentation has a skill level above which it subtracts — exactly the shape the workplace evidence turns out to have. Speculative The deliberate-practice framework and the meta-analytic challenge to its explanatory share of expert performance are the obvious companions; this brief obtained neither and prints no figures from them.

Established Decision support has the largest deployed base and the most sobering result: it improves process far more than outcomes. Two 2005 reviews found clinical decision support improved practitioner performance in 64% of studies and patient outcomes in 13%; a 2014 review found no benefit in risk of death when such systems were combined with the electronic health record. Established The 64%-process against 13%-outcome gap is the number to carry into any claim that a decision aid helps someone think. Established The counterweight is real: the German PRAIM implementation, 463,094 women screened across twelve sites, reported breast-cancer detection of 6.7 per 1,000 with AI support against 5.7 without — a 17.6% relative increase, 95% CI +5.7% to +30.8% — while recall fell from 38.3 to 37.4 per 1,000. Frontier That is the strongest evidence for genuine complementarity in a decision task, and it is observational rather than randomised. Established In dermoscopy, good-quality support beat both the AI and the physician alone and the least experienced gained most — while the same paper found faulty AI misleads the whole spectrum, experts included.

Established Software as a thinking aid is where the recent measurements are, and the two flagship results in the same intervention class have opposite signs. Developers with GitHub Copilot completed a from-scratch JavaScript HTTP server 55.8% faster than controls. Established The METR study randomised sixteen experienced developers across 246 tasks in mature repositories they had worked in for around five years, and measured a 19% increase in completion time — while those same developers forecast a 24% speedup and afterwards still estimated they had been 20% faster. Established The moderator is task context, not model quality, and the self-report gap is a finding in its own right: people cannot introspect on their own tool-assisted performance, which makes user satisfaction worthless as an efficacy measure. Established Where the workplace numbers are positive they are positive at the bottom of the distribution. Across 5,172 support agents output rose 15% per hour, with lower-skilled workers improving both speed and quality while the most experienced saw small gains in speed and small declines in quality. Established Tutor CoPilot, the first randomised trial of an AI-assisted human tutoring system, covered 900 tutors and 1,800 students: 4 percentage points more topic mastery overall, 9 points for students of the lowest-rated tutors, at 20 dollars per tutor per year, with over 550,000 messages showing assisted tutors more likely to ask guiding questions. Frontier The defensible generalisation is that cognitive augmentation which works is variance reduction, not mean elevation: it lifts the bottom of a distribution toward an existing standard and the mean rises as a by-product. Frontier The live counterexample is a Harvard physics crossover trial — 194 eligible students, AI-tutored median post-score 4.5 against 3.5 for in-class active learning, regression effect 0.63 and quantile estimates from 0.73 to 1.3 standard deviations, in 49 minutes rather than about 60 — and that population is not below the mean. Frontier It also used a purpose-designed post-test, which is this section’s opening caveat restated.

Established The ceiling on all of it is a preregistered meta-analysis of 106 studies and 370 effect sizes, and it is negative. Combinations performed worse than the best of humans or AI alone, g = −0.23, 95% CI −0.39 to −0.07; against humans alone they were clearly better, g = +0.64. Established Both are true at once. Established When humans outperformed the machine, combining helped, g = 0.46; when the machine outperformed humans, combining hurt, g = −0.54. Established And the bias tests run opposite to assumption: none detected for the negative headline comparison, clear evidence of it for the positive one. Frontier The pessimistic result is bias-clean and the optimistic one is bias-inflated. The rule to take away is that an aid’s value is set by which party is better at the task, and that is knowable in advance.

3 · Frontier questions

Frontier The first EEG study of language-model-assisted cognitive work exists, it is being over-read in both directions, and its critical comparison rests on eighteen people. Kosmyna and colleagues at the MIT Media Lab had participants write essays under three conditions — language model, search engine, no tools — with 54 participants across the first three sessions and only 18 completing the fourth, in which conditions were swapped. Brain-only participants showed the strongest and most distributed networks, search users moderate engagement, model users the weakest connectivity; participants moved from the model condition to the unassisted one showed reduced alpha and beta connectivity, and model users struggled to accurately quote their own work. Established The abstract does not quantify that quoting failure. Frontier The authors describe potential cognitive costs, not established harm, and they are right to: EEG connectivity is not a learning outcome, the paper is a preprint, and cognitive debt names a construct requiring a longitudinal design the study does not have. Speculative Settling it needs a dose-response relationship between exposure and later unassisted performance, persistence after exposure ends, and a behavioural outcome rather than a connectivity measure.

Frontier Withdrawal effects are the genuinely new hazard and there is one serious result. A field experiment with roughly a thousand high-school mathematics students reported that when access to an unguarded generative assistant was subsequently removed, students performed worse than those who never had access — a grade reduction this brief does not print, because the pack verified the paper’s record and abstract through a metadata registry rather than reading the article. Established The title changed on peer review from a flat claim that generative AI can harm learning to a conditional form naming the absence of guardrails, and the study’s structure includes a guarded tutor arm that did not show the harm. Frontier Nothing in the older technology-in-education literature measured withdrawal, because none of those tools created a dependency worth measuring. Speculative If withdrawal effects are real and general, they are the most policy-relevant finding in this brief, because every deployment decision becomes a commitment rather than a trial.

Frontier The methodological frontier for the whole subject is publication-bias adjustment, and it has already demolished one headline result. A 2025 meta-analysis of fifty-one studies reporting a large positive effect of ChatGPT on learning performance was retracted in April 2026, eleven months after publication. Independently, a robust Bayesian re-analysis reported that the effects greatly diminish once publication bias is accounted for and the evidence in favour of the benefits disappears; a second, separate re-analysis put the higher-order-thinking outcome at a value statistically indistinguishable from zero. Established This brief read the retraction banner and date directly; it did not read the re-analyses, whose records were verified through a registry, and prints their conclusions as reported rather than their numbers. Frontier The same group has demonstrated the method eliminating a well-established behavioural effect across seventy-nine randomised trials, where adjustment did not merely null the effect but produced moderate evidence against it. Speculative Nobody has run this on the brain-training, working-memory-training or stimulation corpora. The software is packaged and free. It is the cheapest decisive study in this cluster and it is unrun.

Frontier One field experiment has instrumented the interaction rather than the output, and it is the closest thing to a mechanism. Across 2,234 participants and 11,024 advertisements, human-AI teams produced 50% more output per worker, with 25% more task-oriented messages, 18% fewer interpersonal ones, 17% more delegation to the machine than to a human partner, and 62% fewer direct text edits — and more homogeneous, self-similar outputs. Speculative Output homogenisation is not cognitive homogenisation, and the inference from one to the other is exactly the sort this brief blocks elsewhere — but no measurement of reasoning diversity under shared augmentation exists at any scale.

Established The restorative neurotechnology results are the strongest thing in the wider augmentation field, and every one of them is n = 1. Two 2023 speech neuroprostheses reached usable rates: intracortical arrays at 62 words per minute with a 9.1% word error rate on a fifty-word vocabulary and 23.8% on 125,000 words; surface recordings at a median 78 words per minute and 25% error on 1,024 words, with synthesised speech in the participant’s pre-injury voice and avatar control from one decoder stack. Established The reference survey of the field records no evidence of brain-computer interfaces used for cognitive enhancement in healthy, non-impaired individuals. Handwave A non-restorative channel — information into or out of a healthy brain faster than existing sensorimotor bandwidth — has no evidence base and no agreed experimental design. See Brain-Computer Interfaces; the point here is that the applied end of cognitive augmentation currently has nothing to do with the brain.

4 · Technological bottlenecks

Established The binding constraints here are measurement practices rather than capabilities, which is unusual and worth saying plainly. Three links bind, in order. Near transfer binds first. The canonical review of brain-training programmes found some evidence of improvement on trained tasks, less evidence that improvement generalises to related tasks, and almost no evidence of generalisation to everyday cognitive performance. Sixty years of work has not cleanly cleared the second tier for any intervention aimed at healthy adults, and everything downstream is speculative until one does. Frontier Note what that does not say: it does not say the interventions do nothing. It says the something they do stays inside the trained task.

Established Far transfer binds second and has bound longest. A gain on an ecological outcome — job performance, error rate, a grade on an instrument the intervention did not shape — is what the field cannot produce. Frontier The workflow results above are the closest anything has come, and they are gains on the task the tool was built into rather than gains in the person using it. The distinction is not pedantic: an agent resolving 15% more tickets with an assistant has not been shown to be a better reasoner, and no study has tested whether they are.

Frontier Absence of a measured decrement binds third, and it is the link most often skipped. Almost no augmentation study is powered for, or even reports, the capacities that got worse. Expertise reversal and the workplace results both predict failure here: the aid that lifts a novice degrades an expert, and in the METR case the degradation was invisible to the people experiencing it. Established Making decrement reporting mandatory requires no new technology and no new money.

Established Two instrument-level bottlenecks sit underneath all three. The first is test alignment: until a study reports both a local and an independent outcome its effect size is uninterpretable, and the 0.73-against-0.13 split says how much is at stake. Frontier The second is publication bias, now shown to inflate at least one adjacent literature to the point of reversal under adjustment. Speculative A third and softer bottleneck is that the field has no agreed battery — no pre-registered set of uncorrelated cognitive domains plus one ecological outcome against which a broad-spectrum claim could fail. Brain training’s entire failure was a failure at that link: the task gains were real and the battery was the wrong one. Handwave None of the three binding links is a capability constraint. All three would yield to money spent on measurement, which is an unusual position for a frontier subject to be in and the main reason this brief is optimistic about the next decade.

5 · Research dependencies

Established This brief waits on one result another brief on this map produces, and it is a legal result rather than a scientific one. The tools described here are increasingly deployed by employers, schools and clinical systems rather than chosen by individuals, and they read and shape cognition as a condition of the job. Whether a person may be required to use a cognitive aid, permitted to refuse one, or assessed on their unassisted performance is upstream of every deployment question in this brief, and it belongs to Cognitive Liberty. Until that question has an answer, the engineering recommendations below are advice rather than specification.

Frontier Three further dependencies are evidential rather than typed. The complementarity boundary — which party is better at the task, and therefore whether pairing helps or hurts — is established in Human-AI Integration, and this brief applies it rather than re-deriving it. Established The pharmacological and stimulation routes are adjudicated in Cognitive Enhancement; their nulls are cited here as background and not re-argued. Frontier And the measurement problem running through everything here is the one Intelligence Measurement treats: a capability claim is only as good as its instrument, and no instrument in this field has been validated the way a psychometric test would be.

Speculative The dependency nobody is producing is a common-metric comparison across intervention classes — schedule against software against stimulation against drug, on one battery, with one durability window and one pre-registered equivalence bound. Handwave Every component of that study is cheap and none of it is technically hard. It has not been run because no discipline owns it, and that is a sociological fact about research funding rather than a scientific obstacle.

6 · Required experiments

Established Experiment one, and it needs no new data. Apply publication-bias adjustment — selection models, PET-PEESE, robust Bayesian multilevel meta-analysis — to the existing brain-training, working-memory-training and stimulation corpora. Measurement: the adjusted pooled effect and its credible interval per corpus. The tooling is packaged in R and in a free graphical statistics package, the method has already eliminated one established behavioural effect across seventy-nine trials, and the expected outcome is that several remaining positives disappear. Frontier This is the highest value-per-dollar action available in the subject and it could start this week.

Established Experiment two: dual-instrument reporting, made a condition of publication. Every study of a cognitive tool reports its outcome on both a local measure and an independent one, with the ratio published. The tutoring literature already knows its ratio is around 5.6 to 1; nobody knows it for note systems, search, or AI assistants, and until they do the field has impressions rather than numbers.

Frontier Experiment three: the silent-failure trial. Randomise participants to an aid that fails silently against an aid of identical accuracy that announces its own uncertainty when it fails, hold residual-skill practice constant, and measure both assisted performance and unassisted performance after withdrawal. Speculative If the harm literature is fully explained by silent failure, the signalling arm shows no decrement and the mitigation programme for the whole subject becomes an interface specification. This is the falsifier for the strongest positive hypothesis in this brief and nobody has run it.

Frontier Experiment four: decrement-powered trials. Pre-register an equivalence bound on every domain the intervention is not aimed at, and power the study to detect degradation rather than only improvement. Speculative Experiment five: withdrawal as a standard endpoint — unassisted performance measured at cessation, six months and twelve, in every deployment study. Speculative Experiment six: the variance test. Pre-register the prediction that gains fall as baseline skill rises, and report the interaction rather than the mean; four independent results already show the pattern and not one of them was designed to test it. Frontier Experiment seven: the retrieval-first interface. Build a note or search tool that withholds a stored answer until the user attempts recall, deploy it against ordinary lookup, and measure retention at six months. Handwave That is the one experiment in this list which could be run by a software company rather than a laboratory, and none has.

7 · Engineering requirements

Frontier If the silent-failure conjecture holds, most of the engineering agenda for cognitive augmentation is uncertainty communication, and it is unglamorous. A tool that reports calibrated confidence, marks the span it is least sure about, and degrades visibly rather than gracefully is buildable today. The obstacle is commercial rather than technical: visible uncertainty makes a product feel worse in a demonstration and better in a decade.

Established Alert budgets are a hard constraint, not a preference. Eight thousand alerts a week desensitised trained dispatchers, and 566 catalogued deaths followed ignored alarms over three years. A decision aid that fires on everything has spent its channel. The design question is not how to notify but how many notifications a human attention budget can carry per shift, and that number belongs in a specification rather than emerging by accident.

Frontier The most useful unbuilt artefact is a note system that schedules retrieval rather than storage. Personal knowledge management software is overwhelmingly optimised for capture and linking; the intervention with the best durability evidence in this brief is scheduled self-testing, and almost no general-purpose note tool does it by default. Speculative A tool that withheld a stored answer until the user attempted recall would implement the testing effect at the point of use, and its measured effect against ordinary lookup is unknown because nobody has built and evaluated it at scale.

Speculative Two further requirements follow directly from the workplace results. An assistant should know the user’s skill level on the current task, because the same support helps a novice and degrades an expert and the crossover point is measurable. Frontier And it should sample unassisted performance periodically, because the self-report gap means neither the user nor the employer can otherwise tell which way the tool is pushing. Handwave Neither requirement appears in any shipped product this brief is aware of.

8 · Adjacent technologies

Established The nearest technical neighbours are the restorative implanted-interface programmes covered in Brain-Computer Interfaces and Neural Interfaces, which supply this subject’s only headline capability numbers and none of its healthy-adult evidence, and the biological routes in Cognitive Enhancement and Neuroplasticity Engineering.

Established Human factors and aviation automation research is the oldest adjacent field and the most useful, because it has fifty years of documented instances of exactly the failure this subject is walking into: automation policy producing skill decay, out-of-the-loop operators, and handoffs designed without regard for what the human was doing thirty seconds earlier. Frontier It also supplies measurement apparatus — situation-awareness probes, recency-based skill assessment — that cognitive-tool studies almost entirely lack, and one finding transfers directly: manual skill decays with recency of practice rather than with total experience, which means an expert who has not done the task unaided in a year is not an expert at doing it unaided.

Established On the systems side, Distributed Cognition and Collective Intelligence carry the group-level form of every question raised here, and Future Education Systems carries delivery at population scale, including the two-sigma target that has driven fifty years of investment and the technology that has never come close to it. Frontier Artificial Creativity is adjacent in an awkward way: the one task class where the anchor meta-analysis found combinations doing better than decision tasks is content creation, and that effect was not itself statistically distinguishable from zero.

Frontier Finally, information retrieval is the adjacent technology nobody in this field treats as one. Search is the most widely used cognitive augmentation in history, its measured effects on memory are the contested literature at the centre of this brief, and its interface design has never been evaluated as a cognitive intervention with a control group. Speculative A ranking change shipped to a billion people is an uncontrolled experiment on distributed human memory, and nobody — inside or outside the companies running it — measures the outcome.

9 · Institutional requirements

Established There is regulatory precedent here and it is consumer-protection precedent, not medical. The US Federal Trade Commission found brain-training marketing deceptive and imposed a judgment against the leading vendor of 50 million dollars, reduced to 2 million for inability to pay; the reduction was financial and the finding stands. Frontier That is currently the main institutional fact about the cognitive-enhancement market: the enforcement route that has actually bitten is advertising law, not device regulation or clinical approval.

Established The field’s own consensus machinery has failed in a documented and instructive way. In 2014 more than seventy scientists signed a statement that brain games cannot be scientifically proven cognitively advantageous, and more than a hundred and thirty signed the opposite; the evidence favoured the seventy. Frontier A signatory count is not evidence, and this is the cleanest available case of a numerically lopsided consensus statement pointing the wrong way in a live field. Speculative Any institution proposing to procure cognitive tools on the strength of expert endorsement should read that episode first.

Frontier Three actions are available immediately and none requires new science. First, mandatory decrement reporting: any study or product claiming a cognitive benefit states what got worse, against a pre-registered equivalence bound. Second, dual-instrument reporting as a condition of publication and of public procurement. Third, pre-registration, which the tutoring and education literatures have already begun adopting through public trial registries and which should make the 2029 evidence base substantially better than the 2025 one.

Speculative A fourth is harder and matters more: workplace deployment of cognitive tools currently carries no audit requirement for skill retention. Aviation learned to mandate manual-handling practice after accidents; medicine is arguing about the equivalent now. Frontier An employer deploying an aid that degrades unassisted competence has created a liability nobody measures, and the endoscopy correspondence exchange is the first visible sign of a profession noticing. Handwave A retention audit — periodic unassisted assessment, reported in aggregate rather than per person — is technically trivial and institutionally radical, because it makes visible a cost that currently falls entirely on the individual and the patient. Speculative The obvious objection is that it converts a tool into a surveillance instrument, which is why it belongs downstream of the cognitive-liberty question rather than upstream of it.

10 · Ethical & societal considerations

Frontier The distributive question here has an unusually clean answer, and it favours augmentation. If the robust positive results really are variance reduction — the lowest-skilled agents, the lowest-rated tutors, the least experienced clinicians gaining most — then cognitive tools are levelling technologies in their measured form, which is the opposite of the standard enhancement worry about entrenching advantage. Speculative That holds only while the tools stay cheap and general; twenty dollars per tutor per year is a fact about one system, not a law of the category.

Established The consent question is not hypothetical. These tools are deployed by employers and schools, and the person whose cognition is being restructured is frequently not the person who chose the tool. Frontier Where an aid produces measurable decay in unassisted competence, requiring its use is a decision about somebody else’s capability taken without measurement and usually without notice. That is the substance of Cognitive Liberty as it arrives in ordinary workplaces rather than through neurotechnology.

Established Responsibility gaps are the concrete harm. The automation-bias structure — a human nominally accountable for catching a failure they have been trained out of the ability to catch — is an accountability arrangement that cannot be met, and it is now standard in medicine, transport and increasingly in professional work. Frontier Assigning blame to the operator inside that arrangement is a category error the aviation investigations named forty years ago and that institutions keep reproducing.

Speculative The enhancement debate proper is fifty years ahead of its evidence, and this brief regards that as its most important ethical observation rather than a complaint. Handwave Arguments about whether a duty to improve exists, or whether the aspiration itself corrodes something, are being conducted over interventions whose measured effects on healthy adults are null, tiny, or unmeasured. Frontier The ethics that would actually bind today concerns silent failure, deskilling and compelled use — and almost nobody is writing it.

11 · Civilizational implications

Speculative Every durable civilisational capability humans have is offloaded, and the alarm about the current round has the same shape as every previous one. Writing, numerals, positional notation, the index, double-entry bookkeeping, the library catalogue and the citation are all external cognitive infrastructure, and each decayed the capacity it replaced while expanding what the composite system could do. Frontier The evidence in this brief is consistent with the current round being the same kind of event, with one candidate disanalogy that keeps recurring: a notebook does not confabulate, and it has no interests in what it tells you.

Speculative The exotic risk is not that individuals get worse but that they get similar. A population’s error-correction capacity depends on people being wrong in different ways; if everyone offloads to the same substrate, that decorrelation degrades, and the failure would be invisible at the individual level because each person is performing normally. Frontier The only evidence pointing at this is output homogenisation in a single field experiment, which measured advertisements rather than reasoning. Handwave Nobody measures reasoning-error decorrelation at population scale, so we would not know if it were happening, and there is no instrument in development that would tell us.

Speculative The optimistic version deserves equal weight. If augmentation is variance reduction, its civilisational effect is compression of the competence distribution — more people at the standard, fewer catastrophic failures at the bottom, and no change at the frontier. Frontier That is a smaller and more plausible story than either the enhancement literature or its critics tell, and it is the one the measurements support. Speculative It also implies something uncomfortable: the returns to augmentation fall as a population gets better, so this is a technology whose value declines with its own success.

12 · Timelines

These horizons track measurement practice rather than capability, because in this subject measurement is what binds.

  • 10 yr: Frontier Publication-bias-adjusted re-analyses of the training and stimulation corpora are published and several surviving positive effects disappear. Dual-instrument reporting becomes normal in education research and starts appearing in workplace-tool studies. Withdrawal effects are measured longitudinally for the first time, in education before medicine because the cohorts are cheaper. Speculative The first randomised comparison of a silently failing against a signalling aid is run; a note tool that schedules retrieval by default reaches a mass market; and at least one professional body issues guidance on unassisted-skill retention, most plausibly in endoscopy or radiology, where the measurement already exists.
  • 25 yr: Speculative Decrement reporting is a condition of publication and of public procurement, and augmentation tools are rated on what they degrade as well as on what they improve. Skill-retention auditing exists in at least one regulated profession. Frontier Whether any intervention has cleanly cleared near transfer in healthy adults is settled either way — a result the field has owed since the 1960s and could deliver in a decade if it chose to. Speculative Search and note tools are by then evaluated the way a drug is, with a stated indication, a stated population and a stated harm profile.
  • 50 yr: Speculative Either a durable, transferring, broad-spectrum gain in healthy adults has been demonstrated and replicated with no offsetting decrement, or the field has a formal statement of why the inverted-U forbids it. Handwave If the second, cognitive augmentation is retired as a research programme and reclassified as interface design, which would be an unusually honest ending for a sixty-year effort.
  • 100 / 250+ yr: Handwave The unit of assessment stops being the person. Composite human-plus-tool systems are certified, audited and insured as units; individual unassisted competence is measured only as a reserve capability, the way manual flying is now; and the question this brief asks — whether the tool made the human better — reads as a category error from a period when the boundary seemed obvious.

13 · Technology tree & dependencies

  • Depends on Cognitive Liberty, and specifically the question of whether a person may be required to use a cognitive aid, permitted to refuse one, or assessed on unassisted performance. Almost every applied augmentation described here is now deployed by an institution rather than chosen by an individual, so the legal answer arrives before the design question does: a retention audit, a decrement disclosure and a right to work unassisted are all instruments that require a settled account of what an employer may do to an employee's cognition. Two further prerequisites are evidential rather than typed. The complementarity boundary — whether the human or the machine is the better component on a given task — is produced by Human-AI Integration and applied here rather than re-derived, and the pharmacological and stimulation nulls come from Cognitive Enhancement.
  • Enables Any programme that needs an honest baseline for what a tool does to the person using it. Future Education Systems inherits the test-alignment discipline directly, since an effect size without its instrument is not a number. AI-Driven Productivity inherits the self-report gap and the variance-reduction pattern, both of which change what a deployment business case should claim. Collective Intelligence inherits the decorrelation question, which is the only route by which individually harmless augmentation could produce a population-level failure. And the interface specification that would follow from a confirmed silent-failure result — calibrated uncertainty, visible degradation, a bounded alert budget — is a prerequisite for deploying decision support in any safety-critical setting at all.
  • Adjacent Human factors and aviation automation research, which holds the failure cases and the measurement apparatus this field lacks, including the finding that manual skill decays with recency of practice rather than with accumulated hours. Information retrieval, which is the most widely used cognitive augmentation in history and has never been evaluated as one with a control group. Personal knowledge management software, optimised almost entirely for capture rather than for the retrieval practice that carries the evidence. The restorative neurotechnology programmes, which supply the subject's headline capability numbers and none of its healthy-adult evidence, and whose results are routinely borrowed to argue for enhancement across a boundary at which the signal, the surgical risk-benefit calculus and the regulatory framework all change. And meta-scientific method — selection models, PET-PEESE, bias-adjusted meta-analysis — which is adjacent to every claim in this brief and is currently doing more to change what the field believes than any new experiment has.

14 · Common misconceptions & speculative claims

Established “Google is destroying our memory — there is a famous Science paper.” The paper is real, and it failed to replicate in 2018 and again in 2020, replicating only conditionally when participants have prior experience that saved information stays accessible. Established The 2024 meta-analysis across thirty-five studies reports changes in how people process and remember information — not damage. Frontier The defensible statement is that people preferentially remember where information sits rather than what it says, which is what a functioning filing system is for.

Established “AI use is measurably rotting our brains — MIT proved it with EEG.” 54 participants in the first three sessions, 18 in the critical fourth; EEG connectivity rather than learning; a preprint; and authors describing potential costs rather than established harm. Frontier A real and well-designed first study, nowhere near the claim made on its behalf.

Established “Brain training works — it is backed by neuroscientists.” A hundred and thirty scientists signed a statement saying so in 2014 and seventy signed the opposite; the evidence favoured the seventy. Established The canonical review found task gains, weak transfer to related tasks and almost no transfer to everyday performance, and the FTC found the leading vendor’s marketing deceptive. Frontier The task gains were real: brain training failed not because nothing happened but because what happened never left the task.

Established “Tutoring produces two standard deviations, so a software tutor will too.” The paper is titled a problem, argues that tutoring is unaffordable at scale, rests on two doctoral dissertations, and was three-armed — conventional instruction, mastery learning, tutoring — so the middle arm the retelling drops carries roughly half the effect. Established Intelligent tutoring systems measure 0.73 on local tests and 0.13 on standardised ones. Frontier Intensive human tutoring under randomisation meta-analyses at 0.288 standard deviations in the peer-reviewed version, having been 0.37 in the working paper, whose title changed from The Impressive Effects to The Promise; both abstracts were read through a metadata registry rather than at the journals. Established Stated carefully: the target that has driven fifty years of educational-technology investment stands in a ratio of roughly fifteen to one against what the resulting technology measures on tests it did not shape, and roughly seven to one against what the intervention it describes measures under randomisation. Full treatment in Future Education Systems.

Established “AI made developers 55.8% faster.” It did, on a from-scratch task with no existing codebase. On mature repositories the same class of tool made sixteen experienced developers 19% slower, while those developers believed they had been 20% faster. Frontier Quoting the first alone is the standard error in the genre, and the self-report gap is the more transferable finding.

Established “The technology is the active ingredient.” PLATO, the largest courseware library ever built, was externally evaluated as essentially equal to an average human teacher, at roughly 300,000 dollars per delivery hour. Established One Laptop Per Child shipped over three million machines; the randomised evaluation across 531 Peruvian primary schools found no significant effects on performance, completion or university enrolment, and the Uruguayan evaluation found no impact on reading or maths with only 4.1% of laptops used all or most days. Frontier In each case the pedagogy was separable from the device and the device was not the part that worked. Speculative The current round should be read against that record rather than against its own marketing.

Established “Personalisation is the mechanism, and learning styles are the evidence.” Learning styles has no adequate evidence base; studies using the necessary crossover design were virtually absent, and among those with proper methodology all but one were negative. One review catalogued 71 models. Established The instructive part is a 2017 survey in which 90% of academics agreed the theory has basic conceptual flaws while 58% agreed students learn better in their preferred style and 33% had used it that year. Frontier Demand for personalisation is not evidence-driven, which is the single most useful fact for assessing any claim that a tool works because it personalises.

Frontier “Checklists cut deaths by a third.” This brief does not know, and says so: the surgical-checklist trials and the population-scale evaluations that complicated them fell outside what its research pack retrieved, so no effect size appears here. Speculative What it will assert is structural: a checklist meets every design criterion this brief arrives at — no hidden state, no silent failure mode, a transferable procedure rather than an answer, no cost. Frontier And if the effects prove smaller than the famous figure, the deflator will almost certainly be the one operating everywhere else here: a result measured on the outcome an intervention was designed around, re-measured on one it was not.

Speculative “Augmentation compounds.” The strong version — tools make people better at making tools, so gains multiply — is the founding intuition of the tradition and belongs to Intelligence Amplification. Frontier The applied record looks more like a plateau with a moving baseline: each tool produces a step change on the task it was built for, the step does not transfer, and the next tool starts from the new throughput rather than from a more capable person. Handwave That does not refute compounding — nobody has measured a person’s capability across two decades of tool adoption against a matched control — but the burden sits with the claim.

Handwave “A genuinely augmented human would be recognisably different.” The honest exotic position is that we do not know what such a person could do, because the field has never specified the target well enough for it to fail. Speculative A serious specification would be a pre-registered battery spanning at least four uncorrelated domains plus one ecological outcome that is not a lab task, a durability window of six months past cessation, and an equivalence bound on every domain not targeted. Handwave Nobody has met that specification with any intervention of any kind, ever — not a drug, not a current, not a training regime, not a piece of software. Frontier Until somebody does, the sentence to hold onto is the unglamorous one: the best-evidenced cognitive augmentation available to a healthy adult is a study schedule, a well-designed piece of paper, and a tool that tells you when it is unsure.