1 · Concept overview

Human-AI integration is not a technology. It is a configuration: a decision about who holds the decision right, whose information enters the decision, what happens when one party declines to act, and what the reversion mode is when the arrangement fails. The same model and the same person can be arranged in a dozen ways, and the arrangements do not perform alike. That is the whole subject, and almost none of the public argument about it is conducted at that level.

The question the arrangement is supposed to answer is complementarity: does the pairing beat the better of its parts? Answering it requires three measurements — the human alone, the machine alone, and the combination. Most published studies of AI assistance measure two. A preregistered meta-analysis that insisted on all three found 106 qualifying experiments in three and a half years of literature, and the pooled answer was that the combination came out below the better component. Against humans alone the same combinations look excellent. Both numbers are true, they answer different questions, and quoting either without the other misleads.

What follows takes that result as the spine: the measured outcomes of AI assistance in software, support work and clinical medicine; the reliance-calibration literature and what explanations actually do to it; the automation-dependence and deskilling record, which now includes a clinical outcome measured before and after routine AI exposure; the handoff literature from aviation, where this field has a body count and fifty years of findings; and, at the far end, the structural argument that for a large class of tasks complementarity is not merely unachieved but unachievable, together with the task properties that would have to hold for it to be achievable at all. Every capability number quoted here inherits the protocol questions set out in intelligence measurement, and this brief does not restate them.

2 · Current scientific position

Established The anchor result: across 106 experiments, human-AI combinations performed significantly worse than the best of humans or AI alone. Vaccaro, Almaatouq and Malone published a preregistered systematic review and meta-analysis in Nature Human Behaviour in 2024 (8: 2293–2303). Established The search covered the ACM Digital Library, Web of Science and the AIS eLibrary from 1 January 2020 to 30 June 2023, and inclusion required an original human-participants experiment reporting all three arms — humans alone, AI alone, and the combination — which yielded 106 studies and 370 effect sizes. Established The pooled effect against the better of the two components was negative: g = −0.23, t = −2.89, two-tailed P = 0.005, 95% CI −0.39 to −0.07. Established Against humans alone, the same combinations were clearly better: g = 0.64, t = 11.87, P = 0.000, 95% CI 0.53 to 0.74. Established Both are results from the same corpus, and the gap between them is the entire subject: adding a machine to a person helps the person and does not, on average, produce a system that beats whichever of the two was already better.

Established The moderator table is the actionable part, and one of its cells is routinely misread. Established By task type, decision tasks come in at g = −0.27 (t = −3.20, P = 0.002, CI −0.44 to −0.10) and creation tasks at g = +0.19 (t = 1.35, P = 0.180, CI −0.09 to 0.48). Established The decision-task deficit is solid; the creation-task advantage is not distinguishable from zero, because its interval spans zero. Established What the paper establishes is that the two task types differ from one another, not that combinations beat the better component on creative work. Frontier That distinction matters commercially, because the creation cell is the one quoted in support of every writing, design and code-generation product on the market. Established The sharper moderator is the relative-baseline split: where humans outperformed the AI alone, the combination gained (g = 0.46, t = 5.06, P = 0.000, CI 0.28 to 0.66); where the AI outperformed humans alone, the combination lost badly (g = −0.54, t = −6.20, P = 0.000, CI −0.71 to −0.37). Established The pairing helps when the human is the stronger component and hurts when the machine is. Frontier That is a moving target rather than a fixed property, because which component is stronger changes with every model release, which implies the measured value of teaming should decay as capability rises.

Established The publication-bias tests run in the direction nobody expects, and this brief has not seen them quoted anywhere. Established For the primary comparison — combination against best-of-either, the negative headline — Egger's regression found no evidence of bias, with a coefficient of −0.67, t = −0.78, two-tailed P = 0.438, 95% CI −2.39 to 1.04, and the rank correlation test agreeing at 0.05, P = 0.121. Established For the secondary comparison — combination against human alone, the positive result — Egger's regression does indicate bias, with a coefficient of 1.96, t = 3.24, two-tailed P = 0.002, 95% CI 0.76 to 3.16. Established The pessimistic finding is the robust one; the optimistic one is inflated. Frontier If that asymmetry generalises, the entire “AI helps professionals” literature is running warm and the “AI alone is better” literature is not, which inverts the usual assumption that syntheses in a hyped field lean toward the hype. Speculative The mechanism would be mundane: a study showing assistance helps is publishable, a study showing the combination underperforms the machine is awkward for everyone who funded it, and the second kind mostly never gets run at all, because most designs omit the AI-alone arm.

Established In professional work the same intervention has produced opposite signs, and the moderator is task context rather than model quality. Established Peng, Kalliamvakou, Cihon and Demirer randomised developers implementing an HTTP server in JavaScript from scratch; the group with an AI pair programmer finished 55.8% faster. Established METR randomised 16 experienced open-source developers across 246 tasks in mature repositories they had worked on for an average of five years, with early-2025 tooling; the measured effect was a 19% increase in completion time. Established Those developers forecast a 24% reduction beforehand and, after finishing, still estimated they had been 20% faster; surveyed economists predicted 39% faster and machine-learning experts 38%. Frontier The gap between the measured effect and the believed effect is the finding, not the slowdown: practitioners here cannot introspect on their own assisted performance, and neither can the experts forecasting for them. Established Sixteen developers is a small sample; 246 tasks is not, and the direction of the self-report error is what generalises. Established Brynjolfsson, Li and Raymond's staggered rollout across 5,172 support agents found a 15% average rise in issues resolved per hour, concentrated among less experienced and lower-skilled workers, with the most experienced seeing small speed gains and small quality declines, and the largest gains on rare problems where human training is thinnest. Established Ju and Aral ran 2,234 participants producing 11,024 advertisements: human-AI pairs produced 50% more output per worker with higher text quality, human-human pairs produced higher image quality, and human-AI outputs were measurably more homogeneous. Frontier Their instrumented mechanisms — 25% more task-oriented messages, 18% fewer interpersonal ones, 17% more delegation to the AI than to a human partner, 62% fewer direct text edits — describe a human role shifting from producer to reviewer. Frontier Noy and Zhang's preregistered experiment with 453 professionals is the most-cited positive result in this literature; its abstract, as it stands in the bibliographic record, reports a 40% reduction in time and an 18% rise in quality, and this brief did not obtain the paper itself.

Established Clinical medicine has the largest deployments and the most careful controls, and its record is a conditional rather than a conclusion. Established PRAIM compared AI-supported double reading against standard double reading in the German population mammography programme: 12 sites, 119 radiologists, women aged 50 to 69, 463,094 screened of whom 260,739 were read with AI support. Established Detection was 6.7 per 1,000 with AI against 5.7 in the control, a relative increase of 17.6% (95% CI +5.7% to +30.8%), while recall fell slightly, 37.4 against 38.3 per 1,000, and the positive predictive value of recall rose from 14.9% to 17.9% and of biopsy from 59.2% to 64.5%. Established Detection up and recall down together is the shape genuine complementarity takes, and this is the strongest existing evidence for it. Frontier It is also observational, framed as a noninferiority study, and compares assisted reading against unassisted reading rather than against the AI alone, so it does not answer the anchor question. Established Tschandl and colleagues found in dermoscopy that good-quality AI support improved accuracy over either party alone and that the least experienced clinicians gained most — and, in the same paper, that faulty AI can mislead the entire spectrum of clinicians, including experts. Established The conditional is the finding: complementarity depends on the quality of the aid, and the downside is not bounded by expertise. Frontier Gaube and colleagues gave 138 radiologists and 127 internal and emergency physicians advice written throughout by human experts, varying only the label attached, and reported that the purported source did not affect performance; this brief read that study at abstract level and could not confirm its better-known secondary finding about differential susceptibility, so treats its own account as incomplete. Established The background rate sets expectations: two 2005 reviews found process improvement in 64% and 68% of decision-support studies against patient-outcome improvement in 13%, a 2014 review found no mortality benefit when decision support was combined with the electronic record, and a 2011 critique concluded that a large gap remains between postulated and demonstrated benefit and that cost-effectiveness has yet to be shown.

Established Appropriate reliance is the mechanism, both failure modes are measured, and explanations make one of them worse. Established Automation bias — favouring automated suggestions and discounting contradictory information even when the contradiction is correct — splits into commission errors, following a directive without weighing other evidence, and omission errors, failing to notice what the automation missed. Established The magnitudes are not subtle: in one breast-cancer study, cancers found in 46% of cases without an aid were found in only 21% when an aid was present and had missed them, a 25-point absolute drop caused by the tool's presence. Established Decision support raised correct answers from 29% to 50% while turning 7% of previously correct answers wrong. Established Every pilot given a false automated instruction to shut down an engine complied, having insisted in interview that they would not; a 2001 study found pilots detected fewer engine malfunctions with an indication system than without one; a 2005 study found that when an air-traffic aid failed late in a simulation, considerably fewer controllers caught the conflict than when working manually. Established Anti-bias training lowers commission errors but not omission errors, and training that injects deliberate errors works better than training that merely warns of them. Established Under-reliance has its own record: Dietvorst, Simmons and Massey documented algorithm aversion, in which people abandon an algorithm after seeing it err even when it outperforms them, and found that allowing even slight modification of its output restores use — perceived control rather than actual control drives adoption. Established Bansal and colleagues then tested the obvious remedy and found it is not one: with an AI of human-comparable accuracy, explanations increased the chance that humans accept the recommendation regardless of its correctness, and did not improve complementary team performance. Frontier Explanations raise compliance, and compliance is the failure mode rather than the fix; this is the most important constraint on explainable AI as applied to teaming, and the most consistently mis-cited result in it. Established The channel itself also saturates: the FDA catalogued 566 deaths from ignored alarms between 2005 and 2008, the Joint Commission documented 80 alarm-related deaths and 13 serious injuries, and in the 2009 Washington Metro collision, track-circuit alerts running at roughly 8,000 a week had, the NTSB concluded, thoroughly desensitised the dispatchers.

Established Automation dependence, deskilling and handoff have real incident data, and the newest of it is clinical. Established Bainbridge stated the structure in 1983: automating most of a task while leaving the operator responsible for the residue causes skill loss through disuse and imposes an exhausting monitoring load at the same time, so operators need more training rather than less, for interventions that almost never come. Established Budzyńn and colleagues turned that into a measured clinical outcome: in a multicentre observational study published in The Lancet Gastroenterology & Hepatology in 2025, adenoma detection rate in standard non-AI colonoscopy fell from 28.4% to 22.4% after endoscopists had routine exposure to an AI polyp-detection system. Established That is deskilling measured in a clinical outcome, not inferred from a proxy. Frontier Those figures reached this brief through a secondary source citing the Lancet paper rather than through the paper itself, and the result is contested: four published correspondence pieces and an authors' reply followed within months, and a 2026 review has already reframed it as a design question. Established Air France 447 lost 228 people on 1 June 2009 when pitot icing produced airspeed inconsistency, the autopilot disconnected and control reverted to alternate law; the BEA's final report of 5 July 2012 identified the crew's lack of practical training in manual handling at high altitude and in speed-indication anomalies as a major contributing factor, found the crew failed to recognise that the aircraft had stalled, and cited poor management of the startle effect. Established Asiana 214 is cleaner still: the NTSB's probable cause names mismanagement of the descent, the pilot flying's unintended deactivation of automatic airspeed control, inadequate monitoring of airspeed and delayed execution of a go-around, and the board further found the autothrottle and autopilot systems inadequately described in the manufacturer's documentation and the airline's training, the pilot's mental model of the automation logic faulty, and the airline's own automation policy emphasising full use of all automation and not encouraging manual flight during line operations. Established That last finding is the cleanest documented case in any industry of an organisational automation policy producing the skill decay that then produced the accident. Established Haslbeck and Hoermann found fleet to be the strongest predictor of manual flying performance, with long-haul crews degraded, which makes recency of practice rather than total experience the active variable. Established And the Uber ATG vehicle that killed Elaine Herzberg in March 2018 had identified an unknown object six seconds out and flagged the need for emergency braking 1.3 seconds before impact; braking under computer control had been disabled to reduce the potential for erratic vehicle behaviour, and the system did not alert the safety driver that it had detected anything. Frontier That is a handoff-design failure rather than an attention failure: the machine saw the hazard, was forbidden to act, and did not say so.

3 · Frontier questions

Frontier The field's central unreconciled fact is that the sign of the assistance effect flips with position in the skill distribution, and nobody has a mechanism. Established Brynjolfsson's lowest-skilled agents improved on both speed and quality while top performers gained slightly in speed and lost slightly in quality; METR's expert developers on their own mature repositories were 19% slower; Tutor CoPilot's lowest-rated tutors produced the largest student gains, at roughly four percentage points of topic mastery overall and nine for the students of the weakest tutors, for about twenty dollars per tutor per year. Frontier The candidate reconciliation — that the sign flips with the ratio of tacit, uncodified context to codified knowledge in the task — is a hypothesis with no direct test. Frontier What does hold across at least five independent studies is variance reduction: every robust positive in this literature lifts the bottom of a distribution and every negative occurs at the top. Speculative If that is the real shape of the technology, the correct framing is floor-raising rather than amplification, and the policy questions change accordingly.

Frontier Homogenisation is the newest and least-discussed finding. Frontier Ju and Aral measured human-AI team outputs as more homogeneous and self-similar than human-human team outputs, in a design with independent human raters and a live field test. Speculative If augmenting everyone with the same system raises individual performance while compressing population diversity, then for problems solved by parallel search the second effect can dominate the first. Handwave No measurement of that trade-off exists at population scale, and the theory that would predict it — standard in collective intelligence — has never been joined to a measured homogenisation result.

Frontier Deskilling has moved from aviation into medicine faster than the professional response. Frontier The colonoscopy result is the first controlled measurement of AI-induced skill decay in a clinical procedure, and the four-way correspondence exchange it triggered is a discipline deciding in public whether to believe it. Frontier Current work on the aviation version continues — a 2025 Applied Ergonomics study of automation level and pilot flying performance, and the natural experiment created by the pandemic's grounding of fleets, now beginning to appear in conference proceedings. Established This brief obtained neither text and states no numbers from either.

Frontier Handoff is where the harm is and where the research is not. Established The anchor corpus is 106 studies of steady-state joint task performance. Established The three best-documented catastrophic failures in the literature — an autopilot disconnect with law reconfiguration, an unintended autothrottle deactivation under a faulty mental model, and a detection the machine was forbidden to act on and did not report — are all transitions. Speculative The claim that research effort in this field is allocated inversely to risk is a claim about a research portfolio and is testable by survey; nobody has run the survey.

Frontier The methodological frontier is preregistration and three-arm reporting. Established Vaccaro's team could use only 106 studies out of a far larger literature precisely because most published human-AI work omits the AI-alone arm. Frontier AI-tutoring trials are now being lodged in the AEA RCT registry rather than run ad hoc, which is a field trying to get ahead of its own publication bias. Frontier A 2026 randomised trial of an LLM Socratic-questioning tutor against faculty-led endodontic instruction reported no statistically significant between-group difference while substantially reducing faculty time; this brief did not obtain the paper and states no figures from it. Speculative Non-inferiority at lower expert cost may be where this literature honestly settles, and it is a different and more defensible claim than superiority.

Frontier And the jagged-frontier concept has entered the field faster than its evidence. Frontier The term originates with a 2023 Harvard Business School working paper on consultant performance and is now used in peer-reviewed work; its headline figures — the participant count and the contrast between performance inside and outside the capability frontier — are very widely repeated. Handwave This brief could not verify the citation at all through any working instrument and therefore prints none of its numbers. Speculative The concept is repeated here as a concept in circulation, which is a different thing from a result, and the distinction is worth making precisely because the concept is a good one.

4 · Technological bottlenecks

Established The chain to a human-AI pairing that reliably beats the best of either component runs through five binding links, and the first two are cheap and almost never attempted. Established Link one is a measurement standard: every human-AI study reports all three arms. Frontier That is a governance action rather than an experiment, it costs a reviewer's insistence, and nothing downstream is measurable without it.

Frontier Link two is the oracle bound, and it is the cheapest decisive computation in the field. Speculative For a given domain, collect paired human and model predictions on identical held-out instances, build the joint error matrix, and compute what a perfect selector that always chose the correct party would achieve. Frontier If the oracle router does not beat the better single arm, the domain is closed and no amount of interface work can open it. Frontier Almost nobody computes this, and it should be the first thing computed before any complementarity claim is made.

Speculative Link three binds: a learnable router that approaches the oracle bound without outcome leakage. Speculative Every complementarity claim currently in circulation is a claim about the oracle bound, not about a deployable system. Frontier If the features that predict human advantage are exactly the features the model already conditions on, no learnable router can extract the gap, and human-AI complementarity is a fact about counterfactuals rather than about systems.

Speculative Link four is the one most likely to kill the programme: show the router's advantage is not just calibration repair. Frontier Much apparent complementarity is the human catching instances the model was already uncertain about. Speculative A well-calibrated model with an abstention option captures that value without a human in the loop. Speculative If the routed system cannot beat model-with-abstention at matched coverage, the correct architecture is selective automation rather than teaming, and the human's remaining function is governance.

Speculative Link five decides whether any of this is stable, and it costs one re-run. Frontier Every published complementarity result is a single-generation snapshot. Speculative Repeat the routing evaluation with the automated baseline upgraded one full model generation and see whether the margin holds. Handwave Nobody has published a two-generation series in any domain, so the field has no evidence at all on whether teaming is an equilibrium or a waypoint. Established Links six and seven — a handoff protocol, and a practice regime that arrests deskilling — are expensive, and they are where the bodies are.

5 · Research dependencies

Established This brief waits on a measurement result rather than a capability result. Established Every claim above that a combination did or did not beat a component depends on the AI-alone arm being measured under a stated protocol, and the protocol questions — which evaluation split, how many attempts, at what cost, who ran it, whether the items were public before the system was built — are set out in intelligence measurement. Frontier A complementarity finding computed against a contaminated or under-specified AI baseline is not a finding about teaming; it is a finding about the baseline. Speculative This is the one dependency that could invalidate the anchor rather than merely weaken it, and it cuts both ways: an inflated machine baseline makes teaming look worse than it is.

Established The second dependency is corpus access, and it is already satisfied. Established The 106 studies and 370 effect sizes are published and specified, which makes an extended moderator analysis a desk exercise rather than a new data collection. Frontier Extending the task taxonomy from two classes to eight or more, with pooled effects per class, is the highest-value reanalysis available anywhere in this subject and requires no new participants, no new funding and no new instruments.

Frontier The third is a literature this brief could not reach. Established The bulk of the anchor corpus comes from human-computer interaction proceedings, and the research instruments available for this brief were skewed toward encyclopedia articles, preprint abstracts, one publisher's journals and registry lookups by identifier. Frontier Human-factors journals, education journals and HCI proceedings are therefore under-sampled here, and a reader should treat this brief's coverage of the interaction-design literature as thinner than its coverage of the outcome literature.

6 · Required experiments

Speculative Compute the oracle-router bound in five structurally distinct domains. Frontier Paired human and model predictions on a common held-out set, the joint error matrix, and the accuracy a perfect selector would achieve. Speculative The output is a per-domain verdict on whether complementarity is even available, and it is roughly a week of work per domain on data most deploying organisations already hold.

Speculative Run the abstention control before running anything else. Frontier Take any published complementarity result and re-run it against a well-calibrated model that abstains at matched coverage instead of one that always answers. Speculative This is the experiment most likely to dissolve the field's positive results, which is exactly the reason to run it early rather than late.

Speculative Publish a two-generation series. Frontier Same task, same routing policy retrained, automated baseline upgraded one generation, margin reported as a function of baseline capability across at least three points. Handwave No such series exists anywhere, and its absence is why the waypoint question is currently unanswerable rather than merely unanswered.

Speculative Build a handoff protocol and test it against a never-automated control. Frontier The endpoint is measured human performance in the first 60 seconds after an unannounced automation failure, against performance in steady manual operation. Established Air France 447, Asiana 214 and the Uber ATG collision are all this measurement failing in the field, and Bainbridge's constraint applies: the operator needs more training, for an event that almost never happens.

Speculative Add a mandated-practice arm to the deskilling design. Frontier The colonoscopy study run prospectively, with unassisted detection rate as the endpoint at 0, 12 and 24 months and one cohort required to perform a quota of unassisted procedures. Speculative The ugly possibility to price in advance is that the practice regime costs more expert time than the assistance saves.

Speculative Re-measure the centaur claim. Frontier A modern engine alone against a competent player with the same engine and unlimited consultation, under controlled conditions, over enough games to matter. Established This comparison appears never to have been run, and the claim it would settle carries more rhetorical load than any other in the field.

7 · Engineering requirements

Frontier The engineering requirements follow from the failure modes, and almost none of them are about model quality. Frontier A system intended to be part of a team should expose calibrated uncertainty rather than reasons, because the one thing explanations are known to do is raise acceptance without raising discrimination. Speculative Calibration displays, cost-of-error framing, and forcing functions that capture the human's independent judgement before revealing the model's are the three interventions with a plausible mechanism and thin evidence behind them.

Speculative Abstention is an architecture, not a feature. Speculative A model that declines on the instances it is least sure of, routing only those to a person, is the configuration that captures most of the value attributed to teaming, and it is the baseline every teaming architecture should be measured against. Frontier Building it requires calibration good enough that the abstention threshold means something, which is a measurement problem before it is an engineering one.

Established Handoff must be designed as a first-class interface, including the case where the machine detects and does not act. Established The Uber ATG system flagged the need for emergency braking 1.3 seconds before impact, had braking under computer control disabled, and did not tell the driver. Frontier A system forbidden to act that does not announce what it has seen has been engineered into the worst available configuration. Speculative The requirement is that any suppression of autonomous action generate a proportionate human-directed signal, and that the signal budget be managed against the saturation data — 566 alarm-related deaths catalogued over four years, and 8,000 alerts a week desensitising a control room, are what an unmanaged signal budget produces.

Speculative Finally, systems must log what routing will require. Frontier Training a router needs paired human and model judgements on the same instances with outcomes attached, and most deployed assistance tools record only the assisted outcome. Speculative Instrumenting for the counterfactual is cheap at build time and impossible to retrofit.

8 · Adjacent technologies

Established This brief sits between a tradition and a measurement problem. Established Intelligence amplification is the sixty-year programme whose central empirical promise the anchor meta-analysis tests, and the two briefs carry the same result read in opposite registers — one as a verdict on a research tradition, one as a design constraint on deployed systems. Frontier Human-machine symbiosis holds the founding framing, including the founding text's own expectation that the arrangement would be temporary. Frontier Distributed cognition supplies the theory under which a person and a tool are one cognitive system rather than two parties to a transaction, which is the frame in which deskilling stops looking like loss and starts looking like transfer.

Established The applied ends are three. Frontier Human cognitive augmentation carries the individual-level enhancement evidence and the offloading literature, where the deskilling argument takes its general form. Frontier Future education systems holds the tutoring evidence, including both the clearest positive teaming results in the whole cluster and the retraction and publication-bias corrections that have cut several of them down. Frontier AI-driven productivity and future labour markets inherit the skill-compression finding directly: if assistance lifts the bottom of a distribution and mildly degrades the top, the labour-market consequence is not the one usually forecast.

Frontier Two further adjacencies matter at the exotic end. Speculative Collective intelligence is where homogenisation stops being a curiosity and becomes a civilizational question. Speculative Brain-computer interfaces are the bandwidth answer to the channel-capacity hypothesis: if the binding constraint on teaming is how much machine output a person can evaluate per unit time, widening that channel is the only structural fix, and the evidence base there is entirely restorative and does not yet reach enhancement in healthy people.

9 · Institutional requirements

Established The cheapest institutional intervention in this entire subject is a reporting standard. Established A meta-analysis that required all three arms found 106 qualifying studies in three and a half years; the literature it drew from is many times larger. Frontier A CONSORT-style extension requiring human-alone, AI-alone and combination arms in every human-AI study, adopted by the venues that publish this work, would convert an unmeasurable literature into a measurable one at the cost of a checklist. Speculative A reasonable target is 80% of new studies reporting all three arms within three years of adoption.

Established Organisational automation policy is a documented causal variable and it is under institutional control. Established The NTSB found that the airline's policy in the Asiana 214 case emphasised full use of all automation and did not encourage manual flight during line operations, and that the pilot's faulty mental model of the automation logic led to the deactivation that began the accident sequence. Frontier If policy is the mediator, then mandated manual-practice regimes should arrest skill decay — and that experiment is already running in every airline that has changed its policy, with data that are collected and not published. Speculative The same instrument is available in medicine: a quota of unassisted procedures, recorded, with unassisted detection rate as the reported endpoint.

Frontier Professional bodies are the deciding institution for deskilling, and they are currently deciding by correspondence. Frontier The response to the first controlled measurement of clinical AI deskilling was four published letters and an authors' reply within months, which is the right kind of argument to be having and not a substitute for a prospective trial. Speculative The institutional question none of them has answered is whether certification and revalidation should be conducted unassisted — testing the practitioner rather than the practitioner-plus-tool — and no professional body has yet had to decide that in public.

Speculative Liability may settle this before measurement does. Frontier Uber's decision to disable emergency braking under computer control, in order to reduce the potential for erratic vehicle behaviour, is a liability-shaped choice that removed the machine's ability to act and preceded a death. Speculative If deployment configurations track liability regimes more closely than they track performance evidence, the configuration that wins will be the one that allocates responsibility acceptably, whatever the meta-analysis says. Frontier That is a testable claim: look for cross-jurisdictional variation in configuration and see whether it tracks law or evidence.

10 · Ethical & societal considerations

Frontier The strongest case for keeping a human in the loop is not an accuracy case, and pretending otherwise sets the arrangement up to fail on its stated metric. Established Human oversight requirements exist to locate responsibility, satisfy due-process norms and preserve contestability, and those are real goods. Speculative They are not accuracy goods, and a measured decision-task effect of g = −0.27 against the better component means a system defended on accuracy grounds will be embarrassed by its own evaluation while succeeding at what it is actually for. Speculative Saying plainly that the human is there for legitimacy would be more honest, and would change what gets measured.

Established Deskilling transfers cost from the institution that adopts the tool to third parties who did not choose it. Established A patient whose endoscopist has been through a period of routine AI exposure and is now working without the tool bears a six-point drop in detection rate that they did not consent to and cannot observe. Frontier The same structure holds for passengers of an airline whose automation policy discourages manual flight. Speculative What follows is a disclosure requirement about configuration rather than outcome: which decisions were made with assistance, and what the unassisted competence of the responsible professional currently is.

Frontier The human-in-the-loop requirement can itself be the harm. Established A system detected a pedestrian, was forbidden by design to brake, did not alert the human it was relying on, and a person died. Speculative This does not license removing humans from loops; it means the claim that oversight makes systems safer has to be argued configuration by configuration rather than assumed from the presence of a person, and that a duty to keep a human responsible carries a duty to keep that human informed and able to act in the time available.

Speculative Finally, the practice regime has a distributional cost nobody has priced. Frontier If arresting deskilling requires practitioners to work unassisted for a quota of cases, then somebody is receiving the unassisted care or flying behind the unassisted crew. Handwave There is no published framework for who that should be or how the burden should be allocated, and a real institution is going to be asked the question soon.

11 · Civilizational implications

Speculative The civilizational question is not whether individuals get smarter but what happens to the variance. Frontier Every robust positive result in this literature lifts the bottom of a skill distribution, and the measured homogenisation of human-AI team outputs suggests the same intervention compresses the top of a diversity distribution. Speculative For problems solved by parallel search across many differently-wrong approaches — which describes most of science and most of markets — a loss of variance can outweigh a gain in mean, and nobody has measured the trade-off at population scale.

Speculative The second asset at risk is a reserve of unassisted competence. Established Manual flying skill decays with recency rather than with total experience, and a clinical detection rate fell six points in the time it took departments to adopt a tool. Speculative A society that has automated a capability retains the ability to perform it only for as long as it deliberately pays to practise it, and the ratchet turns one way unless something is spent against it. Frontier That is a maintenance cost with no natural owner, which is usually how such costs come to be unpaid.

Handwave And the waypoint hypothesis, if true, dissolves the subject. Speculative If every stable human-AI configuration is a temporary artifact of a capability gap, then integration research is the study of a moving boundary rather than of an architecture, and effort spent designing teaming interfaces is effort spent on arrangements that dissolve from below as capability rises. Handwave The counter-position — that some task classes require a human because they require a human, for accountability, preference elicitation or legitimacy — is coherent, is probably right, and is currently defended with no measurement at all.

12 · Timelines

These horizons track the measurement, not the products. What is being forecast is when the complementarity question could be answered rather than argued.

  • 10 yr: Frontier Three-arm reporting becomes normal in at least the venues that produced the anchor corpus, and an extended moderator analysis over eight or more task classes is published from the effect sizes that already exist. Speculative At least one domain has its oracle-router bound computed and published, and at least one two-generation complementarity series exists. Frontier The deskilling question is settled prospectively in one clinical specialty, with or without a practice regime that works, and the answer decides whether professional revalidation goes unassisted.
  • 25 yr: Speculative Either a learnable router that approaches the oracle bound and survives the abstention control exists in at least one domain, or it has been shown not to, and the field's architecture converges on selective automation with human oversight retained on governance grounds rather than accuracy grounds. Speculative Handoff protocols are certified rather than designed ad hoc, on the model of aviation type-rating requirements, and the 60-second post-handoff endpoint is a regulatory number rather than a research one.
  • 50 yr: Speculative The waypoint question has an answer, because a multi-generation series long enough to show a trend exists. Handwave If the margin shrinks monotonically, teaming architectures are historical and the residual human role is entirely institutional. Speculative If it holds in some task class, that class has been characterised and the properties that make complementarity achievable are known rather than guessed — which is the result this whole subject has spent two decades assuming it already had.
  • 100 / 250+ yr: Handwave Either the channel-capacity constraint has been broken by a bandwidth technology, in which case the human contribution to a joint system is limited by something other than how fast a person can read and check, or it has not, in which case the equilibrium division of labour is set by a biological constant and this subject has a fixed ceiling that was visible from the start and that nobody measured.

13 · Technology tree & dependencies

  • Depends on This brief waits on one result another brief produces: a measurement of machine-alone performance that survives its own protocol questions. Every complementarity claim here is a difference between three numbers, and the AI-alone number is the one most often missing, most often produced by the party that built the system, and most often computed on items that may have been in training data. Until that arm is trustworthy the sign of the combination effect is an artifact of the baseline, and it can be wrong in either direction: an inflated machine baseline makes teaming look worse than it is, a contaminated one makes it look better.
  • Enables The variance-reduction finding and the reliance-calibration constraint feed into every brief that assumes assistance improves professional output. The measured record here is the reason productivity and labour-market forecasts built on the greenfield coding result rather than the mature-repository result are forecasting the wrong effect; the reason education and clinical deployment arguments have to name their instrument before quoting an effect size; and the reason a design that surfaces reasons rather than calibrated uncertainty should be expected to raise compliance rather than accuracy.
  • Adjacent Shares its central result with intelligence amplification, which is the same evidence read as the verdict on a research tradition; its founding framing with human-machine symbiosis; its offloading and deskilling arguments with human cognitive augmentation and distributed cognition; its strongest positive results with future education systems; its skill-compression finding with AI-driven productivity and future labour markets; its diversity argument with collective intelligence; its bandwidth hypothesis with brain-computer interfaces; and its oversight-versus-accuracy problem with AI governance, where the human-in-the-loop requirement is written into rules the accuracy evidence does not support.

14 · Common misconceptions & speculative claims

Established “The meta-analysis shows human-AI teams beat humans alone, so complementarity works.” It shows both results at once: g = +0.64 against humans alone and g = −0.23 against the better of the two. Established Beating the weaker component is not complementarity; it is the machine doing the work with the human in the way. Established And the +0.64 figure is the one the authors' own Egger test flags as publication-biased, coefficient 1.96 at P = 0.002, while the negative headline passes the same test at P = 0.438. Frontier A press release or product page that reports the positive number alone is reporting the contaminated half of a two-part result.

Established “The meta-analysis shows AI helps with creative tasks.” The creation-task effect is g = 0.19 with a 95% interval of −0.09 to 0.48 and P = 0.180, which is not distinguishable from zero. Established What is established is the difference between creation and decision tasks, not the creation effect itself. Frontier The abstract's phrase about significantly greater gains on content-creation tasks is a statement about the contrast between task types, and it is being read constantly as a statement about the creation cell.

Established “AI coding assistants make developers 55% faster.” On a greenfield HTTP-server task with no existing codebase, yes: 55.8%. Established On mature repositories with experienced maintainers, the randomised measurement was a 19% increase in completion time, while the developers believed they were 20% faster. Established Both numbers are real, and the moderator between them is task context rather than model quality. Frontier Quoting the first without the second is the commonest error in writing about AI productivity, and the self-report gap is a result in its own right that generalises well beyond software.

Frontier “An amateur with a laptop beat grandmasters with supercomputers, so a weak human with a machine and a better process beats a strong machine.” Established The tournament series happened: PAL/CSS Freestyle, run on the Playchess server by Computer-Schach und Spiele and sponsored by the PAL Group, with prizes totalling 132,000 euros between 2005 and 2008, and an amateur expert in chess software winning against the general expectation that grandmasters would prevail. Frontier The retelling has not been checked. Established The winning pair's names are spelled differently in different accounts and this brief could not resolve which is right; the ratings usually attached to them appear in no source this brief was able to read and are therefore not printed here. Established A bibliographic search for empirical work on centaur or freestyle chess performance, run in September 2026, returned zero relevant studies — every hit was metaphorical use in another domain. Frontier The most-cited exemplar of human-AI complementarity rests on one tournament series that ended in 2008, has produced no peer-reviewed analysis in nearly two decades, and was never a human-plus-engine against engine-alone comparison under controlled conditions in the first place.

Frontier “Centaur chess proves human-AI teams beat AI.” Established Taken entirely at face value, it shows that in one tournament format a team of humans running multiple engines with a good workflow beat other humans running engines. Established The engine context has moved: an SSDF rating of 3361 for one program in 2016, a 2900 performance rating achieved on a smartphone in 2009, FIDE no longer accepting human-computer results in its rating lists, and a grandmaster's 2016 assessment that the computers are just much too good. Established Tyler Cowen said in 2013 that the centaur advantage was already nearly gone and unlikely to persist much longer; Kasparov in 2017 and James Bridle in 2018 assert the contrary, as assertions in trade books rather than as measurements. Speculative The narrow reading — that the advantage existed only in a window where engines were strong enough to compute and weak enough to need strategic correction, and that the window closed around 2010 — fits every dated source and has never been tested.

Established “Explainable AI will fix human-AI teaming.” The best-known controlled study found that explanations increased the chance humans accept the AI's recommendation regardless of its correctness, and did not produce complementary team performance. Frontier Explanations increased compliance, which is the failure mode rather than the fix, and the paper is regularly cited as though it showed the opposite. Speculative The untested alternative is a display that surfaces calibrated uncertainty and cost of error instead of reasons.

Frontier “AI in radiology has been proven to improve outcomes.” Established The strongest result available — 463,094 women, nationwide, a 17.6% relative detection increase with recall slightly lower — is observational and was framed as a noninferiority study. Established The strongest experimental result comes with the finding that faulty AI can mislead the entire spectrum of clinicians, including experts, attached to it in the same paper. Established And the background rate for clinical decision support is 64% of studies showing process improvement against 13% showing patient-outcome improvement. Frontier A detection improvement in a screening programme is real and valuable; it is not the same claim as improved outcomes, and the distance between them is where most of this literature lives.

Frontier “Keeping a human in the loop makes automated systems safer.” Established The Uber ATG vehicle's system flagged the need for emergency braking 1.3 seconds before impact, had that capability disabled to reduce erratic behaviour, and did not tell the human it was relying on. Established The human-in-the-loop requirement was the proximate mechanism of the death. Speculative This does not mean humans should be removed from loops; it means the safety claim needs arguing for each configuration rather than assuming from the presence of a person.

Established “Deskilling is a theoretical worry.” Adenoma detection in unassisted colonoscopy fell from 28.4% to 22.4% after routine exposure to an AI detection system, measured across multiple centres. Frontier That number reached this brief through a secondary source citing the Lancet paper rather than through the paper itself, and it is contested by four published correspondence pieces; it is nonetheless a clinical outcome measured before and after, not an inference from a proxy task. Established The BEA identified lack of practical manual-handling training as a major contributing factor in an accident that killed 228 people, and the NTSB found an airline's automation policy discouraged the manual flight whose absence it then cited.

Established “Practitioners can tell whether AI is helping them.” Developers forecast a 24% speedup, estimated a 20% speedup after the fact, and were measured 19% slower. Established Pilots insisted in interviews that they would not comply with a false automated instruction to shut down an engine, and then all of them complied. Frontier Self-report is not evidence in this domain, in either direction, and practitioner confidence is uncorrelated with the measurement in every case where both have been collected.

Handwave “Augmentation is obviously better than automation.” Established The largest test of the question finds combinations averaging below the better component, and the word doing the work in that sentence is “obviously.” Speculative The harshest available version of the criticism is that augmentation is an unfalsifiable term: any deployed system containing a human can be described as augmentation, so the word functions to make automation politically legible rather than to name an architecture. Frontier Note that the anchor corpus is drawn entirely from studies whose authors described them as human-AI collaboration, and its average effect against the better component is negative — a literature calling substitution-with-friction by a friendlier name. Speculative Producing an operational definition on observable features — who holds the decision right, whose information enters, what the reversion mode is — would settle this, and is an unglamorous task nobody has taken on.