1 · Concept overview

Future legal systems means the entry of computation into adjudication, legal practice and law-making: algorithmic risk assessment in bail and sentencing, online dispute resolution, machine-readable legislation, self-executing contracts, generative models in legal research and drafting, and the proposition that law becomes predictable enough to dissolve the need for litigation. The subject is usually organised by technology. This brief is organised by evidence, because the evidence is distributed very unevenly across those six things and organising by technology hides that.

The split that forces the structure. Algorithmic risk assessment has a field record running to more than a million cases across three American states, and it is largely null. Online dispute resolution has one deployment anywhere that publishes enough to be assessed — British Columbia's Civil Resolution Tribunal — and what it publishes is a record of deterioration under load. Generative AI in filings has an unusually dense global record, entirely on the failure side, running to 2,035 court findings across more than fifty jurisdictions as of 9 September 2026. Rules as code has eight years of pilots and no published outcome evaluation. Smart contracts have a settled legal answer and no measured adoption. And the legal singularity has no outturn evidence at all. Presenting those at a uniform confidence would misdescribe the field, so the four-flag level does most of the structural work below.

What the brief concludes, stated at the top because it is unfashionable. Nothing in the documented record shows computation improving the quality of adjudication anywhere. The measurable changes are of three other kinds: a collapse in the cost of producing plausible legal text, which courts are absorbing; an enlargement of the intake queue at institutions whose adjudicative capacity is set by staffing budgets; and a set of ordinary doctrinal and procedural responses — certification duties, practice directions, costs orders — that are working roughly as procedural rules usually work. The strongest single comparison available, drawn from one tribunal's own annual report and set out in the position section, has the law outperforming the software by a factor of five on the same platform in the same year, because a legislature wrote a better statute.

The scope boundary with siblings. This brief owns adjudication and the institutions of dispute resolution: courts, tribunals, the rules of procedure, the professional duties of those who appear, and the instruments used to decide cases. AI-Assisted Governance owns the same instruments inside the administrative state, where the decision-maker is an official rather than a judge. AI Governance owns the regulation of the models themselves. Distributed Governance owns machine-readable rules as a mode of collective organisation; smart contracts appear here only as a question about what a court will do with code, which is the question the Law Commission of England and Wales actually answered.

Canada is at the frontier of two of these and is used accordingly. The Civil Resolution Tribunal is the best-documented online tribunal in the world, and Canadian courts have built the densest layer of judicial AI practice directions anywhere. Neither fact makes the brief Canada-specific; both make Canada the place where the evidence is.

2 · Current scientific position

Established Start with the algebra, because it dissolves the argument most people think they are having about algorithmic risk assessment. Chouldechova's identity is the whole of it: FPR = (p/(1 − p)) × ((1 − PPV)/PPV) × (1 − FNR), where p is the prevalence of reoffending in the group. Four quantities — prevalence, positive predictive value, false-positive rate and false-negative rate — are algebraically bound, so fixing two determines the others. In Broward County the prevalences were 51% for Black defendants and 39% for White defendants, and COMPAS satisfied predictive parity at a positive predictive value of ≈ 0.591 across groups at the high-risk threshold. Given those two inputs, the false-positive rate for Black defendants was not free to be equal; it was determined, and it came out at roughly double the White rate. A separate and stronger result by a disjoint author team establishes the three-way version — calibration within groups, balance for the positive class and balance for the negative class cannot hold together except under perfect prediction or equal base rates. Conflating the two theorems is a citation error and is common.

Established So the most famous dispute in this field was not an empirical disagreement, and no better dataset would have settled it. ProPublica measured error-rate imbalance; the vendor measured predictive parity; both measured correctly and the incompatibility is proved. Two things are usually lost. The first is that ProPublica itself published the impossibility result, under the headline that bias in criminal risk scores is mathematically inevitable, quoting Kleinberg’s plain statement that with unequal base rates “you can't satisfy both definitions of fairness at the same time” — an outlet reporting the theorem that complicates its own story. The second is that the vendor's rebuttal is the vendor auditing its own product, and it is routinely cited in the academic literature as a technical contribution of equal standing.

Established The remedy the mathematics points at is one the law may not permit, and almost nobody quotes it. Chouldechova's own recommendation is group-specific thresholds rather than a uniform cut-off; re-thresholding to equalise error rates raised accuracy for Black defendants from 63% to 69% while leaving White defendants at 59%. That is an explicitly race-conscious decision rule written into a criminal-justice instrument. The technical literature's silence about whether any constitutional or human-rights framework would accept it is conspicuous, and this brief records the gap rather than closing it.

Established The best-known finding against these tools is real and conditional, and the replication is the more important paper. The result that untrained laypeople matched a commercial risk instrument holds under the original conditions. A four-experiment replication with 645 participants and 32,250 responses reproduced it exactly — humans 64%, the instrument 65%, a simple regression 68% — and then removed one condition at a time. Without trial-by-trial accuracy feedback, at a realistic low base rate, humans fell to 60% while the algorithm reached 89%. Given more information, humans got worse and the algorithm improved. The slogan that a risk instrument is no better than random people on the internet is an accurate summary of an experiment that does not resemble a bail hearing.

Established And now the field record, which is a million cases deep and largely null. Kentucky made risk assessment mandatory in bail decisions. Non-financial release jumped 13 percentage points, then declined steadily; by five years later more than half the increase had disappeared. The share released within three days rose by 4 percentage points, and a later tool added a barely perceptible further increment. Failures to appear and pretrial crime rose alongside. There was no effect on racial disparities once regional trends were controlled. The compliance gap is the finding: had judges followed the recommendations, 90% of defendants would have received immediate non-financial release, and the actual figure was 29%.

Established Virginia used risk assessment in felony sentencing and got threshold effects without outcome effects. Judges were measurably influenced at the cut-offs — about six percentage points either way, with sentences 23% shorter below the diversion threshold and 34% longer above the upper one. But there was no robust evidence that the reshuffling reduced recidivism, judges used the scores less over time, and the instrument reduced neither incarceration nor reoffending.

Established Illinois settles what the operative variable is, because Illinois had no algorithm. The Pretrial Fairness Act abolished cash bail without mandating any instrument. Pretrial jail populations fell 14% in urban counties and 25% in rural ones in the first year; by two years later the effect had substantially reversed, with the largest counties exceeding pre-reform levels, while the total number of people under some form of pretrial control rose 17% and rural pretrial supervision rose 138%. Same initial effect, same reversion, same net expansion of control, no model anywhere. Whatever drives these outcomes, it is not the presence of a model.

Established The single most useful comparison in this brief is not about risk assessment at all, and it comes from one tribunal's own annual report. In 2024/25 British Columbia's Civil Resolution Tribunal resolved its new intimate-images claims in 58 days on average, with 69% reaching a final decision. Everything else on the same platform, decided by the same members in the same year, averaged 250 days for small claims, 352 days for strata disputes and 378 days for motor-vehicle injury, and reached a final decision 22% of the time. The intimate-images caseload was also the fastest-growing, up 326% to 222 applications, so the speed is not an artefact of a jurisdiction nobody was using. Nothing technological distinguishes the two. What distinguishes them is that British Columbia designed an expedited statutory cause of action with protection orders attached and gave it to the tribunal, and did not do that for the other heads of jurisdiction. The law outperformed the software by a factor of five on the same software.

Established Meanwhile the same tribunal got two and a half times slower in two years, and says why. Average time to resolution ran 108.4 days in 2022/23, 152.8 days in 2023/24 and 277 days in 2024/25; the median in the most recent year was 217 days. Satisfaction with timeliness fell 67% → 60% → 51%, and the proportion who would recommend the tribunal fell from 82% to 75%. The tribunal's own explanation is not technological. New disputes have grown 66% since 2021/22; staffing “didn't keep pace due to budget constraints”; and the report states in terms that “timeliness remains a challenge due to rising dispute volumes and limited funding, and we unfortunately have a significant backlog.” The binding constraint on the world's leading computational-justice institution is the number of people its treasury will pay to decide things.

Established The one unambiguous global change caused by the technology is a failure mode, and it is densely measured. An independently maintained database recorded 2,035 cases across more than fifty jurisdictions as of 9 September 2026 in which a court explicitly found or implied that a party relied on fabricated AI-generated content. The growth curve is the story: two to three cases a month before 2025, exponential growth from April to July 2025, and roughly five a day by February 2026. Its inclusion criterion is strict — findings, not allegations — and its maintainer describes it as “necessarily an undercount” for exactly that reason.

Established And the distribution inverts how it is usually reported. Self-represented litigants account for 1,171 of those cases against 810 for lawyers, with 31 judges and 15 experts. This is substantially a phenomenon of people without a lawyer filing confidently wrong law, rather than primarily one of professional misconduct. The maintainer adds a qualification that belongs in the same breath: a “substantial minority” of entries are vexatious litigants and reckless counsel for whom fabrication is the visible surface of conduct courts were already sanctioning.

Frontier The tools sold to professionals to prevent this fail at rates their marketing denied. On 202 preregistered queries, Lexis+ AI was 65% accurate and hallucinated 17% of the time; Westlaw AI-Assisted Research was 41% accurate and hallucinated 33%; Ask Practical Law AI was 19% accurate, incomplete 62% of the time and hallucinated 17%. The vendor claims quoted in the study were “100% hallucination-free linked legal citations”, tools that “avoid hallucinations by relying on trusted content”, and a system that does not “make up facts, or ‘hallucinate’”. The authors' conclusion is that “the providers' claims are overstated”. This brief's reading is narrower and harder: retrieval grounding traded fabrication for refusal in one product and produced a tool accurate on fewer than one query in five.

Handwave The legal singularity has no outturn evidence — not weak evidence, none — and its peer-reviewed reception says its methods could not produce any. There is no jurisdiction where legal predictability has been measured before and after the introduction of predictive systems. A law-review assessment finds the thesis conflates logic-based and data-driven approaches and that probabilistic methods deliver only “soft” predictability rather than the hard certainty the argument requires. The thesis's leading proponent is the chief executive of a legal-AI company: an interested party authoring the field's manifesto.

3 · Frontier questions

Established Judicial override is not friction in the deployment; it is the substantive finding, and it has been misread for a decade. In Virginia, judges systematically granted leniency to young defendants despite age being among the most heavily weighted inputs to the instrument. Full compliance with the algorithm would have produced a 15 percentage-point relative increase in incarceration for defendants under 23 and a 53% relative increase in their sentence lengths. The humans are not failing to implement the technology. They are implementing a competing normative commitment — youth as mitigation — that the instrument does not encode and was never asked to. Any account treating override as an adoption problem has the argument backwards.

Frontier The leading appellate authority permits a trade secret to inform a sentence, and that is still the law. Wisconsin's Supreme Court held in State v. Loomis that a proprietary instrument may be considered at sentencing provided it is not determinative, and provided every presentence report carries written warnings: that the methodology is undisclosed, that the scores rest on group rather than individual data, that there is potential disparate impact on minority defendants, that no local validation study exists, and that the tool was designed for corrections rather than sentencing. A decade on, that remains the state of the law. Note what the warnings concede. A court told judges that the instrument in front of them had never been validated locally and was being used outside its design purpose, and then permitted its use.

Frontier The genuinely open question in online dispute resolution is whether the front end diverts people or loses them. The Civil Resolution Tribunal's Solution Explorer — a free guided-pathway diagnostic — was used 35,146 times in 2024/25, and only 25% of those explorations resulted in a claim being filed. The figure is normally reported as triage success: most users are diverted before litigation. It is equally consistent with roughly twenty-six thousand people bouncing off a self-help tool and giving up. Nothing in the report distinguishes the two, and the tribunal is the only body in Canada positioned to find out. This is the single most useful unmeasured quantity in Canadian civil justice.

Frontier Whether ODR reaches the party it exists to reach is open and the early answer is bad. The independent usability evaluation of Utah's small-claims platform reports that 64% of Utah defendants never log in to the ODR platform at all, against a stated baseline that 70% of debt-collection defendants do not respond to their lawsuits by any means. Those are different populations and the comparison is suggestive rather than a measured effect — the arithmetic of setting them side by side is this brief's, the figures are the evaluators'. But a platform designed to raise defendant engagement, whose non-engagement rate sits within a few points of the paper baseline, has not yet demonstrated the thing it was built for. The same evaluation found that nearly every participant would prefer a website to a courthouse, which is a finding about preference and not about use.

Frontier Rules as code has produced better legislative drafting and no encoded law, and the obstacle is constitutional rather than technical. New Zealand's flagship programme delivered a new-parent service, a rates-rebate calculator, and encoded rules used by three life-event services. The movement's own advocacy site names no pilots, no agencies beyond a single ministry, no dates and no quantified outcomes. The intergovernmental primer states there are “no known large-scale approaches currently embedded in and across governments” and offers no impact data. The most serious national assessment titles a part “Code should not be Legislation”, holds that “only the Judiciary makes authoritative interpretations”, warns against legislation “that, by its architecture, permits no interpretation”, and finds “a lack of understanding around the constitutional implications” in the movement's own work. What the pilots actually produced were concept models, decision flow diagrams and rule statements: genuinely useful instruments for clarifying policy intent during drafting, and not law.

Established The smart-contract question is closed and was closed by declining to change anything. The Law Commission of England and Wales advised Government on 25 November 2021 that “the current legal framework in England and Wales is clearly able to facilitate and support the use of smart legal contracts”, requiring only “incremental and principled development of the common law in specific contexts”. The substantive move is the interpretive test. The Commission proposed a “reasonable coder” standard — what a person with knowledge and understanding of code would take a coded term to mean — and expressly rejected asking how a functioning computer would execute the code, because that “would entail necessarily conflating the meaning of a coded term with its performance or output”. Put plainly: a law-reform body was asked whether execution is the agreement, and answered no. Residual work was identified on deeds, where witnessing and attestation sit awkwardly with code, and on private international law, where jurisdiction rules struggle with contracts that operate unilaterally or through autonomous program interaction.

Frontier What is open about smart contracts is not the doctrine but the volume. No source consulted for this brief supplies an adoption figure or a body of decided disputes in England and Wales in the five years since the advice. A practitioner commentary argues the proposed workaround for the conflict-of-laws problem, express choice-of-law clauses, is a patch rather than a reform, and notes the Commission's own concession that “legislation or regulation may be needed in the future”. The sceptical reading is available and this brief cannot verify it either way: the law may have been declared adequate for a phenomenon nobody has measured.

Frontier The frontier that is genuinely moving is doctrinal, and it is moving fast in Canada. Between June 2023 and September 2025, Canadian courts and tribunals built the densest layer of judicial AI practice directions anywhere: Manitoba's Court of King's Bench on 23 June 2023, Yukon on 26 June 2023, Nova Scotia from October 2023, the Federal Court on 7 May 2024, Ontario's O. Reg. 384/24 in force 1 December 2024, Tribunals Ontario in April 2025, the Trademarks Opposition Board in June 2025 after two incidents, amended Federal Court practice guidelines on 20 June 2025 adding a contempt warning, and Federal Court interim principles governing the Court's own AI use on 29 September 2025. Quebec, Newfoundland and Labrador and Alberta issued guidance without mandatory disclosure. This is what a legal system changing actually looks like, and it is entirely made of practice directions and a regulation.

4 · Technological bottlenecks

Established The first bottleneck is adjudicative labour, and the leading deployment says so in its own annual report. The Civil Resolution Tribunal's average time to resolution rose from 108.4 to 277 days in two years while new disputes grew 66% since 2021/22. Its stated cause is that “claim volumes rose but staffing levels didn't keep pace due to budget constraints”, producing what it calls “a significant backlog”. Digitisation lowered the cost of filing and did nothing to the cost of deciding, which is a salaried tribunal member reading a file. Every technology in this brief acts on intake; none of them acts on the binding constraint.

Established The second is that nobody independently evaluates any of this, and the field's own coordinating bodies concede it. The founding statistic of the entire online-dispute-resolution literature — sixty million disputes a year, ninety per cent resolved by software — comes from a single slide presented by the operating company's own dispute-resolution director in 2010, with no outcome data and no verification, and it has been recycled ever since. The National Center for State Courts' resource on measuring ODR success recommends no metrics, publishes no outcome data, and advises courts to define success early and engage outside evaluators later. Utah's ODR pilot “Final Report”, filed under a State Justice Institute grant, contains no outcomes at all: it is pre-pilot planning that describes a launch scheduled for April 2018, lists hoped-for indicators — “a drop in default judgments and an increase in settlement agreements” — and documents development delays caused by IT miscommunication and a “moving target” of design changes.

Established The third is that the independent user research is negligible in volume. The only non-operator study of Civil Resolution Tribunal users this brief could reach surveyed 49 people, in 2019. It found 75% rating Solution Explorer at least somewhat helpful and 25% not, 43% saying the information did not address their problem, 20% feeling anxious entering a legal problem online, 6% with internet-access difficulty, and roughly one-third finding the system difficult and confusing; of 21 participants with recent legal problems, 67% found the tribunal easier than court. Those are useful findings and the sample is forty-nine. That is the independent evidence base on the best-documented online tribunal in the world.

Established The fourth is that the fabrication problem is measured only where court records are open. The database's maintainer is explicit that United States dominance is an artefact of PACER and CourtListener, while European jurisdictions impose “artificial frictions” and anonymisation requirements that produce an international undercount. So the geographic distribution of the 2,035 cases measures the transparency of court records, not the distribution of the behaviour. Any national comparison drawn from it is a comparison of publication regimes.

Established The fifth is that the largest court system in the world does not count its own incidence. Two United States federal district judges issued opinions in July 2025 containing fabricated quotations, citations to non-existent passages, parties who were not parties and misstated case outcomes, disclosed only under congressional inquiry. The judiciary's administrative office confirmed it is aware of incidents anecdotally and “does not systematically track such activity at the national level”. Thirty-one judges appear in the independent database. Neither number is the true one and only one of the two bodies could produce it.

Frontier The sixth is that sanctions are too small to be the mechanism, whatever the commentary says. The Canadian record's most-cited outcomes are personal costs of $100 in one Federal Court matter, a small costs award against a self-represented respondent in another, contempt purged through professional development in a third, and a struck motion record in a fourth. Against 2,035 findings worldwide, the deterrent operating here is professional reputation and judicial displeasure, not money. That may be sufficient. It is not what the phrase “courts are cracking down” implies.

Established What is not a bottleneck, stated plainly. Model capability is not the constraint on anything in this brief. Nothing here waits on a research result; the models are ordinary and the platforms are ordinary software. Every binding constraint is a budget line, a procedural rule, a disclosure regime or a constitutional allocation of interpretive authority.

5 · Research dependencies

Established Nothing in this brief waits on a research result. The models are ordinary and the platforms are ordinary software. That sentence is unusual in a frontier-research corpus and it is the correct one here: no capability described above is blocked by anything a laboratory could produce.

Established What this field waits on is four things, all institutional, and their absence is a choice rather than a constraint. Local validation of any instrument before deployment, whose absence the leading appellate authority requires judges to be warned about rather than remedied. Disclosure of methodology, currently defeated by trade-secret claims that courts have accepted. Published override rates, without which the human-in-the-loop cannot be assessed and where the one indirect reconstruction available gave 29% actual against 90% modelled. And independent outcome evaluation of dispute resolution, which the field's own coordinating body recommends and which has never been performed on any major deployment.

Established It depends, more than anything else, on adjudicative headcount funded from general revenue. The leading online tribunal spends $12.2 million of a $14.6 million budget on salaries against $1.1 million on information systems, and attributes its performance collapse to staffing that did not keep pace with volume. Any projection of computational justice that does not carry a staffing line is projecting the wrong variable.

Established It depends on open court records, which is a prerequisite for knowing anything at all. The fabrication database exists because PACER and CourtListener exist; its international undercount exists because other jurisdictions impose access frictions and anonymisation. Measurement of legal systems is downstream of publication policy, and publication policy is set for unrelated reasons.

Established What depends on it. AI-Assisted Governance inherits every finding here about override, validation and disclosure, because the administrative version of the problem is the same problem with a weaker appeal route. Distributed Governance inherits the smart-contract answer: the “reasonable coder” test is the reason a legal system will not treat execution as agreement, and any scheme premised on the contrary is premised on a proposition a law-reform body has already rejected. And access to justice depends on this brief in the direction nobody wanted — the tools reaching self-represented litigants first are currently damaging their cases.

6 · Required experiments

Established The cheapest decisive experiment is to publish override rates and outcomes by defendant group. Two of the field studies underpinning this brief reconstructed judicial compliance indirectly from administrative data — that is how the 90%-modelled against 29%-actual figure for immediate non-financial release was obtained. No jurisdiction publishes it directly. It is the single number that separates a decision-support tool from a decision, it exists already inside case-management systems, and its absence after a decade of deployment is a choice.

Frontier The second is already half-run and needs only a follow-up survey: find out what happened to the 75%. Solution Explorer was used 35,146 times in one year and produced a filed claim 25% of the time. Whether the other twenty-six thousand users resolved their problem, were correctly diverted, or abandoned a valid claim is unknown, and the tribunal holds the contact details. A single outbound survey would convert the field's most-quoted triage statistic from an ambiguity into evidence, in either direction.

Frontier The third is a natural experiment that has already produced its result and that nobody has written up. The Civil Resolution Tribunal's intimate-images jurisdiction resolved claims in 58 days against 250 to 378 days for every other head, on the same platform, in the same year, with the same staff, while growing 326%. The design difference is statutory. Decomposing that gap — how much comes from the expedited procedure, how much from the availability of protection orders, how much from the simplicity of the legal test — would tell legislatures more about how to make civil justice fast than any further work on platforms.

Frontier The fourth is to independently evaluate one online tribunal against a comparable court. Volume, time to resolution, outcome distribution, default rates and appeal rates all exist in administrative records on both sides. That nobody has done this for any major deployment is a striking gap for a field that presents itself as evidence-led, and the field's own coordinating body recommends exactly this and reports no instance of it.

Established A negative result worth recording as an experiment in its own right. Utah built a small-claims ODR platform to raise defendant engagement and reduce default judgments. The independent usability evaluation found 64% of defendants never log in, against a paper baseline of 70% non-response among debt-collection defendants. Different populations, so not a clean measurement — but the pilot's own “Final Report” named the drop in default judgments as its target indicator and then contained no outcome data. A programme that names its own success criterion and never reports against it has produced a result, and the result should be recorded.

Frontier The fifth is the machine-readable-law comparison that eight years of pilots have not produced: error rates and processing times for a benefit administered from encoded rules, against the same benefit administered conventionally. Rates rebates and parental-entitlement services were encoded; the counterfactual arm was never run, and the movement's own site publishes no outcome figures of any kind.

Established And an experiment the profession is running on itself without a protocol. Commercial legal-research tools are now in general professional use with published accuracy of 65%, 41% and 19% on a preregistered 202-query benchmark. Every filing made with them is a trial. The result is being collected, badly and involuntarily, by the fabrication database — which is why that database is the most valuable evaluative instrument in the subject and why it is maintained by one researcher rather than by any court system.

7 · Engineering requirements

The engineering content of this subject is mostly institutional design, because the software is ordinary. What follows is the record of what has been built, what it cost, and what it produced.

Established The Civil Resolution Tribunal is the only online tribunal in the world that publishes enough to be engineered against, and this is what its 2024/25 year looked like. The numbers below are the operator's own, from an annual report that criticises its own performance.

Measure2022/232023/242024/25
New applications7,8158,806
Disputes closed6,1266,928
Open at year end6,1668,044
Average time to resolution108.4 days152.8 days277 days
Median time to resolution133 days217 days
Satisfied with timeliness67%60%51%
Would recommend82%75%
Cost per dispute$2,016$2,109
Survey response rate~6% (465 of ~8,000)4% (487 of ~11,000)

Established The composition of the caseload and the composition of the outcomes are both instructive. Of 8,806 new applications in 2024/25, small claims accounted for 5,737, motor-vehicle injury 1,338, strata 977, accident benefits 245, accident responsibility 206, intimate images 222 and societies and co-operatives 78. Of the disputes closed, excluding intimate-images claims, 55% were resolved by consent or withdrawn, 22% ended in a final decision, 16% ended in default or non-compliance, 3% the tribunal refused to accept and 4% it refused to resolve. The monthly detail sharpens it. In January 2025 the tribunal closed 632 disputes: 306 withdrawn, 176 by final decision, 95 by default or non-compliance and 568.9% — resolved by agreement. The addition and the percentage are this brief's; the components are the tribunal's. An online tribunal whose modal outcome is withdrawal, and whose negotiated-agreement rate is under one in ten, is not obviously doing what the online-dispute-resolution literature says such bodies do.

Established The cost structure is the least discussed and most important engineering fact here. Total expenses of $14,608,976 against revenue of $634,835, of which $12,164,737 is salaries and benefits and $1,137,366 is information systems and technology. Staffing is 112 full-time equivalents: 18 tribunal members, 25 case managers, 38 information and intake support, 14 decision support, and the rest management, counsel and corporate services. Salaries are roughly eleven times the technology line. A system that is described as a software platform is, financially, a hundred-and-twelve-person office with a website, and its throughput scales with the office.

Established Engineering lessons from the failure side, which is where the dense evidence is. Verify citations mechanically before filing, because the failure mode is an authority that looks correct: in one leading English judgment, five non-existent cases were cited and a neutral citation that did exist belonged to an unrelated case on a different subject. Distinguish the two hallucination types when designing a check, because they need different checks — an incorrect answer is factually wrong or non-responsive and a competent reader catches it, whereas a misgrounded answer supplies a real, clickable citation that does not support the proposition, and that one survives a spot-check by design. Build disclosure into procurement rather than litigating it afterwards, since the trade-secret question was settled in the vendor's favour by a court that then required judges to be warned about the consequences. And instrument the override: a system that cannot report how often it was overruled cannot be evaluated at all.

Established The best procedural design in the record is the one that does not mention the technology. Ontario's O. Reg. 384/24 requires certification that authorities cited in factums are authentic, and it applies to every party regardless of whether AI was used. Every disclosure-based direction, by contrast, must answer whether spell-check, machine translation, voice recognition or a research database's newly added AI feature triggers the duty — and the Federal Court has already had to issue clarifications doing exactly that. A technology-neutral procedural duty outperforms a technology-specific disclosure duty, and does not go stale when the tools change. That is a general lesson about regulating by artefact rather than by harm, and this brief regards it as the most transferable engineering finding in the subject.

8 · Adjacent technologies

AI-Assisted Governance is the nearest neighbour and the relationship is a shared instrument with a different decision-maker. Risk scores, eligibility engines and generative drafting appear in both; what changes is that an administrative decision is reviewable on different grounds, by a different body, usually later and more weakly. Every finding here about override, local validation and undisclosed methodology transfers, and transfers worse, because the administrative version has a thinner appeal route.

AI Governance supplies the frame and does not decide anything in this brief. Model regulation operates on developers and deployers; the operative rules in adjudication are procedural duties on parties and practice directions on judges, which is a different legal instrument reaching different people. The Ontario authenticity-certification regulation is the clearest illustration: the most effective response to a generative-AI harm in a Canadian court did not regulate AI at all.

Distributed Governance takes up machine-readable rules as a mode of organisation. This brief supplies the doctrinal answer that side needs and rarely quotes: England and Wales considered whether a coded term means what a computer does with it, and said no, adopting instead a “reasonable coder” standard that keeps interpretation with the court.

Access-to-justice measurement is adjacent as the source of this field's motivating number and none of its outcome measures. Roughly 36% of people worldwide report a justiciable problem in a two-year window and 49% of those cannot meet the resulting legal need; on the broadest framing, 5.1 billion people face at least one justice issue. Nothing built in this brief has been connected to a measured movement in any of those figures.

Off the map entirely: criminal procedure and the law of evidence, which govern how any of this may be used; algorithmic fairness, which is where the impossibility results were proved and where the group-threshold remedy sits unexamined against constitutional law; and court administration, which produces almost every usable number in this subject and is almost never treated as a research field.

9 · Institutional requirements

Established The institution that has actually changed legal systems in the last three years is the practice direction, and Canada has built the densest layer of them anywhere. Manitoba's Court of King's Bench on 23 June 2023 requiring disclosure of how AI was used; Yukon on 26 June 2023 requiring the tool and its purpose; Nova Scotia's Provincial Court from October 2023 and its Registrar in Bankruptcy from October 2024; the Federal Court on 7 May 2024; Ontario by regulation on 1 December 2024; Tribunals Ontario in April 2025; the Trademarks Opposition Board in June 2025 after two incidents; amended Federal Court guidelines on 20 June 2025; and Federal Court interim principles on the Court's own AI use on 29 September 2025. Quebec, Newfoundland and Labrador and Alberta issued guidance without mandatory disclosure. Nine instruments, eight jurisdictions, twenty-seven months.

Established The Federal Court's design is worth reading closely because it is the most carefully drawn. A declaration must appear “in the first paragraph of the document in question” identifying AI-created content, and it is required only where content was “directly provided by AI” — not where AI merely “suggest[s] changes, provide[s] recommendations, or critique[s]” human-authored work. Three principles govern: caution, requiring “well-recognized and reliable sources” including official court websites, commercial publishers and CanLII; human in the loop, requiring verification before submission; and neutrality, under which a declaration draws no adverse inference and responsibility for accuracy stays with the signing party under the existing Rules. The Court also commits to avoiding automated decision-making in judgments “without first engaging in public consultation” — a court placing a procedural condition on its own future conduct, which is an unusual and under-noticed instrument.

Established Ontario's response is the structurally different one and, on this brief's reading, the better one. O. Reg. 384/24 requires certification that authorities cited in factums are authentic, and it applies to every party regardless of whether AI was used. It is Canada's only legislative response, it needs no definition of “artificial intelligence”, and it does not go stale when the tools change. Disclosure-based directions, by contrast, must keep answering whether translation, voice recognition or a research database's newly added AI feature triggers the duty, and the Federal Court has already issued clarifications doing exactly that.

Established The doctrinal centre of gravity has moved from the tool to the declaration. As the Federal Court's approach has been summarised in Canadian commentary, “the real issue is not the use of generative artificial intelligence but the failure to declare that use”. In Hussein v. Canada the aggravating fact was concealment: counsel used an immigration research tool that fabricated two cases and mis-cited a third, did not declare it, and the Court characterised the non-disclosure as “an attempt to mislead the Court”, adding that AI use “must be declared and as a matter of both practice, good sense and professionalism, its output must be verified by a human”.

Frontier The sanctions, however, are small, and the gap between rhetoric and consequence should be stated. Personal costs of $100 in the Federal Court; a small costs award against a self-represented respondent in British Columbia; contempt purged through professional development in Ontario; a motion record struck from the file as a sanction “necessary to preserve the integrity of the Court's process”. Trade commentary counts 27 Canadian cases tracked, against 22 United States cases in July 2025 alone. Courts are responding consistently and are not, on this record, imposing costs that would change a rational actor's behaviour.

Established The institution that measures this does not sit inside any court system. The global count — 2,035 cases, 50-plus jurisdictions, the pro se and lawyer split, the growth from two or three a month to five a day — is maintained by a single independent researcher using referrals, scrapers, keyword searches and tip-offs from legal editors, and is described by him as “necessarily an undercount” because a court must say so before a case can be counted. Its geographic distribution measures publication policy: the United States dominates because PACER and CourtListener exist, while European anonymisation requirements create “artificial frictions”. The most important dataset in this subject is a by-product of two American court-records systems and one person's time.

Frontier The institution that does not exist is an evaluator. No independent body has evaluated any major online dispute resolution deployment against a comparable court. The National Center for State Courts recommends exactly this and reports no instance of it, publishing process guidance and no outcome data. Utah's pilot filed a “Final Report” containing no results. The only independent user study of the Civil Resolution Tribunal this brief could reach has n = 49 and dates from 2019. A field that has been running for two decades has no evaluative institution, and its coordinating bodies are membership organisations for the operators.

Established And the institution that decides everything is a treasury. The Civil Resolution Tribunal's caseload grew 66% since 2021/22 and its staffing did not, because of budget. Its resolution times followed. Whatever else is true about computation in law, the throughput of the leading computational tribunal in the world is set by a line in a provincial budget, and no software on the market changes that.

10 · Ethical & societal considerations

Established The impossibility theorem is an ethical result disguised as a technical one. Once base rates differ between groups, a choice must be made about which kind of unfairness to accept, and no amount of engineering removes it. That choice is currently made implicitly by vendors and left to courts to warn about, which is the least accountable place it could sit. The sharper point is that the mathematics recommends a remedy — group-specific thresholds, which raised accuracy for Black defendants from 63% to 69% while leaving White defendants at 59% — that is an explicitly race-conscious decision rule in a criminal-justice instrument. Nobody in the technical literature consulted here asks whether a constitutional or human-rights framework would permit it. This brief does not know the answer and records the question as open.

Frontier The trade-secret holding remains the sharpest institutional problem. A defendant may be sentenced with reference to a score whose weightings the defence cannot examine, on the reasoning that the score is not determinative — a distinction the field evidence shows is real, since judges override constantly, but which offers a particular defendant nothing. The accompanying warnings concede that no local validation exists and that the tool was designed for a different purpose, and then permit its use anyway.

Established The fabrication record raises an access-to-justice question nobody planned for. The tools are used most by litigants without representation — 1,171 of 2,035 recorded cases — and what they produce is a filing that looks like law and is not. A technology adopted disproportionately by the least-resourced party, which systematically damages that party's case, is a distributional harm even though every individual use was voluntary. Canadian courts have held self-represented litigants to the same verification standard as counsel, which is defensible in principle and falls on people who cannot buy verification.

Established Evidence quality in this subject is dominated by interested parties, and the brief marks them. The founding ODR statistic is a vendor's slide. The most detailed deployment data in the world comes from an operator reporting on itself — candidly, which is why it is used, but it is self-report. The most-quoted satisfaction figures come from a survey with a 4% response rate, 487 completed from about 11,000 invitations, skewed 71% toward applicants. Any citation of that tribunal's satisfaction figures should carry the response rate in the same sentence. The vendor rebuttal in the COMPAS dispute is the vendor auditing its own product. The legal-singularity thesis is authored by a legal-AI chief executive. Most of the Canadian practice-direction chronology comes from law-firm commentary, whose facts are checkable and whose framing sells advisory services.

Established Public accountability has a specific failure here that is worth naming. The largest court system in the world does not count its own incidence of AI-fabricated material: its administrative office confirmed it is aware of incidents anecdotally and “does not systematically track such activity at the national level”, disclosing two federal judges' contaminated opinions only under congressional inquiry. The count that exists — 31 judges among 2,035 cases — is maintained by one independent researcher. A public institution's incidence data being produced entirely outside it is an accountability finding, not a curiosity.

Frontier And an opportunity-cost question that is rarely posed. The leading online tribunal spends $12.2 million on salaries against $1.1 million on information systems and attributes its performance collapse to staffing. If the constraint on civil justice is adjudicative labour, then money spent on platforms is money not spent on the constraint, and the burden of proof sits with the platform. No evaluation exists anywhere that discharges it.

The brief's own unresolved questions, stated as an obligation. Whether the 75% who use a triage tool and do not file were diverted or lost. Whether any legal system would accept the group-threshold remedy the fairness mathematics recommends. What the true incidence of AI-fabricated material is in jurisdictions that anonymise their judgments. Whether smart legal contracts exist in any measurable volume. And whether the Civil Resolution Tribunal's deterioration reverses when it is funded, which is the cleanest available test of this brief's central claim and which British Columbia alone can run.

11 · Civilizational implications

Established The civilisational finding is that legal systems are not changing through technology, and the best evidence for that comes from inside the flagship technological deployment. On one platform, in one year, with one staff, the Civil Resolution Tribunal resolved intimate-images claims in 58 days and motor-vehicle injury claims in 378. The difference is a statute. Set beside it: Kentucky's mandate produced a four-point effect that halved within five years; Virginia's sentencing tool reduced neither incarceration nor recidivism and was used less over time; Illinois got the same trajectory with no algorithm at all; eight years of machine-readable-law pilots produced eligibility calculators on the movement's own admission; and the smart-contract question was resolved by a law-reform body declining to change anything. Every measurable improvement in this subject traces to a legal-design or resourcing decision, and every measurable deterioration traces to volume meeting fixed staff.

Established What genuinely changed is narrower and faster than the transformation thesis, and it is entirely negative in the documented record. The marginal cost of producing plausible legal text collapsed, and courts are absorbing the consequences: 2,035 findings across more than fifty jurisdictions, from two or three a month to about five a day in fourteen months, borne disproportionately by people without lawyers. That is a real, fast, cross-jurisdictional change in legal systems caused by a technology. It is not the change anybody was arguing about.

Established The reassuring half, which deserves saying because this corpus is not in the business of decline narratives. The response has been fast, ordinary and largely competent. Practice directions in eight Canadian jurisdictions inside twenty-seven months; a technology-neutral certification duty written into Ontario's rules; costs orders, contempt jurisdiction and struck filings applied under existing powers; and a court publishing principles constraining its own use of the technology before anyone required it to. The institutional immune system worked, using instruments that predate the internet. Anyone arguing that legal systems cannot adapt to computation is contradicted by the fastest-moving part of the record.

Speculative The larger civilisational stake is distributional and is being decided by default. Roughly two-thirds of humanity faces at least one justice issue and about half of those with a civil problem cannot meet the need. The technologies in this brief reach that population first and unmediated — a self-represented litigant with a chatbot is the modal user — and the measured result so far is a filing that looks like law and is not. Whether that inverts, and reliable assistance for unrepresented parties becomes the durable change, is the only version of this subject with civilisational weight, and it is currently unevidenced in either direction.

Handwave Projections of computationally settled law describe no jurisdiction. They rest on a claim its own reviewers say the methods cannot establish, advanced by a party who sells the software.

12 · Timelines

Established What already happened, because this timeline usually starts too late. Wisconsin decided State v. Loomis in 2016, and the trade-secret holding has not moved since. ProPublica's COMPAS analysis and the two impossibility theorems all landed in 2016–17, which means the mathematical core of the algorithmic-fairness argument is a decade old and settled. British Columbia's Civil Resolution Tribunal opened in 2016. New Zealand's rules-as-code work began around 2018. The Law Commission of England and Wales closed the smart-contract question on 25 November 2021. Kentucky, Virginia and Illinois had all produced their field results before generative models entered legal practice at all.

Established 2023 to 2026: the only fast-moving thing in the subject. Judicial practice directions arrived from Manitoba and Yukon in June 2023, Nova Scotia from October 2023, the Federal Court of Canada on 7 May 2024, Ontario by regulation on 1 December 2024, Tribunals Ontario in April 2025 and the Trademarks Opposition Board in June 2025; the Federal Court added a contempt warning on 20 June 2025 and published interim principles on its own AI use on 29 September 2025. Over the same window the fabrication count went from two or three cases a month to about five a day, reaching 2,035 by 9 September 2026. Doctrine moved at roughly the speed of the problem, which is not the usual complaint about legal systems.

Frontier Next, and resolvable without anyone building anything: whether the Civil Resolution Tribunal's deterioration reverses. Its timeliness collapse is attributed by its operator to funding and staffing against a 66% volume rise. British Columbia can test the claim directly by funding it. If resolution times fall with headcount, the labour-constraint reading in this brief is confirmed; if they do not, something else is wrong and the field should be told what.

Frontier Ten years: mechanical citation verification becomes routine in filing systems, because the failure mode is cheap to detect and expensive to ignore, and because Ontario has already shown the rule can be written technology-neutrally. Over the same period, risk assessment continues in wide use with override rates still largely unpublished unless a jurisdiction is compelled to publish them, and online tribunals keep growing on volume metrics with no independent outcome evaluation.

Speculative Twenty-five years: the plausible settlement is computation doing intake, triage, scheduling and disclosure, where errors are cheap to reverse, while contested adjudication stays human — because the override evidence shows judges enforcing normative commitments the instruments do not encode, and because the Law Commission has already refused to let execution stand for meaning. Whether machine-readable law ever governs anything beyond eligibility calculation is genuinely open and has not moved in eight years.

Speculative Fifty years: the durable change, if there is one, is in access rather than adjudication. Reliable assistance for self-represented litigants would be a larger shift than anything on the sentencing side. The current record is the opposite of that: 1,171 of 2,035 recorded fabrication cases involve people without lawyers.

Handwave Beyond that, and specifically any date attached to computationally settled law. The claim that law becomes perfectly predictable is a conjecture whose reception in the legal literature is that its own methods cannot deliver it, and whose leading proponent sells the software. This brief supplies no date.

13 · Technology tree & dependencies

  • Depends on Nothing on this map. The models are ordinary, the platforms are ordinary software, and no capability described in this brief is waiting on a result another brief produces. The constraints are budgetary, procedural and constitutional, and they are recorded below.
  • Requires (not on this map) Local validation of an instrument before it is used, the absence of which the leading appellate authority requires judges to be warned about rather than fixed. Methodology disclosure that survives a trade-secret claim, which courts have so far declined to require. Published override rates, without which “the judge decides” is unfalsifiable, and where the one indirect reconstruction gave 29% against a modelled 90%. Independent outcome evaluation of dispute resolution, recommended by the field's own coordinating body and never performed on a major deployment. Sustained funding for adjudicative headcount, which is the constraint the leading online tribunal names in its own annual report and which consumes eleven dollars for every one spent on its technology. And open, machine-readable court records, without which nothing in this subject can be counted at all — the global fabrication database is a by-product of two American publication systems. None is a research result; all six are things the field could have and has mostly chosen not to build.
  • Enables Faster, cheaper handling of disputes where a wrong answer is cheap to reverse — intake, triage, scheduling, disclosure. No typed enabling edge is claimed, and the field's own record does not yet support a stronger claim than that.
  • Adjacent AI-Assisted Governance, which meets the same instruments in administration rather than adjudication; AI Governance, which supplies the regulatory frame for the models themselves; Distributed Governance, where machine-readable rules pose the same authority question from the other direction. Off-map: criminal procedure and the law of evidence; the algorithmic-fairness literature, where the impossibility results live; access-to-justice measurement, which supplies the motivating statistic and no outcome measure; and court administration, which produces nearly all the usable data.

14 · Common misconceptions & speculative claims

“The AI-in-law story is about lawyers cutting corners.” Established It is mostly about people without lawyers. Of 2,035 recorded cases, 1,171 involve self-represented litigants against 810 involving lawyers, with 31 judges and 15 experts. The database's maintainer characterises the typical case as a self-represented litigant or “a lawyer who is surprised that an existing tool suddenly has an AI component”. The distributional consequence is the ethical one: a technology adopted disproportionately by the least-resourced party systematically damages that party's case. One qualification belongs in the same breath and cuts the other way — a “substantial minority” of entries are vexatious litigants and reckless counsel for whom the fabrication is the visible surface of conduct courts were already sanctioning.

“Vendor sanctions trackers are stale copies passed off as proprietary research.” Established This brief said that in its previous version and it was overstated; the correction is recorded here rather than quietly dropped. The tracker examined for this pass cites “roughly 1,490 court decisions worldwide” with “more than 1,000 of them in the United States as of May 2026”, and it credits and links the underlying independent database explicitly and repeatedly, calling it the running worldwide count. It is stale — 1,490 against 2,035 four months later — and its publisher sells legal AI to in-house counsel, so its framing is interested. The accurate criticism is staleness and commercial framing, not misattribution.

“Retrieval-augmented generation solved legal hallucination.” Established On 202 preregistered queries, three leading commercial tools hallucinated between 17% and 33% of the time, against marketing that promised “100% hallucination-free linked legal citations”. And the worse number is not the hallucination rate: Ask Practical Law AI answered 19% of queries accurately and returned incomplete answers 62% of the time. Grounding moved some failures from fabrication to refusal without producing reliability. The specifically dangerous failure is the misgrounded one — a real, clickable citation that does not support the claim made — because it is precisely the failure that survives the verification a careful reader performs.

“Online dispute resolution makes justice fast and cheap.” Established The best-documented instance in the world got two and a half times slower in two years — 108.4 to 152.8 to 277 days average — while satisfaction with timeliness fell 67% to 60% to 51% and the cost per dispute rose to $2,109. Its own explanation is that volume grew 66% while staffing did not, producing a “significant backlog”. Meanwhile 64% of Utah's ODR defendants never log in. The honest description of what online dispute resolution has achieved so far is that it moved intake online and enlarged the queue.

“The Civil Resolution Tribunal resolves small claims in about eight weeks.” Frontier A 57-day median for small claims appears in a justice-innovation NGO's case study and circulates widely. The tribunal's own 2024/25 report gives 250 days average for small claims and a 217-day median across all claim types. The figures are not strictly contradictory — different years, and a type-specific median against an all-type average — and this brief does not resolve the conflict. What it does say is that the number in circulation comes from a better period and a narrower measure, and that quoting it in 2026 misdescribes the institution.

“The COMPAS fight was a disagreement about the definition of fairness.” Established It was an algebraic constraint. FPR = (p/(1 − p))((1 − PPV)/PPV)(1 − FNR). With prevalence at 51% for Black defendants and 39% for White defendants in Broward County, and predictive parity holding at PPV ≈ 0.591, the false-positive rates were not available to be equalised. Describing this as a definitional dispute implies somebody could have chosen better words; nobody could have chosen better data. The part that is almost never reported is the author's own remedy — group-specific thresholds, an explicitly race-conscious decision rule — and the fact that no source consulted here examines whether any legal system would permit it.

“The algorithm is no better than random people on the internet.” Established That is a correct summary of an experiment that does not resemble a bail hearing. The four-experiment replication with 645 participants and 32,250 responses reproduces the result under the original conditions — humans 64%, instrument 65% — and shows it vanishing once you remove trial-by-trial accuracy feedback, use a realistic base rate, or supply more information. In the realistic condition it is 60% against 89%, and additional information made the humans worse.

“The flagship advocacy reversal was an evaluation.” Established It was political and said so. The organisation that had championed pretrial risk assessment withdrew support citing no study, on the stated ground that it “did not have the right people at the table”. It is regularly reported as though a body of evidence had turned. The evidence had already turned, years earlier and in the peer-reviewed literature, and the reversal is a fact about advocacy rather than about instruments.

“The famous online-dispute-resolution statistic is well established.” Established Sixty million disputes a year with ninety per cent resolved by software comes from a single slide presented by the operating company's own dispute-resolution director in 2010, with no outcome data, no independent verification and no update since. It anchors most of the literature.

“Machine-readable law is being rolled out across governments.” Handwave The movement's own advocacy site names no pilots, no agencies beyond a single ministry, no dates and no results. The intergovernmental primer states there are “no known large-scale approaches currently embedded in and across governments” and offers no impact data. The most serious national assessment titles a part “Code should not be Legislation”, holds that “only the Judiciary makes authoritative interpretations”, and finds “a lack of understanding around the constitutional implications” in the movement's own output. What eight years of pilots produced were concept models, decision flow diagrams and rule statements — useful drafting instruments, and not law.

“Smart contracts will need new law.” Established England and Wales decided in November 2021 that they do not, and the reasoning matters more than the conclusion. The Law Commission adopted a “reasonable coder” test — what a person who understands code would take a coded term to mean — and expressly rejected asking how a functioning computer would execute it, because that “would entail necessarily conflating the meaning of a coded term with its performance or output”. Code is law was put to a law-reform body and refused on the merits, without legislation. What remains open is deeds and private international law — and the fact that no source consulted here can show a body of decided smart-contract disputes or any adoption figure at all.

“The justice gap proves we need legal technology.” Frontier The justice gap is real and enormous: 5.1 billion people, roughly two-thirds of humanity, facing at least one justice issue; 1.5 billion with unmet civil, administrative or criminal needs; 36% reporting a justiciable problem within two years, of whom 49% could not meet the need. It is also entirely self-reported household-survey data with no before-and-after component, and no source consulted for this brief connects any computational deployment to a measured movement in it. The gap motivates the field; it does not evaluate it.