1 · Concept overview
A digital twin is usually sold as a live model of a physical thing: a simulation that receives data from the machine it represents, stays in step with it, and sends something back — a setpoint, a maintenance order, a warning. The consensus technical definition makes that coupling constitutive rather than optional. Most things called digital twins in manufacturing do not have it. They are three-dimensional visualisations, historian dashboards, or offline simulations refreshed by hand, and the word has expanded to cover all of them.
That definitional spread is not a semantic complaint; it is why the value question cannot be answered from the literature. When an operator reports that a twin paid for itself, the reader cannot tell whether the payback came from a coupled model that changed a control decision or from finally putting the plant’s data in one place — a benefit real enough, and one that owes nothing to simulation. This brief separates the two and asks what evidence exists for each.
Three briefs on this map hold adjacent ground. Smart Cities holds the strongest precedent: two decades of urban instrumentation at enormous cost with no evaluated outcome attributable to the sensing. Autonomous Supply Chains holds the finding that the measured gains in logistics automation are overwhelmingly simulated rather than observed. Automated Construction Systems holds the case where the machines work and the institutions do not. This brief owns the factory-and-infrastructure question none of them takes: what a twin has to be to change a physical decision, what it costs to keep one true, and how anyone would know it worked.
The answer this brief lands is that twins have delivered real, measured value in a narrow and identifiable class of settings — continuous processes with good first-principles models, high-value rotating assets, and semiconductor-style production with dense in-line metrology — and that in discrete manufacturing the dominant deployed artefact is a dashboard with a rendering engine. The binding constraint is not simulation speed. It is data integration and the absence of any standard that says how accurate a twin has to be before a decision may rest on it.
Established A note on sourcing. This brief was commissioned in September 2026 from the Institute’s research base. Reading-list entries without links are cited from the bibliographic record rather than re-fetched, and claims are dated no later than early 2026 unless carried by a linked source.
2 · Current scientific position
Established The consensus definition requires a live feedback loop, and it is stricter than industry usage. The United States National Academies’ 2024 consensus study on digital twins defines one as a model that is dynamically updated with data from its physical twin — bidirectional coupling is part of the definition, not a maturity level. Established On that definition a rendered three-dimensional model of a plant, however detailed, is not a digital twin, and neither is a simulation calibrated once at commissioning.
Established The same study names verification, validation and uncertainty quantification as the open problem, and says the reporting standards do not exist. Its findings require that verification, validation and uncertainty quantification be continual rather than one-off, and state plainly that there is a lack of standards in reporting them and a lack of consideration of confidence in modelling outputs, with continual monitoring required to establish trust. Established It also declines to claim that twins substitute for physical testing, framing the roles as complementary. Frontier That is the most authoritative statement available and it is a statement about an absence.
Established The manufacturing standard that does exist standardises architecture, not correctness. ISO 23247, the digital twin framework for manufacturing, published in four parts in 2021, supplies an overview and principles, a reference architecture, a way to describe manufacturing elements digitally, and an information-exchange specification. Established It tells an integrator how to structure a twin and how its components should talk. Frontier It does not set a required fidelity, an accuracy threshold, a calibration interval or an acceptance test, so two conformant twins can differ arbitrarily in whether they are right. Frontier Conformance to ISO 23247 is therefore not evidence that a twin predicts anything, and it is routinely cited as though it were.
Established Standards that do grade model credibility exist in other sectors and are not used here. The American space agency’s standard for models and simulations carries an explicit credibility assessment scale — levels for verification, validation, input pedigree, uncertainty characterisation and the history of the model’s use — and requires the credibility level to be reported with the result. Established The mechanical engineering profession’s verification-and-validation series does something similar for computational solid and fluid mechanics. Frontier Neither is normal practice in factory twins, and the gap between “this model conforms to an architecture” and “this model has a stated credibility level” is the single largest institutional gap in the subject.
Established Where twins demonstrably work, the technology usually predates the word. Model-based real-time optimisation and advanced process control in refineries and chemical plants — a first-principles or empirical plant model running against live measurements and writing setpoints — has been in production since the 1980s and is measured in yield and energy per tonne. Established Per-asset health models for aero engines and turbines, driven by fleet telemetry and used to schedule maintenance, are similarly old and similarly measured. Established Semiconductor virtual metrology and run-to-run control, where a model predicts a measurement that is expensive to take and adjusts the next run, is the cleanest industrial case of a coupled model changing a decision. Frontier All three share the same three preconditions: a process with a usable model, dense instrumentation already installed for other reasons, and a decision the model output maps onto directly.
Frontier Where twins have not demonstrably worked is plant-wide discrete manufacturing. A factory that assembles varied products with human labour, changeovers and supplier variability has no tractable first-principles model, and the twin becomes a discrete-event simulation whose parameters are fitted to historical throughput. Frontier Such models reproduce the past well and extrapolate poorly, which is exactly the regime where confident wrongness is cheap to produce. Speculative This is where the published evidence thins to vendor case studies.
Frontier The measured gains in the adjacent automation literature are overwhelmingly simulated. The warehouse and logistics optimisation literature reports large improvements — computation-time reductions above ninety per cent and rack-movement reductions of tens of per cent in one goods-to-person study, and up to 59.8% fewer robotic tasks in another — and they are simulation results, in models, with the denominators chosen by the authors. Autonomous Supply Chains makes this its central point about the field. Frontier A digital twin is a simulation; a benefit demonstrated inside a simulation of a simulation is not an operational measurement, and the distinction is collapsed in most of the promotional literature.
Established The instrumentation precedent from cities is discouraging and directly relevant. Smart Cities records the largest urban instrumentation programme in the world — a hundred cities, some seventeen and a half billion dollars — producing no evaluated city-level outcome attributable to sensing, and an independent assessment of a decade of that programme describing gains as unbalanced and inconsistent. Frontier The mechanism that produced that result — instrumentation installed without a decision it feeds, and metrics that make sophistication legible while leaving outcomes unmeasured — is available in a factory exactly as it is in a city.
Frontier The productivity statistics cannot settle the question either. Measured manufacturing productivity growth in the United States has been shown to be biased upward by the treatment of offshored inputs in the price indexes, with the bias concentrated in the industries that offshored most. Established That is a result about measurement, not about factories. Frontier It matters here because the natural test of a plant technology — did sector productivity move — runs through a statistic whose construction is contested, so any claim that twins lifted manufacturing productivity needs its denominator stated before it is worth arguing about.
Established The most-cited concrete twin successes in construction and infrastructure are measurement automation. A vision model scoring bricklaying progress against a building information model, feeding a live twin, is the clearest published instance: what is automated is the survey, not the work. Frontier That is a genuine and useful capability, and it is a narrower claim than the one the word carries.
3 · Frontier questions
Frontier How accurate does a twin have to be? Nobody has proposed a general answer, and the question is decision-relative: a model good enough to trigger a bearing inspection is not good enough to authorise skipping one. Speculative A usable formulation ties required fidelity to the cost asymmetry of the decision the twin feeds, which would make fidelity a contract term rather than a marketing adjective.
Frontier Can model-form error be bounded in a running plant? Parameter uncertainty is tractable; the dominant error is usually that the model has the wrong structure — an unmodelled wear mechanism, an unrepresented operator behaviour. Frontier Continual validation, which the consensus study requires, has no accepted procedure when the plant never stops.
Frontier Does a twin degrade gracefully or silently? Sensors drift, instruments are replaced with different ones, a line is rebalanced and nobody updates the model. Speculative The failure mode that matters is the twin that stays confident while becoming wrong, and no deployed system this brief could identify publishes a staleness metric alongside its outputs.
Frontier Is the write-back path worth its risk? The definition’s feedback loop is also a control path into operational technology. Frontier The cyber-security literature has no settled account of how to expose one safely, and the honest position is that most deployed twins avoid the question by not writing back — which also means they are not twins.
Speculative Can twins substitute for physical qualification? The consensus study declines to claim substitution. Frontier Certification by analysis is nonetheless the direction regulators in aerospace and medical devices are moving, and whether a credibility-graded model can replace a test article is the highest-value open question in the subject.
4 · Technological bottlenecks
Established Data integration is the binding constraint and it is unglamorous. A brownfield plant runs controllers of several generations speaking several protocols, with tags named by whoever commissioned the line, no common time base, and gaps whenever a network segment was down. Established Building a twin means first building a semantic model of the plant’s own data, and that work dominates project cost in every honest account.
Frontier Time synchronisation is a specific and underrated failure. A coupled model needs to know the order of events to milliseconds; historians frequently store values with timestamps applied at collection rather than at measurement. Speculative A twin fitted to mis-sequenced data will learn a causal structure that is not there.
Established Sensor drift breaks the coupling quietly. Calibration intervals are set for process control, not for model fidelity, and a slowly drifting instrument produces a slowly wrong twin that continues to report confidence. Frontier No standard requires a twin to declare the calibration status of its inputs.
Frontier There is no accepted acceptance test. Software is delivered against a functional specification; a twin’s real specification is predictive, and integrators are rarely paid against a prediction. Speculative Until a contract can say what the model must forecast and to what tolerance, the market cannot price accuracy.
Established Skills are scarcer than software. The scarce person is the one who understands both the process and the model well enough to notice when they diverge, and that person is the same person the plant already relies on for everything else. Frontier Deployments stall on that individual’s calendar more often than on computation.
5 · Research dependencies
Established The twin depends on the interoperability stack, and the stack is plural. Machine-tool and process telemetry, asset description and information exchange each have credible standards with partial and competing adoption, and a plant typically has several in use at once. Frontier ISO 23247 assumes such interfaces rather than supplying them, so the integration problem is delegated to whichever stack the site already has.
Established It depends on instrumentation that was installed for another purpose. Every successful case in the record inherited dense measurement from process control, safety or metrology. Frontier Where that inheritance is absent the twin requires a sensing retrofit whose cost is rarely counted inside the twin’s business case.
Frontier It depends on model credibility practice imported from elsewhere. The credibility scales and verification-and-validation procedures that would make a twin’s output auditable exist in aerospace and computational mechanics. Speculative Adopting them is an institutional choice, not a research problem, and the fact that it has not happened is the most informative observation in this section.
Established It depends on operational-technology security. A bidirectional twin widens the boundary of the control system to include a data platform, and the state’s own guidance on software security — the memory-safe roadmaps published by the American and allied cyber agencies — applies to that platform as much as to anything else. Memory-Safe Computing Transition holds that argument; the point here is that a twin inherits it.
Frontier It depends on an evaluation culture that infrastructure does not have. The megaproject literature shows that forecasts in built-environment projects are systematically inaccurate and that inaccuracy persists across decades because nobody is scored. Frontier A twin is a forecasting instrument installed in exactly that culture.
6 · Required experiments
Frontier The decisive experiment is a stepped-wedge deployment with pre-registered outcomes, and no operator has published one. Take a firm with several comparable lines or sites, randomise the order in which the twin is switched on, register the primary outcome before the first switch — unplanned downtime hours, first-pass yield, energy per unit — and publish the result whichever way it falls. Established The design is standard in health services research and needs no new technology. Frontier It is the result that would most change this brief’s assessment, because every operational claim in the field currently rests on before-and-after comparisons at single sites where the twin was installed alongside other changes.
Frontier Second: score the twin as a forecaster. Require the model to log dated, quantified predictions — this bearing fails within the next 400 operating hours, this batch lands within tolerance — before the outcome is known, and report calibration over a year. Established The instrumentation for this already exists in every deployment; what is missing is the decision to keep the log and publish it. Speculative The expected finding is that twins are well calibrated on the failure modes present in their training history and badly calibrated on everything else.
Frontier Third: the drift experiment. Freeze a validated twin, then track divergence between model and plant over twelve to twenty-four months with no re-calibration, and publish the decay curve. Speculative This single measurement would convert twin maintenance from a vague overhead into a budgeted interval, and nobody has published one for a production system.
Speculative Fourth: the substitution trial. For a component qualified by physical test, run the qualification in parallel on a credibility-graded model and compare the two decisions across a batch of designs. Frontier That is the experiment that would tell regulators whether certification by analysis is safe in a given domain, and Additive Manufacturing Qualification describes the same missing evidence from the parts side.
Speculative Fifth: the adversarial test. Have a red team attempt to induce a wrong but plausible twin state through sensor manipulation, and measure whether operators detect it. Frontier The result bounds how much authority a twin can safely be given, and it is cheap compared with the deployments it would govern.
7 · Engineering requirements
Established The simulation is the easy part. Physics solvers, discrete-event engines, reduced-order surrogates and the compute to run them in real time are mature and commercially available. Established No deployment in the published record is blocked by solver performance.
Established The hard engineering is the data plane. Edge collection with accurate timestamps, buffering through network outages, a semantic layer that maps tags to plant entities, and versioning so that a model can be replayed against the data it actually saw are the components that decide whether a twin is auditable. Frontier Replay capability in particular is what separates a twin that can be investigated after an incident from one that cannot.
Frontier The write-back path needs to be engineered as a safety function. If a model can change a setpoint, the model is part of the control system and belongs inside the same hazard analysis, with rate limits, envelopes and an interlock that a human can exercise without argument. Speculative Treating the twin as an information-technology system that happens to touch the plant is the architecture most likely to produce the first serious incident.
Frontier Surrogate models introduce a second validation problem. Real-time operation usually requires replacing the physics with a learned approximation, which is accurate where training data was dense and unconstrained elsewhere. Established The approximation error is additional to the model-form error of the physics it approximates, and the two are rarely reported together.
8 · Adjacent technologies
Established The urban precedent is the one to read first. Smart Cities is the same argument at a different scale: instrumentation bought at enormous cost, an evaluation record that does not exist, and a metrics culture that rewards sophistication over outcome. Frontier Every mechanism it identifies is available inside a factory fence.
Established The logistics case supplies the evidence-quality warning. Autonomous Supply Chains finds that the field’s headline improvements are simulated rather than observed, which is the precise hazard for a technology whose product is a simulation.
Frontier Construction and infrastructure supply the deployed instances. Automated Construction Systems holds the pattern where the machines work and the institutions do not, and Robotics in Infrastructure holds the progress-monitoring twins that are the clearest working examples in the built environment.
Frontier Qualification and security are the two disciplines a twin must borrow from. Additive Manufacturing Qualification holds the model-versus-test evidence question, and Memory-Safe Computing Transition holds the software-assurance argument that a control-adjacent platform inherits.
9 · Institutional requirements
Established The standards body has standardised the wrong half. An architecture standard makes systems interoperable; a credibility standard makes their outputs admissible. Frontier Manufacturing has the first and not the second, and the second is what a regulator, an insurer or an auditor would actually need.
Frontier Insurance is the most likely forcing function. An insurer that priced machinery breakdown cover differently for a plant with a validated predictive model would create demand for validation overnight, because the discount would have to be earned against evidence. Speculative No such product is visible in the public record, and its appearance would be a better indicator of real maturity than any vendor announcement.
Frontier Procurement currently buys software, not accuracy. Contracts specify functionality, integration scope and support, and almost never a predictive tolerance with a penalty attached. Established A market cannot reward fidelity it does not purchase.
Frontier The evaluation obligation should sit with the buyer, not the vendor. The operator holds the counterfactual — the other lines, the other sites, the years before — and is the only party able to run the trial that would settle the question. Speculative A public-sector operator of infrastructure could make publication of such an evaluation a condition of any twin procurement at essentially zero cost.
Established Regulated sectors are already moving without the evidence. Certification by analysis is an explicit direction of travel in aerospace and in medical devices, where verification-and-validation practice for computational models is furthest advanced. Frontier Manufacturing twins are being described in the same register without the same practice behind them.
10 · Ethical & societal considerations
Frontier The twin is also a worker-monitoring system, and that is rarely in the business case. A model that tracks cycle times, operator interventions and station dwell resolves to individuals in most real factories. Established The instrumentation needed for a useful production twin is the same instrumentation needed for granular performance management, and nothing in the technical standards distinguishes them.
Frontier Automation bias is the operational risk. A confident wrong model in front of a tired operator at three in the morning is a known failure pattern in every automated domain. Speculative Systems that display uncertainty honestly are harder to sell and safer to run, and the commercial pressure runs the wrong way.
Established A high-fidelity twin is a reconnaissance product. It contains process parameters, layouts, capacities and control logic in one place, which is what an attacker or a competitor would otherwise have to infer. Frontier Concentrating that in a cloud-hosted platform is a decision with security consequences that is usually taken as an information-technology convenience.
Speculative Accountability after an incident is unsettled. If a model recommended the setpoint that produced a failure, responsibility is shared between the operator who accepted it, the integrator who fitted the model and the vendor who supplied the solver, and no jurisdiction has tested that division. Frontier Replay capability is the practical prerequisite for ever apportioning it.
11 · Civilizational implications
Frontier The realistic prize is maintenance, not production. Unplanned downtime and over-conservative maintenance intervals are enormous, diffuse costs across industry, and a calibrated predictive model attacks both. Speculative That is a large aggregate gain and an unglamorous one, and it does not require any of the visualisation the field spends its money on.
Speculative The second-order gain is capital efficiency. If a model can certify that an asset has remaining life, the asset is not replaced, and the embodied emissions and capital of the replacement are avoided. Frontier This is the strongest environmental argument for twins and it depends entirely on the credibility problem being solved, because nobody defers a replacement on an uncalibrated forecast.
Frontier A plant-level twin is a step toward the legible factory, which cuts both ways. Full instrumentation makes production auditable — for emissions, for provenance, for safety — and equally makes it remotely controllable and remotely attackable. Speculative Which of those dominates is an institutional question settled long before the technology matures.
Handwave The strong version — the twinned economy — works by assertion. The vision of a coupled model of an entire supply network, updating in real time and optimising globally, assumes solved data sharing between competitors, solved model composition across organisational boundaries, and solved validation of the composite. Speculative Each is unsolved individually; the composite has no published prototype at industrial scale.
12 · Timelines
These horizons track credibility — whether anyone can say how right a twin is — rather than deployment counts, which will rise regardless and say nothing.
- 10 yr: Frontier A credibility-grading practice for industrial models is published and adopted first in regulated sectors, not in general manufacturing. Frontier At least one stepped-wedge or staggered-rollout evaluation of a factory twin appears, probably from a public-sector or academic-industrial partnership rather than a vendor. Speculative Predictive maintenance consolidates as the dominant profitable use and the plant-wide visual twin quietly becomes a presentation layer.
- 25 yr: Speculative Coupled models are ordinary in continuous process industries and in high-value asset fleets, and remain marginal in low-margin discrete assembly, for the same reason advanced process control did: the model has to be worth more than the instrumentation it requires. Frontier Certification by analysis is accepted for defined component classes with graded model credibility and rejected elsewhere, and the boundary is drawn by accident history rather than by theory.
- 50 yr: Speculative The word disappears into the practice, as advanced process control did, and what remains is an expectation that any significant industrial asset ships with a maintained model and a stated confidence. Speculative Insurance and procurement, not standards bodies, are the mechanisms that made it so.
- 100 / 250+ yr: Handwave A continuously coupled model of industrial production as a whole, composable across firms and jurisdictions, is the horizon the rhetoric already occupies. Handwave Nothing in the measured record constrains it, and the obstacles are commercial and legal rather than computational.
13 · Technology tree & dependencies
- Depends on Three briefs hold results this one waits on. Smart Cities holds the evaluation record for large-scale instrumentation and the finding that spending does not produce attributable outcomes without a decision the sensing feeds. Autonomous Supply Chains holds the distinction between simulated and observed gains that this brief applies to twins themselves. Additive Manufacturing Qualification holds the model-versus-test question on which any claim that a twin can replace a physical qualification depends.
- Requires (not on this map) A pre-registered stepped-wedge trial of a factory twin, because every operational claim in the field currently rests on uncontrolled before-and-after comparisons. A credibility grade for industrial models that a regulator or insurer will accept as evidence, which exists in aerospace and does not exist here. Time-synchronised telemetry out of controllers that were never designed to emit it, because a model fitted to mis-sequenced data learns causal structure that is not there. A sensing retrofit for the installed base of older plant, whose cost is almost never counted inside a twin’s business case. A contract form that prices predictive tolerance rather than software licences, because a market cannot reward accuracy it does not buy. And a safety-case regime for the write-back path, because a model that changes a setpoint is part of the control system whether or not anyone has said so.
- Enables A credible twin enables three things the current artefact does not: maintenance scheduled on remaining life rather than on calendar intervals, which is the largest measurable prize in industry; qualification evidence generated by analysis for component classes where the model has been graded, which compresses development cycles; and auditable production, in which emissions, provenance and safety claims can be checked against a recorded model state rather than against a self-report. All three depend on the same missing thing, which is a number expressing how much the model should be trusted.
- Adjacent Automated Construction Systems and Robotics in Infrastructure hold the built-environment instances, including the progress-monitoring twins that are the clearest working examples. Infrastructure Resilience holds the failure-mode argument that a monitored asset is not a resilient one. Memory-Safe Computing Transition holds the software-assurance case for anything with a path into a control system.
14 · Common misconceptions & speculative claims
Established “We have a digital twin of the factory.” On the consensus definition, a twin is dynamically updated with data from its physical counterpart. Frontier A rendered model, a historian dashboard or a simulation refreshed manually at quarter end fails that test, and most deployed artefacts sold under the name are one of those three.
Established “It conforms to ISO 23247, so it is validated.” The standard supplies an architecture, a way of describing manufacturing elements and an information-exchange specification. Established It sets no accuracy requirement and no acceptance test, so conformance says how a twin is built and nothing about whether it is right.
Frontier “Digital twins reduced downtime by a large percentage.” Almost every such figure is a single-site before-and-after in which the twin arrived with new sensors, a new maintenance process and management attention. Established The design that would separate those causes exists and is standard in other fields; nobody in this one has published it.
Frontier “The simulation is the hard part.” Solvers are mature and real-time performance is routinely achievable. Established What consumes projects is building a semantic model of a plant’s own data out of inconsistent tags, mixed protocols and unreliable timestamps.
Frontier “Twins can replace physical testing.” The most authoritative consensus study on the subject declines to claim substitution and frames modelling and physical testing as complementary. Frontier Certification by analysis may well arrive for defined component classes; it will arrive with a credibility grade attached, not because a twin looked realistic.
Frontier “More sensors make a better twin.” Additional measurement without a decision it feeds is the exact pattern that produced two decades of urban instrumentation with no attributable outcome. Speculative The binding question is which decision changes, and adding channels does not answer it.
Established “The twin is an information-technology project.” If it writes back to the plant it is part of the control system and belongs in the hazard analysis. Frontier Treating it otherwise is the governance error most likely to produce a serious industrial incident attributable to a model.
Speculative “Nothing here works, it is all dashboards.” The opposite overcorrection, and also wrong. Established Model-based real-time optimisation in process industries, per-asset health models for engines and turbines, and virtual metrology in semiconductor fabrication are coupled models that demonstrably change decisions and have done so for decades. Frontier The honest reading is that the technology works where the preconditions hold and that the word has been extended far beyond them.