1 · Concept overview
Established A robot foundation model is a single neural network trained on large curated corpora of robot demonstrations, usually alongside internet-scale vision and language data, intended to control many robots across many tasks from natural-language instructions. The field calls the dominant family vision-language-action models (VLAs): a pretrained vision-language backbone is fine-tuned to emit motor commands, on the bet that the semantic knowledge in web-scale pretraining will let physical skills generalise the way text generation did. Google DeepMind's RT-2 and RT-X models, the Open X-Embodiment consortium dataset, Stanford's OpenVLA, Physical Intelligence's pi-0 series, and Gemini Robotics are the load-bearing artefacts; Figure, Tesla, Agility Robotics, 1X, Unitree and AgiBot are the companies staking capital on the same bet in humanoid form.
Frontier The research question this brief examines is whether curated demonstrations can become transfer — across embodiments, into recovery from failure, at control-loop latency, under physical-safety constraints, and out into open-world reliability. Each of those five capabilities now has at least one published proof of concept. None of them has a demonstrated, independently replicated transfer at deployment scale. The distance between those two states is the subject of this brief, and the field's own measurements — not its critics' — are what document it.
Established The seed result illustrates the pattern exactly. In Nature Machine Intelligence (vol. 7, pp. 592–601, March 2025), Mon-Williams and colleagues at Edinburgh presented ELLMER: a Kinova Gen3 seven-degree-of-freedom arm driven by GPT-4 with retrieval-augmented generation and closed-loop force and vision feedback, making coffee and decorating plates in a kitchen while people jostle the mug. It is a genuine advance in long-horizon robustness — and it is one arm, one kitchen, one curated knowledge base, with the authors themselves listing the assumptions (accurate object identification, a pre-built affordance map of every utensil) that confine it there.
Frontier Money has moved far ahead of the measurements. Figure sought a $39.5 billion valuation in 2025, roughly fifteen times its February 2024 mark, on deployment claims journalists could not verify; Morgan Stanley projects a $5 trillion humanoid market by 2050. Against that: the largest disclosed real deployment of a learning-based humanoid moved 100,000 totes at one warehouse, and the best-selling humanoid maker's own IPO filing shows about 5,500 units delivered in 2025, mostly to research and demonstration buyers. Reader beware is the correct posture on every number in this field that comes from a party raising money on it.
2 · Current scientific position
Established Cross-embodiment transfer has been measured once at scale, and the result was positive, partial, and direction-dependent. The Open X-Embodiment collaboration (21 institutions, 60 datasets, 22 robot embodiments, over one million trajectories covering 527 skills across 160,266 tasks) trained RT-X models on the pooled data and evaluated them at the contributing labs. RT-1-X improved mean success rate by roughly 50 percent over each lab's original method in small-data domains — the headline transfer result. In data-rich domains the same model underperformed the specialists it replaced: 27 percent versus 40 percent on the Bridge evaluation, 73 percent versus 92 percent on the RT-1 kitchen tasks. Only the 55-billion-parameter RT-2-X recovered the deficit, and its emergent-skill evaluation (75.8 percent versus RT-2's 27.3 percent, roughly a threefold gain) is the strongest single piece of evidence that pooled robot data buys something real.
Established Open replication then found that scale is not the whole story. OpenVLA, a 7-billion-parameter open-source model trained on 970,000 Open X-Embodiment trajectories for 21,500 A100-hours, beat the closed 55B RT-2-X by 16.5 absolute percentage points across 29 WidowX tasks (roughly 71 percent versus 55 percent). It also documented the costs the demo videos omit: about 6 Hz inference on an RTX 4090, single-image input only, incompatible with 50 Hz control, and below 90 percent success even on tested tasks — the authors' own list.
Frontier Open-world generalisation has one strong vendor-reported result. Physical Intelligence's pi-0.5 (vendor) performed multi-minute cleaning tasks — dishes to sink, beds made, laundry into baskets — in homes absent from training, from roughly 400 hours of mobile-manipulation data plus co-training on other embodiments and web data; the company reports out-of-distribution instruction-following and success figures in the 83–94 percent range on its own evaluation slices, and states plainly that the model “often makes mistakes both in terms of its high-level semantic deductions and motor commands.” Google DeepMind's Gemini Robotics 1.5 (vendor), announced 25 September 2025, claims motion transfer in which tasks trained only on the ALOHA 2 platform “just work” on Apptronik's Apollo humanoid and a bimanual Franka. Neither claim has an independent replication.
Established Where independent evaluators have re-measured frontier VLAs, reported numbers did not survive contact. A Macquarie University study (arXiv:2511.11298, November 2025) benchmarked pi-0, OpenVLA-OFT, RDT-1B and the ACT baseline on four bimanual tasks on an ALOHA Mobile platform. In-distribution macro-averages: pi-0 at 72.3 percent, ACT at 48.3 percent, RDT-1B at 13.5 percent, OpenVLA-OFT at 13.0 percent — two of four celebrated models in the low teens. Under unseen objects with randomised placement, pi-0 fell to 58.5 percent and ACT collapsed to 5.5 percent. Three researchers with one robot produced the most informative distribution-shift data in the public record, which says as much about the field's evaluation habits as about the models.
Established The instruction-following that gives VLAs their name is partly illusory. The LIBERO-PRO evaluation (arXiv:2510.03827) found models posting 90-plus percent benchmark success whose outputs did not change when the instruction was corrupted, swapped for a paraphrase about a different object, or replaced with meaningless tokens. A separate 2026 study measured drops of 22 to 52 percentage points from mere paraphrasing of the instruction. Cheap fixes exist — relabelling trajectories with discriminative language lifted one single-task setup from 0 to 90 percent success, and counterfactual relabelling added 27 points without new data — which is encouraging precisely because it shows the models had been passing benchmarks by memorising trajectories, not by grounding language.
Established Evaluation practice itself is the documented weak point. A Toyota Research Institute and Cornell team (Kress-Gazit et al., Robot Learning as an Empirical Science) showed the field's dominant metric — success rate — is routinely published with no trial counts, no initial-condition control, and no stated success criteria; in one of their own examples three researchers disagreed on whether an ice-pouring run had succeeded, and two policies with similar aggregate rates (72 versus 80 percent) behaved entirely differently across initial conditions. Until reporting improves, most published VLA success rates are not comparable with each other, and few are reproducible even in principle.
Established Real deployments exist and are real — and narrow. Agility Robotics (vendor) reported on 20 November 2025 that its Digit fleet had moved over 100,000 totes at GXO's Flowery Branch facility, the first humanoid deployment with a disclosed cumulative-throughput number under a multi-year commercial agreement; the announcement discloses no uptime, mean-time-between-failure, fleet size or intervention rate. Unitree's March 2026 IPO filing — a regulatory document, but still the company's own accounting — reports over 5,500 humanoids delivered in 2025 against output above 6,500, revenue of 1.71 billion yuan (about $250 million, up 335 percent year on year), first profitability, and an average humanoid price collapsing from 593,400 yuan in 2023 to 167,600 yuan (about $25,000); the filing does not say what buyers do with them, and the visible answer — research labs, showcases, entertainment — is not labour.
Established The flagship humanoid programmes have publicly slipped. Tesla targeted 5,000–10,000 Optimus units in 2025; supply-chain reporting relayed by Electrek in July 2025 indicated parts secured for roughly 1,000, deployment limited to moving batteries in Tesla's own workshops at less than half human efficiency, supplier feedback citing joint-motor overheating, low hand load capacity, short transmission life and limited battery, and the programme head, Milan Kovac, departing weeks after promotion. Figure's CEO declined on stage in June 2025 to answer questions about the scope of its BMW deployment after threatening to sue an outlet that reported on it. 1X began taking $20,000 orders (or $499 per month) for its NEO home humanoid in October 2025 with delivery in 2026; in a Wall Street Journal test, every household task was performed by a remote human teleoperator — five minutes to load three dishes — and the CEO's stated rationale was data: “If we don't have your data, we can't make the product better.”
Frontier Summed honestly: transfer has been demonstrated as a training-time effect, not as a deployment property. Pooling demonstrations measurably helps low-data robots (RT-X, and DROID's 22-point in-distribution and 17-point out-of-distribution co-training gains). Nothing yet shows a robot foundation model holding its performance when the developer is not present, the environment is not curated, and the evaluator is not invested — the three conditions that jointly define open-world reliability.
3 · Frontier questions
The live questions are the ones the proofs of concept were designed around rather than through.
Frontier Is cross-embodiment transfer interpolation or generalisation? RT-X's gains came on embodiments inside the training pool, evaluated by the contributing labs. Whether a genuinely held-out morphology — different kinematics, sensor suite, gripper — benefits, and by how much, has no controlled public measurement; Gemini Robotics 1.5's zero-shot motion-transfer claim is exactly this experiment, reported so far only by its vendor.
Frontier Do robot-data scaling laws exist, and what is the exponent? Language models had a measured loss-versus-data curve before the capital arrived. Robotics has nothing comparable: the largest open teleoperation corpora (DROID at 76,000 trajectories and 350 hours; AgiBot World at over one million trajectories from 100 robots) are minute against the text corpora that made LLMs work, and nobody has published a defensible curve saying how much demonstration data buys how much open-world success.
Frontier Can recovery from failure be learned from demonstrations of success? Teleoperators rarely demonstrate error states; curated corpora are success-biased by construction. ELLMER's force-feedback re-planning and pi-0.5's vendor-reported mid-task corrections are early evidence that some recovery is learnable, but no benchmark currently scores a policy on what it does after its first mistake, which is where deployed robots live.
Frontier Is vision-plus-proprioception the right sensor basis at all? Rodney Brooks' September 2025 argument is that it is not: human dexterity rests on roughly 17,000 mechanoreceptors per hand plus muscle-spindle and tendon force sensing, and demonstration video — especially the human-video-only data Tesla and Figure say they will train on — simply does not contain the contact signals the skill consists of. Every prior AI breakthrough, he notes, rode on domain-appropriate input engineering; vision-only robot learning assumes dexterity is the exception.
Frontier What does a success rate mean? The Macquarie replication, LIBERO-PRO's corrupted-instruction probes, and the TRI critique together make evaluation science itself a frontier: the field does not yet agree on trial counts, initial-condition control, or even what counts as success, so the headline numbers that drive investment are not yet measurements in the ordinary scientific sense.
Speculative Whether the humanoid form factor is the right chassis is genuinely open. The argument for it — the built world fits humans — is an assertion about economics, not physics; Brooks predicts the winning general robots will have wheels and two-fingered grippers within fifteen years, and current deployments (tote handling, fixed workcells) do not yet exercise the capabilities that justify legs and five-fingered hands.
4 · Technological bottlenecks
Established Data is the binding constraint, and it is priced in human hours. Every frontier VLA is trained on teleoperated demonstrations collected one robot-hour per robot-hour: DROID took 50 collectors twelve months to gather 350 hours; pi-0.5's mobile-manipulation corpus is about 400 hours; AgiBot World's million-plus trajectories required a dedicated facility running 100 robots (programme's own figures). Commercial collection services (vendor) quote $8–35 per episode depending on contact richness — at which price a trillion-token-equivalent corpus of physical interaction does not exist and cannot be bought.
Established Latency is a second hard constraint. OpenVLA runs at about 6 Hz on a desktop GPU; contact-rich control wants 50–500 Hz. Every current system bridges the gap with action chunking or a fast low-level controller under a slow deliberative layer, which reintroduces exactly the hand-engineered hierarchy end-to-end learning promised to remove, and caps reactive performance at the speed of the slowest necessary layer.
Established Touch sensing at scale does not exist as a component. No commercially available robot hand carries tactile sensing within two orders of magnitude of human mechanoreceptor density, and the fragility of what does exist (AgiBot's array-based visual-tactile fingers are the exception that proves the rule) means most fleet data being collected today contains no contact channel to learn from.
Established Hands and actuators are failing before the learning does. Tesla's 2025 supplier feedback — overheating joint motors, hand load capacity, transmission lifespan, battery life — is a list of mechatronic problems, not machine-learning ones; the reported production pause over hand and forearm design in late 2025 stalled the largest planned humanoid fleet on hardware grounds.
Established There is no certification route for a learned controller. ISO 10218:2025 (published March 2025 after eight years of drafting, folding in the ISO/TS 15066 collaborative provisions) still assumes the machine's safe state is reachable by removing power. A 2026 Siemens feasibility study names the resulting fail-passive gap: a walking biped cannot be de-energised safely, so no existing standard can even express its safe state, let alone certify a neural policy's contribution to it; their demonstrator's worst-case stop response was about 1.1 seconds, and the authors explicitly decline any PL e / SIL 3 claim.
Frontier Evaluation cost quietly rations progress. Statistically meaningful real-robot evaluation takes hundreds of trials per condition per policy; almost nobody pays it. The result is a literature of 10-to-30-trial success rates with no confidence intervals — a bottleneck on knowing, which is upstream of every other bottleneck.
5 · Research dependencies
Established The field free-rides on vision-language model progress. Every VLA inherits its semantic generalisation from a pretrained backbone (PaLI, PaliGemma, Llama-2-plus-DINOv2/SigLIP in OpenVLA's case); gains in VLM spatial reasoning and video understanding transfer downstream at near-zero marginal cost to robotics, and this dependency is the strongest argument the optimists have.
Frontier Simulation fidelity determines whether data economics can be escaped. If contact-rich manipulation can be learned in simulation and transferred, the teleoperation-hour constraint dissolves; sim-to-real for dexterous contact remains unreliable enough that every frontier lab still pays for real demonstrations, which is the market's own estimate of simulator adequacy.
Frontier Tactile sensor manufacture is a materials-and-packaging dependency — robust, cheap, high-density skin at the fingertip — that sits outside machine learning entirely and currently has no volume producer.
Established Standards work is a live dependency with a visible schedule. ISO 25785-1, the first standard drafted specifically for dynamically stable humanoids, was in working-group development through late 2025 (the Barcelona meeting hosted by Novanta); until it lands, every humanoid deployment is negotiated ad hoc against standards written for machines that can safely lose power.
Established Deployment depends on batteries, actuators and thermal management developed for and by adjacent industries — electric vehicles above all — and humanoid cost curves (Unitree's 72 percent average-price fall in two years) are largely inherited from those supply chains rather than produced by robotics itself.
6 · Required experiments
Frontier The decisive test is independent, statistically powered replication: a frontier vision-language-action model, run by evaluators its developer does not employ, in facilities it has never seen, holding its reported success rates within stated confidence intervals. The Macquarie study is a first, small instance — four tasks, one platform, three authors — and two of the four models it tested landed in the low teens in-distribution. The full version needs no new hardware and no new science: released model weights, a handful of labs, pre-registered success criteria, hundreds of trials per condition. Nobody has funded a standing version of it, and until one exists every capability claim in this field is calibrated by parties with a position.
Frontier The cross-embodiment ablation that would settle transfer is equally concrete. Hold out an entire morphology from a pooled corpus, fine-tune on N local demonstrations with and without the pooled pretraining, and publish the curve of transfer benefit against embodiment distance. RT-X measured the with/without contrast only for embodiments inside the pool; the held-out version, at several values of N, is the experiment the phrase “robot foundation model” presumes and does not have.
Established Instruction-perturbation audits should be attached to every reported benchmark. LIBERO-PRO's protocol — corrupt, paraphrase and nullify the instruction, and report the success delta — costs almost nothing and already reclassified several 90-percent models as trajectory memorisers. It is the cheapest decisive test in the field.
Established Reliability has a natural experiment already running. Agility's GXO deployment accumulates exactly the statistic the field lacks — interventions, faults and throughput per robot-hour under commercial pressure — and the 100,000-tote announcement proves the counter exists. Publication of the denominator (robot-hours, intervention rate, MTBF) would convert a marketing milestone into the first field-reliability dataset for a learning-based legged system; its absence is itself informative.
Frontier Safety needs a falsifiable demonstration, not a benchmark. The Siemens fail-passive study sketches it: a dynamically stable robot executing a certified-style monitored standstill from worst-case states, thousands of trials, documented response-time distribution. Extending that protocol to a learned controller — showing a neural policy's stop behaviour can be bounded tightly enough to enter a functional-safety argument — is the experiment on which every human-adjacent deployment quietly depends.
Frontier The teleoperation-to-autonomy conversion rate is the economic experiment. 1X's NEO fleet will generate the first public longitudinal record of how fast in-home teleoperation data actually converts into autonomous task completion. A published autonomy fraction over time — rising or flat — would test the entire data-flywheel thesis on which the consumer humanoid business model rests.
7 · Engineering requirements
Established Control architecture: every working system is a hierarchy, and the interfaces are load-bearing. ELLMER runs GPT-4 deliberation above 100 Hz force servoing; pi-0.5 runs semantic subtask selection above a flow-matching action head; Gemini Robotics 1.5 pairs an embodied-reasoning model with a separate VLA executor. Engineering the contract between layers — what the slow layer may promise the fast layer — is where these systems succeed or fail, and it is specified by hand in all of them.
Established Onboard compute must fit a power budget the arm race ignores. A 7B-parameter policy at useful rates currently means a desktop-class GPU; a mobile robot carries perhaps 500 W for everything. Quantisation, distillation and on-device variants (Gemini Robotics runs an on-device version (vendor)) are the practical frontier, and each compression step is a fresh, usually unmeasured, hit to the success rates quoted from the datacentre version.
Established Falling is an energy-management problem that scales badly. Kinetic energy grows roughly eightfold with a doubling of size; a 60-kilogram biped falling is a hazard to bystanders and to itself, which is why Brooks advises three metres of standoff from any walking humanoid and why fail-passive stops — the industrial default — are unavailable. Controlled-descent behaviours, protective structure, and joint compliance are required engineering with no standard to design against until ISO 25785-1 lands.
Established Fleet infrastructure is the unglamorous majority of the work. Digit's GXO operation runs on charging docks, tote-conveyor integration and a cloud orchestration layer (vendor); the learned policy is a component in a logistics system, and the system engineering — not the model — is what the 100,000-tote figure actually measures.
Frontier Hands remain unsolved as manufactured objects. Durable five-finger hands with useful load capacity, six-plus degrees of freedom, and any tactile sensing at all do not exist at fleet quality; Tesla's 2025 redesign pause and its multi-supplier search for hand technology show the constraint operating on the best-capitalised programme in the field.
8 · Adjacent technologies
Established This brief sits in a cluster the site already maps. Robotics in Infrastructure covers the deployment surface — inspection, construction, maintenance — where narrow, certifiable autonomy is already economic and where robot foundation models would first matter if they hold up; Automated Construction Systems and Autonomous Supply Chains own the sector-specific cases. Cognitive Architectures is the forty-year-old version of this brief's central architectural question — what structure, if any, must be built in rather than learned — and the VLA hierarchy debate is that literature replayed with gradient descent.
Established Teleoperation is the bridge technology in both directions. It is how training data gets made, and — as 1X's NEO shows — how products ship before autonomy exists; Human-AI Integration covers the human-in-the-loop half of that arrangement, which current humanoid economics quietly depend on.
Frontier World models and video-prediction systems are the nearest substitute technology. If action-conditioned video models become accurate physics simulators, demonstration data loses its monopoly; several labs are betting the data bottleneck dissolves this way rather than through teleoperation scale. The scaling-and-generality debate that frames all of this is owned by Artificial General Intelligence.
Established Component adjacencies decide the cost curve: EV batteries and motor supply chains (from which Unitree's price collapse is largely inherited), MEMS and tactile sensing, and edge inference silicon. None is robotics-specific; all gate robotics timelines.
9 · Institutional requirements
Established Standards bodies are the institution under most pressure. ISO 10218:2025 took eight years and arrived assuming fail-passive machines; ISO 25785-1 for dynamically stable robots began working-group life in 2025 with vendors hosting the meetings. Rob Gruendel, formerly Figure's head of robotics safety and an ANSI/ISO committee member, wrote publicly in 2025 that proving a dynamically stable robot safe is “an order of magnitude more difficult” than for statically stable machines — a rare inside estimate, from someone whose incentive ran the other way while employed.
Frontier No institution currently produces trusted capability numbers. Benchmarks are developer-run, deployment metrics are press releases, and the closest thing to an auditor is three academics with an ALOHA rig. A standing independent evaluation body — the crash-test institute this field lacks — is an institutional requirement stated repeatedly in the methods literature (TRI's empirical-science paper is effectively its charter) and funded by no one.
Established The capital structures differ by bloc and select for different failure modes. US humanoid work is venture-funded against demo-driven valuations (Figure's fifteenfold markup in eighteen months); Chinese work rides industrial policy, municipal buyers and a components ecosystem, which is how Unitree reached profitability on $25,000 units while US flagships slipped. Which structure better survives a long capability winter is an open institutional experiment.
Established Data governance is becoming a labour institution. Teleoperation workforces — 1X's “Turing” operators, AgiBot's collection facility staff, commercial demo-collection services — are the industry's actual production line; their pay, consent regimes and working conditions are unregulated and mostly undisclosed, and they are simultaneously the people whose recorded labour is meant to make their jobs unnecessary.
Frontier Liability allocation is unsettled and will shape architecture. Whether responsibility for a learned controller's failure lands on the manufacturer, the integrator, the data supplier or the deployer is undecided in every jurisdiction; the first serious injury involving a learned policy in a collaborative setting will produce case law that current engineering choices are silently betting on.
10 · Ethical & societal considerations
Established The privacy question is concrete now, not prospective. A home humanoid is a mobile camera platform operated, for years yet, substantially by remote strangers: 1X ships with face blurring, no-go zones and scheduled operator access precisely because the product cannot work without humans watching your kitchen. The CEO's own framing — no data, no product improvement — makes the trade explicit; what is not explicit is the retention, reuse and security of millions of hours of in-home video that doubles as training capital.
Established Hidden human labour is presented as machine capability. When every NEO task in independent testing was teleoperated, the product's apparent autonomy is a staffing arrangement; the same pattern (unbadged teleoperation behind demos) recurs across the sector and is rarely disclosed with the demo. This is a truth-in-advertising problem before it is a philosophy problem.
Frontier Physical safety ethics precede deployment ethics. A 60-kilogram dynamically stable machine among children, the elderly and pets is a hazard class with no governing standard; marketing home humanoids ahead of ISO 25785-1 shifts risk onto buyers who cannot evaluate it. The industrial record is more honest: fenced cells, trained workers, and even there the fail-passive gap is unresolved.
Frontier Labour displacement claims should be weighed against measured capability, in both directions. Warehouse tote handling is genuinely being automated at one facility; that is real displacement pressure on real jobs, and it will grow with each certified narrow deployment. But the $5 trillion labour-replacement narratives price in general dexterity that the measured record does not show, and policy made against vendor projections rather than deployment data will misallocate in whichever direction it errs.
Speculative Cultural knowledge transfer through demonstration raises a quieter question. ELLMER's authors describe retrieval-augmented generation as giving robots “a cultural milieu of knowledge”; corpora of human demonstrations encode the habits, shortcuts and norms of the specific people recorded, and whose kitchen practice becomes the default motor culture of deployed machines is a choice being made implicitly by whoever is cheapest to record.
11 · Civilizational implications
Frontier If curated demonstrations do become open-world competence, the economic object created is a general substitute for physical labour — the largest single input to every economy — and the owning institutions capture rents on a scale that makes current AI concentration look mild. Morgan Stanley's $5 trillion-by-2050 figure is an interested projection, but the direction is right conditional on the capability; the measured record's job is to keep the conditional visible.
Frontier The bloc-level industrial race is already resolving in one direction at the hardware layer. China ships most units, holds the component supply chain, and hosts the largest data-collection operations; the US holds the frontier models and the capital. A world where the machine bodies and the machine minds are made by rival blocs has obvious dependency and security consequences, whichever side of the software/hardware split proves scarcer.
Speculative Demographic ageing gives the optimistic scenario its strongest demand story. Care work, logistics and construction face structural labour shortages in every rich country; a reliable manipulation platform arriving over 25 years would land into demand rather than displacing settled employment, which changes the politics of adoption entirely.
Frontier The pessimistic scenario is a capital-formation event, not a catastrophe. If open-world reliability stays out of reach, the present cycle resolves as prior robotics winters did: tens of billions written down, consolidation around the narrow deployments that pay (tote handling, inspection, fixed workcells), and a decade of reputational drag on the underlying research — which would itself slow the eventual real thing. Bubbles in general-purpose technologies have historically still left useful infrastructure behind; teleoperation rigs, cheap actuators and demonstration corpora would be this one's residue.
12 · Timelines
These horizons track the distance from curated demonstration to independently verified open-world reliability, on the current evidence base.
- 10 yr: Frontier Narrow certified deployments scale: tote-class logistics, inspection and workcell tasks reach thousands of units with published reliability statistics; ISO 25785-1 and successors land; independent evaluation of released VLAs becomes routine, and the gap between developer-reported and replicated success rates is a tracked number. Teleoperation-heavy home robots exist as a service category; their autonomy fraction is the metric to watch.
- 25 yr: Speculative General household manipulation at useful reliability is plausible if tactile sensing, recovery learning and data economics each deliver — three independent bets. Cross-embodiment foundation models either become the standard control stack (the VLM dependency keeps paying) or survive as pretraining for per-platform specialists, which the RT-X negative transfers already foreshadow.
- 50 yr: Speculative Physical labour scarcity ceases to bind rich-country logistics, construction and care if the 25-year branch resolved positively; the form factor question resolves empirically, and on Brooks' record the winning bodies look like neither today's humanoids nor humans.
- 100 / 250+ yr: Handwave Claims at this range — fully self-improving embodied AI, robot-built civilisation off-Earth — assert the compounding of capabilities whose first instance has not been independently measured; nothing in the demonstration-to-deployment record licenses extrapolation past the 50-year line.
13 · Technology tree & dependencies
- Depends on nothing on this map as a hard blocker: the inputs are vision-language models, teleoperation data and actuator supply chains, none owned by another brief. The architectural question it inherits — how much structure must be built in rather than learned — is the live thread of Cognitive Architectures, and the scaling-generality debate that would most change its trajectory is owned by Artificial General Intelligence.
- Requires (not on this map) fingertip tactile sensing manufactured near human mechanoreceptor density, since demonstration video does not contain the contact signals dexterity consists of; independent, statistically powered replication of vision-language-action success rates, without which the field cannot distinguish capability from marketing; demonstration collection at million-hour scale for contact-rich tasks, against current economics of $8–35 per teleoperated episode; a certification route by which a learned controller on a dynamically stable robot can enter a functional-safety argument, closing the fail-passive gap; and a paying customer base for humanoids beyond research labs and pilot showcases, which the Unitree filing shows does not yet exist at scale.
- Enables the deployment briefs downstream: reliable learned manipulation is the missing input for Robotics in Infrastructure, Automated Construction Systems and Autonomous Supply Chains, each of which currently engineers around its absence with narrow autonomy and human oversight.
- Adjacent to Human-AI Integration, which owns the teleoperation and shared-control arrangements this field runs on today, and to Cognitive Architectures, whose forty-year literature on hierarchical control the VLA stack is rediscovering under a different loss function.
14 · Common misconceptions & speculative claims
The claims in circulation deserve answers with numbers attached.
Handwave “Robots will learn dexterity by watching video of people.” This is the stated data strategy of Tesla and Figure, and it works by assertion at the decisive step: video contains no forces. Brooks' case — roughly 17,000 mechanoreceptors per human hand, demonstrable collapse of human dexterity under fingertip anaesthesia, 65 years of manipulation research, and the observation that every prior AI success rode on domain-appropriate input engineering — has no published rebuttal with data; the burden currently sits entirely on the claim.
Handwave “Robotics is about to have its ChatGPT moment.” The analogy fails on its own arithmetic. Language models pretrained on trillions of tokens of the target behaviour; the largest open robot-demonstration corpora hold hundreds of hours (DROID: 350) or around a million episodes (Open X-Embodiment, AgiBot World), collected at $8–35 per episode. The web contained the training data for language; nothing comparable contains physical contact, and building it is priced in human labour, one hour per hour.
Frontier “Cross-embodiment transfer is solved.” Overstated in both directions. RT-X measured real pooled-data gains (about 50 percent mean improvement in small-data domains; a threefold emergent-skill gain for RT-2-X) — and real negative transfer where local data was rich (27 versus 40 percent on Bridge). Gemini Robotics 1.5's zero-shot claim across ALOHA 2, Apollo and Franka is a vendor report awaiting any independent replication. Transfer is a demonstrated training-time effect with unknown reach, not a solved property.
Established “A 90 percent benchmark success rate means a 90 percent capable robot.” Refuted by the field's own audits: LIBERO-PRO found 90-plus percent models whose outputs did not change under corrupted instructions; paraphrasing alone costs 22–52 points; the Macquarie replication reproduced two celebrated models at 13 percent on a real bimanual platform; and the TRI critique documents that most published rates come without trial counts, initial-condition control or stated success criteria. The number is not yet a measurement.
Frontier “Millions of humanoids are shipping imminently.” The delivered-unit record for 2025: about 5,500 from the volume leader (its own IPO filing), roughly 1,000 built against a 5,000–10,000 target at Tesla, one disclosed cumulative-throughput commercial deployment (100,000 totes, one warehouse). Forecasts that doubled twice in a year (Morgan Stanley's China shipment outlook) are extrapolating a base measured in thousands. Rapid growth is real; the “millions” are unscheduled.
Frontier “Teleoperation is a brief bootstrap phase.” This is the business model, not a finding. The conversion rate from teleoperation data to autonomy has never been published by any consumer robot company, and the first public longitudinal test (1X's NEO fleet) is only now beginning; until an autonomy fraction is reported rising, a $20,000 teleoperated robot is a remote-labour service with a hardware deposit.
Established “The robots in car factories are doing production work.” The two flagship cases both thin under inspection: Figure's CEO declined to describe the BMW deployment's scope while threatening litigation over reporting on it, and Tesla's Optimus units were moving batteries in Tesla's own workshops at less than half human efficiency by supply-chain accounts. Pilot presence is documented; production economics are not.
Speculative “General embodied competence requires humanoid form.” A coherent hypothesis, unproven and contested from inside the field: the built environment argument cuts both ways (wheels and grippers already serve most of it), the working deployments do not use hands, and Brooks' fifteen-year prediction is that the successful machines will look like neither today's humanoids nor humans. The form factor is a bet, and it is the most capital-intensive bet in the industry.