Brilliant at chess. Blind to door handles.
In 1997 a machine beat the world champion at chess. In 2011 one won at Jeopardy. In 2016, Go fell — the game with more positions than atoms in the universe. By 2023, models were passing the bar exam, writing working code, and reading medical scans at specialist level. And through all of it, not one of these systems could fold a towel, load a dishwasher, or open an unfamiliar door.1
The roboticist Hans Moravec named this inversion in 1988: the things we consider hard — logic, calculation, strategy — are easy for machines, and the things a toddler does without thinking — seeing, grasping, walking on gravel — are among the hardest problems ever attempted.2 Evolution spent half a billion years engineering perception and movement, and a few thousand on symbolic reasoning. We automated our newest ability first, and mistook it for the summit.
So intelligence grew up behind glass. Every gain of the computing century — every model, every network, every interface — was a gain in what could be computed, displayed, and transmitted. The internet moved information at the speed of light and moved matter not at all. Consider the strangeness of the result: humanity built more than four million industrial robots, and nearly every one of them is blind — a playback machine repeating coordinates it cannot see, in cages built to protect people from its ignorance.3 We had automation without intelligence on the factory floor, and intelligence without hands everywhere else.
The screen was not a stage of progress. It was a confinement.
The data behind Figure I
| Milestone | Year | Domain |
|---|---|---|
| Deep Blue defeats Kasparov | 1997 | Chess |
| Watson wins Jeopardy | 2011 | Language trivia |
| AlphaGo defeats Lee Sedol | 2016 | Go |
| AlphaFold solves protein structure | 2020 | Biology |
| LLMs pass the bar exam | 2023 | Professional reasoning |
| General manipulation & locomotion | 2024– | The physical world — falling now |
Four curves crossed. Quietly, in the same two years.
The body got cheap. A humanoid platform that cost a research budget in 2023 — a quarter of a million dollars and up, hand-assembled, waitlisted — now has commercial peers listing near the price of a used car. One Hangzhou manufacturer put a full humanoid on sale for sixteen thousand dollars.4 The cost of a capable machine body has collapsed roughly 70 percent in two years, and it is on the same curve that took the mobile phone from a briefcase to a birthday present. Bodies are becoming what handsets became: commodity.
The mind learned to act. This is the part the screen era said was impossible. The same architecture that learned language by reading the internet turned out to learn action by watching and doing — one model, camera to motor, perception to movement. The lab results stopped being demos and started being shifts: warehouse picks, kitchen tasks, unfamiliar objects in unfamiliar rooms.5 Moravec's paradox is not being argued with. It is being retired.
The capital arrived. Money is a lagging indicator of belief, and belief has moved: a record $27 billion went into physical AI in a single year — into bodies, into minds for bodies, into the factories that make both.6 One major bank now models a multi-trillion-dollar humanoid economy by mid-century.7 The turn is no longer a thesis. It is a line item.
And the tools turned on themselves. This is the curve the first three depend on, and the one almost nobody prices. The product-development cycle — design, prototype, tool, test, manufacture — was the slowest machine of the industrial century; it is why Bell's fuses burned for decades. That cycle is now collapsing. Models write the code. Generative systems draw the parts. Digital twins replace the pilot plant, and a mind can live ten thousand simulated years of practice before its body is bolted together.8 A team of ten with today's toolchain out-builds yesterday's division — which means, for the first time since 1925, you do not need to be Bell to do what Bell did. The cycle that took decades now compounds in weeks. That is not an efficiency. It is a change in who gets to exist.
The data behind Figure II
| Point | Value | Basis |
|---|---|---|
| Research humanoid platform, 2023 | $150K–$250K+ | Hand-built, limited availability |
| Commercial humanoid, 2024–25 | $16K list | Unitree G1 class |
| Two-year cost decline | ~−70% | Indexed, capable-platform class |
| Capital into physical AI | $27B | Record single year |
A word from 1969 is about to run the decade. Mechatronics.
In 1969, an engineer at Yaskawa Electric named Tetsuro Mori needed a word for what his team actually built — machines in which the mechanism and the electronics were one design, not two — and coined mechatronics.9 For half a century it stayed a specialist's word: the craft of servo motors, sensors, and control loops, practiced in pockets. Japanese factory automation. German machine tools. The autofocus in a camera, the read head flying nanometers above a hard-drive platter, the anti-lock brake. Every object that moves with precision has mechatronics inside it. Almost no one outside the guild has ever said the word aloud.
It stayed a guild for one reason: the mind was scripted. Classical mechatronics meant control laws derived and tuned by hand, machine by machine — months of engineering per degree of freedom, frozen the day the product shipped. Motion stayed expensive, iteration stayed slow, and the discipline stayed a priesthood with a decade-long apprenticeship.
Both constraints just broke, from opposite directions. The learned mind replaces the hand-tuned law — one model, retrained overnight, controlling joints it has never met. And the parts went commodity along two supply chains built for other products: the electric car industrialized motors, batteries, and power electronics; the smartphone industrialized cameras, inertial sensors, and radios. A humanoid is, in its bill of materials, the child of a car and a phone.10
Now run the arithmetic of the era. A capable humanoid carries roughly thirty actuated joints. One bank's base case reaches a million units a year in the 2030s11 — thirty million precision joints annually, a component industry born at aerospace tolerances and consumer-electronics volumes. The actuator is the transistor of this era — and Dispatch No. 002 is about what happened to the people who gave the last one away.
So mechatronics is about to do what software engineering did after 2010: leave the guild and become the default discipline. Every serious product company becomes, quietly, a mechatronics company — the way every company became a software company whether it planned to or not. The defining engineer of the next decade is fluent in all three layers at once: mechanism, electronics, learned mind. We are hiring exactly them.
There is no internet of touch.
Here is the asymmetry the screen era never had to think about. Language models got their education for free: humanity had spent thirty years writing the internet, and a mind could read on the order of fifteen trillion words of it before its first conversation.12 Action has no such inheritance. Nobody uploaded a trillion grasps. There is no archive of what a hand feels when an egg is about to crack, no Wikipedia of friction. The single largest input to physical intelligence — experience of the world — does not exist yet. It has to be manufactured.
It gets manufactured in two places. The first is simulation: GPU-parallel physics engines run thousands of worlds at once, orders of magnitude faster than reality — a mind can accumulate years of practice per wall-clock hour, falling ten thousand times before its body is ever bolted together.13 The second is the world itself: every deployed body is a data factory, and its output — real contact, real failure, real edge cases no simulator imagined — is the scarcest substance in the field.
Now notice what kind of substance it is. Text leaked; that was the whole tragedy of Dispatch No. 002 — you cannot fence a corpus everyone can read. But a fleet's experience is rivalrous and private. It exists exactly once, on the ledger of whoever fielded the fleet. Which closes the first true compounding loop in the history of AI: deploy bodies, harvest experience, train a better mind, deploy more bodies. Scale begets skill begets scale. In the screen era, data leaked. In the physical era, experience is titled. Whoever holds the fleet holds the corpus — and the corpus is the moat.
The data behind Figure IV
| Corpus | Scale | Character |
|---|---|---|
| Text available to language models | ~15T tokens | Public — leaked to everyone at once |
| Recorded robot-action data | ~10⁸ episodes, generously | Scarce — must be manufactured |
| Simulation leverage | ~10³–10⁴× real time | GPU-parallel physics, thousands of environments |
| Fleet experience | Compounds with units deployed | Rivalrous, private — behaves like property |
The screen economy was the small one.
Here is the number the software century never liked to say out loud. After fifty years of digital triumph — after eating media, retail, finance, and every industry with the word "information" in it — the digital sector amounts to perhaps fifteen percent of world output.14 Everything else is atoms. Food, buildings, logistics, care, energy, manufacture: the overwhelming majority of the roughly $110 trillion world economy is earned by moving matter, and it pays out more than $50 trillion a year in human wages to do it.15
The screen era competed for attention, because attention was all a screen could reach. Physical intelligence competes for work. The addressable market of a mind with hands is not advertising — it is labor itself, the largest expenditure of the human species. When the machines that think can finally act, the prize is not another app economy. It is the economy.
And the demand is not hypothetical — it is demographic, and it is already here. The world is running out of hands before the machines arrive, not because of them: American manufacturing projects 1.9 million jobs unfillable within a decade; the WHO counts a shortfall of 10 million health workers by 2030; Japan — the oldest large economy, and the preview of everyone else's future — models an 11 million worker deficit by 2040.16 The hardest political question of automation has quietly inverted: in the industries that feed, house, heal, and move us, there is no one left to displace.
The data behind Figure V
| Market | Annual size | Basis |
|---|---|---|
| Global advertising | ~$0.8T | Industry estimates, mid-2020s |
| Global enterprise software | ~$1T | Industry estimates, mid-2020s |
| Global labor income | $50T+ | ILO labor-income share of ~$110T world GDP |
The data behind Figure VI
| Shortfall | Scale | Source & horizon |
|---|---|---|
| US manufacturing, unfillable roles | 1.9M | Deloitte / NAM projection, 2033 |
| Global health workers | 10M | WHO projection, 2030 |
| Japan, economy-wide worker deficit | 11M | Recruit Works Institute, 2040 |
Minds, made here. Held here.
So this is the founding read. Intelligence is leaving the screen — not gradually, but on three converging curves, in plain sight, priced daily. What comes next will not be decided by who writes the cleverest model. The screen era already taught that lesson: cleverness disperses. It will be decided by who builds and holds — because a mind with a body is not a file to be copied but an asset that compounds, accumulating hours of the real world that exist exactly once, on exactly one ledger.
ARBX is built for that reading. We build minds that live in the real world — machines that sense, reason, and act. And we hold what we build: every mind we deploy stays on our ledger, learns on our network, and compounds in our name. Made here, held here. Why holding is the whole game — why the greatest laboratory in history died a line item while a firm that invented nothing came to hold fourteen trillion dollars — is the subject of Dispatch No. 002.
The century that ends did its thinking behind glass. The century that begins will be done by hand — and built with tools that build themselves. Moves matter.
Notes & Sources
- Deep Blue–Kasparov, 1997; IBM Watson, Jeopardy, 2011; AlphaGo–Lee Sedol, 2016; AlphaFold, 2020; GPT-4 bar-exam performance, 2023. milestones ↗ ↑
- Hans Moravec, Mind Children, 1988 — "it is comparatively easy to make computers exhibit adult-level performance… and difficult or impossible to give them the skills of a one-year-old." moravec's paradox ↗ ↑
- International Federation of Robotics, World Robotics: over four million industrial robots in operation worldwide. ifr — world robotics ↗ ↑
- Unitree G1 humanoid, listed from $16,000, 2024 — against research-grade humanoid platforms at $150K–$250K+ two years prior. unitree ↗ ↑
- Vision-language-action models: single networks mapping camera input to motor output across unfamiliar tasks and objects, 2023–. robot learning ↗ ↑
- Capital deployed into physical AI — bodies, embodied models, and their supply chain — record single year. ARBX research series; treated in full in Dispatch No. 003. arbx dispatches ↗ ↑
- Bank research now models a humanoid-robot economy measured in the trillions by 2050. bank research ↗ ↑
- Simulation-first robot learning (GPU-accelerated physics, domain randomization, sim-to-real transfer), model-written software, and generative engineering design — the compression of the design-build-test loop, 2023–. sim-to-real ↗ ↑
- Tetsuro Mori, Yaskawa Electric, coins "mechatronics," 1969; trademarked by Yaskawa in 1971 and later released for public use. mechatronics ↗ ↑
- Humanoid bills of materials draw on the EV powertrain stack (motors, power electronics, batteries) and the smartphone sensor stack (cameras, IMUs, radios); ~28–40 actuated joints per platform. humanoid actuation ↗ ↑
- Bank research base cases for humanoid deployment reach one million units per year in the 2030s. bank research ↗ ↑
- Frontier language models are trained on corpora on the order of 15 trillion tokens. llm training corpora ↗ ↑
- GPU-parallel physics simulation (Isaac Gym class): thousands of concurrent environments at multiples of real time — three to four orders of magnitude of practice leverage, with sim-to-real transfer closing the gap. simulation-first learning ↗ ↑
- Digital-economy share of global GDP commonly estimated near 15%. oecd digital ↗ ↑
- ILO labor-income share of roughly half of world GDP (~$110T, IMF) implies global wages above $50T annually. ilo ↗ ↑
- Deloitte / National Association of Manufacturers: ~1.9M US manufacturing roles unfillable by 2033. WHO: projected shortfall of ~10M health workers by 2030. Recruit Works Institute: ~11M worker deficit in Japan by 2040. who — health workforce ↗ ↑
What does "intelligence is leaving the screen" mean?
For seventy years, machine intelligence could only compute, display, and transmit — it lived behind glass. It means models can now perceive and act in the physical world through machine bodies: minds that move matter, not just information.
What is physical intelligence?
The fusion of a learned mind with a machine body — systems that sense, reason, and act in the real world. Unlike software, a physical mind accumulates real-world experience that cannot be copied, making it an asset rather than a file.
Why is this happening now?
Four curves converged: machine bodies collapsed roughly 70% in price in two years; foundation models learned to control bodies directly from perception; capital deployed a record $27 billion into physical AI in a single year; and the toolchain turned on itself — simulation, model-written code, and generative design compressed the product-development cycle from decades to weeks, so a small institution can now build what once required fifteen thousand people.
What is mechatronics and why does it matter now?
Mechatronics — coined at Yaskawa in 1969 — is the discipline that fuses mechanism, electronics, and control into one design. It stayed a specialist craft for fifty years because control had to be hand-tuned per machine. With learned minds replacing scripted control, and EV and smartphone supply chains commoditizing its parts, mechatronics is becoming the default engineering discipline — the way software engineering did after 2010.
What is Moravec's paradox?
Named by roboticist Hans Moravec in 1988: the skills humans find hard — chess, logic, calculation — are easy for machines, while the skills a toddler finds effortless — seeing, grasping, walking on gravel — are among the hardest problems ever attempted. Evolution spent half a billion years engineering perception and movement, and only millennia on abstract reasoning. Figure I of this dispatch draws the paradox — and its ending.
What is the difference between AI and physical AI?
Screen-era AI reads, writes, and predicts — it moves information. Physical AI (also called embodied AI) adds a body: it senses, reasons, and acts — it moves matter. The economic difference is total: software competes for attention; physical intelligence competes for work.
What is a vision-language-action (VLA) model?
A single neural network that maps what a machine sees — and what it is told — directly to motor commands: perception to movement, one model. It is the architecture that ended hand-tuned robot control. The same approach that learned language by reading the internet learns action from demonstration and simulation.
How much does a humanoid robot cost?
Two years ago, a capable research humanoid ran $150,000–$250,000+, hand-built and waitlisted. Today a commercial-class humanoid lists near $16,000 — a roughly 70 percent collapse in the cost of a machine body, riding the EV and smartphone supply chains. The curve that took the handset from briefcase to pocket is now running through robotics.
How big is the physical AI market?
Capital deployed a record $27 billion into physical AI in a single year, and bank research models a multi-trillion-dollar humanoid economy by 2050. The deeper measure is the addressable market of a mind that can act: not attention, but labor itself — more than $50 trillion a year in global wages.
Will physical AI replace human jobs?
It arrives first where humans are scarce, at risk, or already absent — factories short millions of workers, care systems short of hands, logistics running on overtime. As in every automation wave, tasks change before jobs vanish: the dull, dangerous, and understaffed work goes first. What is genuinely new is the question of who holds the machines — which is the subject of Dispatch No. 002.
Where does the training data for robots come from?
There is no internet of action — nobody uploaded a trillion grasps. The corpus is manufactured in two places: GPU-parallel simulation, where a mind accumulates years of practice per hour across thousands of virtual worlds, and deployed fleets, where every body harvests real experience. Fleet experience is rivalrous and private — the first AI training data that behaves like property.
What is ARBX?
The institution of physical intelligence. ARBX builds minds that live in the real world and holds what it builds — every mind it deploys stays on its ledger, learns on its network, and compounds. Made here, held here. New York, London, New Delhi.
How do I read the ARBX Dispatches?
The Dispatches are ARBX's founding papers — the doctrine, the precedent, the build — published at arbx.com/dispatches. Become a founding reader at arbx.com — day-one readers remembered. Open a channel: contact@arbx.com.