Perception, Not Actuation
Why the more important bet in Physical AI is building better senses, not better hands
Prologue: The Robot That Can Fold a Shirt
Every few weeks now, a new video crosses my feed: a humanoid robot folding laundry, pouring a drink, walking a warehouse floor without falling over. The comments are always some version of the same sentence. Look how far we’ve come. And it is genuinely impressive, the actuator, the balance, the grip strength calibrated to hold an egg without crushing it.
But I keep having the same reaction, and it took me a while to name it properly. The hands are not the hard part anymore. They haven’t been for a while. The hard part, the part almost nobody is funding, is getting a machine to actually understand what it’s looking at, with the depth and context a person brings to the same room without even noticing they’re doing it. We’ve spent a decade racing toward better hands while the harder problem, better eyes, sat mostly unaddressed a few feet to the left of the demo stage.
That’s the argument I want to make here, and it turns out there’s a name for the paradox at the center of it, one that’s almost forty years old.
Part I: The Paradox Nobody Fully Explains
In 1988, the roboticist Hans Moravec wrote a sentence that has aged better than almost anything else said about artificial intelligence that decade: it is comparatively easy to make a computer perform at an adult level on an intelligence test, and comparatively difficult or impossible to give it the perceptual and physical skills of a one-year-old. Chess, it turns out, was never the hard problem. Recognizing your mother’s face across a crowded room, at any angle, in any light, having never been told the rule for how to do it, that was the hard problem, and it still is.
Moravec’s own explanation was evolutionary. Reasoning, he argued, is the thinnest, most recent veneer on top of a billion years of sensorimotor knowledge, the stuff evolution spent unimaginable time optimizing because an organism that couldn’t see a predator or catch its food didn’t survive long enough to do abstract math. Abstract reasoning is a few hundred thousand years old at most. Depth perception is closer to five hundred million. Machines, with no evolutionary history to inherit, ended up with the order reversed: reasoning came cheap because we could specify it as rules, while perception stayed expensive because nobody, including us, ever fully worked out the rules to begin with.
I want to be honest about a real critique of this idea before I lean on it, because the honest version of this argument is stronger than the tidy one. The computer scientist Arvind Narayanan has pointed out that Moravec’s paradox may say less about what’s objectively hard for machines than about what the AI research community has historically chosen to work on, chess over dexterity, benchmarks over bodies, because chess was legible and fundable in a way that dexterity wasn’t. That’s a fair challenge, and it matters for how we read the pattern now. Whether perception is hard because it’s intrinsically harder, or hard because we spent seventy years pointing our best researchers and our funding at reasoning instead, the practical result today is the same either way. We have systems that pass the bar exam and still struggle to reliably fold a shirt the first time they’ve seen that particular shirt.
If anything, the newest generation of models has sharpened the paradox rather than closed it. A current large language model can pass a bar exam or place at an international math olympiad in the time it takes to load the page, exactly the “hard adult” territory Moravec’s contemporaries assumed would be the last frontier. Put the same class of model in front of a purpose-built physical reasoning benchmark, judging how objects will fall, balance, or collide, and performance drops sharply, closer to guessing than to competence. Herbert Simon, one of the field’s founders, said in the 1960s that everything interesting about cognition happens above the hundred-millisecond mark, roughly the time it takes to consciously recognize your mother’s face, which was his way of saying the unconscious perceptual work underneath that recognition wasn’t worth studying closely. Sixty years and several AI winters later, that hundred-millisecond floor turned out to be hiding most of the actual difficulty, not the least of it.
Part II: Where the Money and the Hype Are Actually Going
Here’s what makes this moment specifically worth writing about. Humanoid robotics funding cleared roughly six billion dollars in 2024 alone, according to Stanford’s AI Index, spread across Figure, Tesla’s Optimus program, 1X, Boston Dynamics, and a growing list of competitors racing to put a body in a warehouse or a factory. Nvidia, Google DeepMind, and Physical Intelligence have all shipped robotics foundation models aimed squarely at closing what researchers now openly call the sensorimotor gap. This is, almost entirely, a bet on actuation. Can we get the arm, the hand, the leg, to do the thing.
It’s the more fundable bet, and I understand why. A robot folding a shirt makes a compelling demo. A system that simply understands its environment more completely doesn’t produce the same viral video, even though it’s very plausibly the harder and more valuable half of the same problem. Self-driving cars taught this lesson already, a decade earlier and more expensively than anyone expected. Steering, braking, and accelerating a car precisely were solved problems years before the technology was ready for the road. What actually held the entire industry back, year after year, delay after delay, was perception: reliably identifying what a shadow, a plastic bag, a stopped truck, or a pedestrian half-obscured by rain actually was, fast enough and confidently enough to bet a life on the answer. Nobody was funding “better windshields.” Everyone was funding better driving. The bottleneck turned out to be the seeing, not the driving.
Part III: The Argument Inside Perception Itself
It’s worth pausing on how much more unsettled the perception half of this bet actually is, even among the companies most committed to it. Actuation, for all its engineering difficulty, has more or less converged: a humanoid robot today has two arms, a gripper or a hand, legs or wheels, and the design space, while hard to execute, is not seriously contested. Nobody is arguing about whether a robot should have hands.
Perception has no such consensus. The self-driving industry has spent the better part of two decades disagreeing, expensively, about the most basic question in the field: what should a machine actually use to see. One camp, led for years by Waymo, bet on layering lidar, radar, and cameras together, redundant senses cross-checking each other the way human vision is itself backed up by balance, proprioception, and hearing. The other camp, led publicly by Tesla, bet that cameras alone, processed well enough, should be sufficient, on the argument that humans manage the same roads with two eyes and no laser rangefinder. Both camps have spent billions. Neither has definitively won. That’s not a sign of an easy problem quietly being solved in the background while actuation gets the headlines. It’s a sign that the field still doesn’t agree on the basic architecture of faithful sensing, a decade after it converged on the far simpler question of how to build a hand.
Part IV: The Older, Quieter Attempt to Solve This
The strange part is that someone tried to solve the perception half of this problem decades before anyone thought to attach it to a robot, and almost nobody remembers it now.
In 1945, Vannevar Bush, who had run American science policy through the Second World War, published an essay called “As We May Think” imagining a device he called the Memex: a desk that could store and instantly retrieve everything a person read, wrote, and thought, linked by association the way a human mind actually works rather than the rigid folders and indexes libraries used. It was, in effect, the first serious proposal for a second brain, and it was written before the transistor existed.
Fifty years later, a computer scientist named Gordon Bell decided to actually build it, on himself. Starting in 1998, Bell wore a device called a SenseCam around his neck that took a photograph every time its sensors detected a meaningful change in light or motion, alongside a system that archived every document, email, phone call, and piece of correspondence he touched. He called the project MyLifeBits, and by the time he wound it down in 2007, it held more than a million document pages, well over a hundred thousand photographs, and years of continuous ambient record of an actual human life. It was, quite literally, a perception-first bet: capture the raw sensory record faithfully, and worry about what to do with it after.
Bell’s own conclusion, stated plainly in later interviews, is the detail I think matters most for everything I’m arguing here. Capturing the material, it turned out, was the tractable part. The part nobody had solved, the part that made the archive more curiosity than tool, was turning a decade of raw perceptual record into something a person could actually use, structured, retrievable, alive to what mattered and quiet about what didn’t. Bell had built extraordinary senses and no real judgment to sit behind them. He said later that the smartphone quietly ended the project, because everyone started accidentally lifelogging through their camera roll without meaning to, capture had become ambient and free, while the harder second half, making sense of it, stayed almost exactly as unsolved as it was in 1998. It’s a small footnote of history that Microsoft, the same company that employed Bell for the entire life of that project, announced its own version of ambient capture for the PC within days of his death in 2024.
Part V: What the Body Already Knew
There’s a biological reason this order of difficulty, perception first, hard; action after, comparatively easy, shouldn’t actually surprise anyone. Roughly thirty percent of the human cerebral cortex is given over to visual processing alone, more real estate than touch and hearing combined, which get roughly eight and three percent respectively. Some estimates, accounting for the full downstream network that visual signals recruit across the parietal and temporal lobes, put the number closer to half the entire cortical surface.
Think about what that split is actually telling you. The brain does not treat seeing as a light preprocessing step before the real thinking starts. Seeing largely is the thinking, or close enough to it that evolution built roughly a third of our most expensive tissue to do it well, while the deliberate, verbal reasoning we’re inclined to think of as intelligence itself runs on comparatively little of the remaining budget. We built our machines backwards relative to how we ourselves are actually built, cheap perception bolted onto expensive reasoning, when the biological template has always been the other way around.
Part VI: The Actual Bet
Put Moravec, Bell, and the visual cortex together and a specific, contrarian position falls out of it fairly cleanly. The industry’s current wave of enthusiasm for Physical AI is, almost entirely, a bet on actuation: robots that act in the world. That’s a real and valuable problem. It is not, on the evidence above, the harder or more differentiated one. The harder problem, the one with a forty-year paper trail of researchers running into it and a quieter, less-funded history of people trying to solve it on their own bodies, is faithful perception: capturing what actually happened, in context, richly enough that something downstream, human or machine, can make good judgment out of it later.
That’s the bet behind the long-term roadmap we’re building toward at Twelfth Brain, stated as plainly as I can put it. We are not building hands. We’re building senses, ambient and wearable capture that gives the second brain a more faithful record of an actual life to work from, the way Bell tried to build for himself twenty-five years too early, before there was any judgment layer capable of doing something useful with what he’d captured. The judgment is the part the rest of this series has been about. This piece is about the other half: judgment has nothing to compound from if the perception underneath it was never faithful to begin with.
A Few Honest Distinctions
First: when you see a demo, ask which half of the problem it’s actually solving. A robot that acts smoothly in a rehearsed warehouse aisle has not necessarily solved perception; it may just have memorized that aisle.
Second: don’t confuse the fundable problem with the important one. Actuation gets the funding and the viral clip. That’s a fact about incentives, not a fact about which problem is harder or matters more.
Third: capture is not the same achievement as understanding. Gordon Bell proved you could capture a life faithfully in 1998. He also proved, by his own account, that capture alone doesn’t get you anywhere close to a usable second brain without a judgment layer sitting on top of it.
Fourth: look at where the biology actually spent its budget. Nature had five hundred million years and no ideology about which problem was more prestigious. It spent nearly a third of the cortex on seeing. That allocation is itself a kind of argument.
Fifth: before betting on a Physical AI thesis, ask whether it’s a hands bet or an eyes bet. Both are real. Only one of them has a forty-year paper trail of turning out to be the harder, more foundational problem underneath the other.
Further Reading
On the paradox itself
Hans Moravec’s Mind Children (1988) is the original source. Arvind Narayanan’s public writing on the paradox is the most useful corrective, arguing it reflects research incentives as much as any intrinsic law of difficulty.
On the older attempt to solve perception directly
Vannevar Bush’s 1945 essay “As We May Think” is the founding document of the entire idea of a second brain, written before the transistor existed. Gordon Bell and Jim Gemmell’s Your Life, Uploaded is the first-person account of what it actually took to build a faithful perceptual record of one human life, and what was still missing once he had.
On where the funding and the research are actually going right now
Stanford’s AI Index report is the most complete public accounting of how much capital is currently chasing humanoid actuation, and against what benchmarks.


