Prologue: The Robot That Can Fold a Shirt
Every few weeks now, I come across a new video on one of my social feed: a humanoid robot folding laundry, pouring a drink, or walking a warehouse floor without falling over. The comments are always some version of the same sentence. Look how far we've come. It is genuinely impressive, the actuator, the balance, the grip strength calibrated to hold an egg without crushing it.
But I keep having the same reaction which took me a while to name it properly. The hands haven’t been the hard part for a while now. The hard part, the part almost nobody is funding, is getting a machine to actually understand what it’s looking at, with the depth and context a person brings to the same room without even noticing they’re doing it. We’ve spent a decade racing toward better hands while the harder problem, better eyes, sat mostly unaddressed a few feet to the left of the demo stage.
That’s the argument I want to make here, and there turns out to be a name for the paradox sitting at the center of it, one that’s almost forty years old.
Part I: The Paradox Nobody Fully Explains
In 1988, the roboticist Hans Moravec wrote a sentence that has aged better than almost anything else said about artificial intelligence that decade: it is comparatively easy to make a computer perform at an adult level on an intelligence test, and comparatively difficult or impossible to give it the perceptual and physical skills of a one-year-old. Chess, it turns out, was never the hard problem. Recognizing your mother’s face across a crowded room, at any angle, in any light, having never been told the rule for how to do it, that was the hard problem, and it still is.
Moravec’s own explanation was evolutionary. Reasoning, he argued, is the thinnest, most recent veneer on top of a billion years of sensorimotor knowledge, the stuff evolution spent unimaginable time optimizing because an organism that couldn’t see a predator or catch its food didn’t survive long enough to do abstract math. Abstract reasoning is a few hundred thousand years old at most. Depth perception is closer to five hundred million. Machines, with no evolutionary history to inherit, ended up with the order reversed: reasoning came cheap because it could be specified as rules, while perception stayed expensive because nobody, including us, ever fully worked out the rules to begin with.
It’s worth being honest about a real critique of this idea before leaning on it, because the honest version of the argument is stronger than the tidy one. The computer scientist Arvind Narayanan has pointed out that Moravec’s paradox may say less about what’s objectively hard for machines than about what the AI research community has historically chosen to work on, chess over dexterity, benchmarks over bodies, because chess was legible and fundable in a way dexterity wasn’t. That’s a fair challenge, and it matters for how the pattern reads today. Whether perception is hard because it’s intrinsically harder, or hard because the field spent seventy years pointing its best researchers and its funding at reasoning instead, the practical result is the same either way. We have systems that pass the bar exam and still struggle to reliably fold a shirt the first time they’ve seen that particular shirt.
If anything, the newest generation of models has sharpened the paradox rather than closed it. A current large language model can pass a bar exam or place at an international math olympiad in the time it takes to load the page, exactly the “hard adult” territory Moravec’s contemporaries assumed would be the last frontier. Put the same class of model in front of a purpose-built physical reasoning benchmark, judging how objects will fall, balance, or collide, and performance drops sharply, closer to guessing than to competence. Herbert Simon, one of the field’s founders, said in the 1960s that everything interesting about cognition happens above the hundred-millisecond mark, roughly the time it takes to consciously recognize your mother’s face, his way of saying the unconscious perceptual work underneath that recognition wasn’t worth studying closely. Sixty years and several AI winters later, that hundred-millisecond floor turned out to be hiding most of the actual difficulty.
Part II: Where the Money and the Hype Are Actually Going
Here’s what makes this moment specifically worth writing about. Humanoid robotics funding cleared roughly six billion dollars in 2024 alone, according to Stanford’s AI Index, spread across Figure, Tesla’s Optimus program, 1X, Boston Dynamics, and a growing list of competitors racing to put a body in a warehouse or a factory. Nvidia, Google DeepMind, and Physical Intelligence have all shipped robotics foundation models aimed squarely at closing what researchers now openly call the sensorimotor gap. This is, almost entirely, a bet on actuation. Can the arm, the hand, the leg, be made to do the thing.
It’s the more fundable bet, and I understand why. A robot folding a shirt makes a compelling demo. A system that simply understands its environment more completely doesn’t produce the same viral video, even though it’s very plausibly the harder and more valuable half of the same problem. Self-driving cars taught this lesson already, a decade earlier and more expensively than anyone expected. Steering, braking, and accelerating a car precisely were solved problems years before the technology was ready for the road. What actually held the entire industry back, year after year, delay after delay, was perception: reliably identifying what a shadow, a plastic bag, a stopped truck, or a pedestrian half-obscured by rain actually was, fast enough and confidently enough to bet a life on the answer. For years, all the money went toward better driving. The actual bottleneck was the seeing.
Part III: The Argument Inside Perception Itself
It’s worth pausing on how much more unsettled the perception half of this bet actually is, even among the companies most committed to it. Actuation, for all its engineering difficulty, has more or less converged: a humanoid robot today has two arms, a gripper or a hand, legs or wheels, and the design space, while hard to execute, isn’t seriously contested. Nobody is arguing about whether a robot should have hands.
Perception has no such consensus. The self-driving industry has spent the better part of two decades disagreeing, expensively, about the most basic question in the field: what a machine should actually use to see. One camp, led for years by Waymo, bet on layering lidar, radar, and cameras together, redundant senses cross-checking each other the way human vision is itself backed up by balance, proprioception, and hearing. The other camp, led publicly by Tesla, bet that cameras alone, processed well enough, should be sufficient, on the argument that humans manage the same roads with two eyes and no laser rangefinder. Both camps have spent billions. Neither has definitively won. A decade after the field converged on the comparatively simple question of how to build a hand, it still hasn’t agreed on the basic architecture of faithful sensing.
Part IV: The Older, Quieter Attempt to Solve This
The strange part is that someone tried to solve the perception half of this problem decades before anyone thought to attach it to a robot, and almost nobody remembers it now.
In 1945, Vannevar Bush, who had run American science policy through the Second World War, published an essay called “As We May Think” imagining a device he called the Memex: a desk that could store and instantly retrieve everything a person read, wrote, and thought, linked by association the way a human mind actually works rather than the rigid folders and indexes libraries used. It was, in effect, the first serious proposal for a second brain, written before the transistor existed.
Fifty years later, a computer scientist named Gordon Bell decided to actually build it, on himself. Starting in 1998, Bell wore a device called a SenseCam around his neck that took a photograph every time its sensors detected a meaningful change in light or motion, alongside a system that archived every document, email, phone call, and piece of correspondence he touched. He called the project MyLifeBits, and by the time he wound it down in 2007, it held more than a million document pages, well over a hundred thousand photographs, and years of continuous ambient record of an actual human life. It was, quite literally, a perception-first bet: capture the raw sensory record faithfully, worry about what to do with it after.
Bell’s own conclusion, stated plainly in later interviews, is the detail that matters most for everything argued here. Capturing the material turned out to be the tractable part. The part nobody had solved, the part that made the archive more curiosity than tool, was turning a decade of raw perceptual record into something a person could actually use, structured, retrievable, alive to what mattered and quiet about what didn’t. Bell had built extraordinary senses and no real judgment to sit behind them. He said later that the smartphone quietly ended the project, because everyone started accidentally lifelogging through their camera roll without meaning to, capture had become ambient and free, while the harder second half, making sense of it, stayed almost exactly as unsolved as it was in 1998. It’s a small footnote of history that Microsoft, the same company that employed Bell for the entire life of that project, announced its own version of ambient capture for the PC within days of his death in 2024.
Part V: What the Body Already Knew
There’s a biological reason this order of difficulty, perception hard, action comparatively easy, shouldn’t actually surprise anyone. Roughly thirty percent of the human cerebral cortex is given over to visual processing alone, more real estate than touch and hearing combined, which get roughly eight and three percent respectively. Some estimates, accounting for the full downstream network that visual signals recruit across the parietal and temporal lobes, put the number closer to half the entire cortical surface.
Think about what that split is actually telling you. Seeing functions as thinking itself, or close enough to it that evolution built roughly a third of our most expensive tissue to do it well, while the deliberate, verbal reasoning we’re inclined to call intelligence runs on comparatively little of what’s left. We built our machines backwards relative to how we ourselves are actually built, cheap perception bolted onto expensive reasoning, when the biological template has always run the other way.
Part VI: The Actual Bet
Put Moravec, Bell, and the visual cortex together and a specific, contrarian position falls out fairly cleanly. The industry’s current wave of enthusiasm for Physical AI is, almost entirely, a bet on actuation: robots that act in the world. That’s a real and valuable problem, though a less differentiated one than the evidence above suggests. The harder problem, the one with a forty-year paper trail of researchers running into it and a quieter, less-funded history of people trying to solve it on their own bodies, is faithful perception: capturing what actually happened, in context, richly enough that something downstream, human or machine, can make good judgment out of it later.
That’s the bet behind the long-term roadmap we’re building toward at Twelfth Brain, stated as plainly as I can put it. It puts senses before hands: ambient and wearable capture that gives the second brain a more faithful record of an actual life to work from, the way Bell tried to build for himself twenty-five years too early, before there was any judgment layer capable of doing something useful with what he’d captured. The judgment is the part the rest of this series has been about. This piece is about the other half: judgment has nothing to compound from if the perception underneath it was never faithful to begin with.
A Few Honest Distinctions
When a demo crosses your feed, the useful question is which half of the problem it actually solved: general perception, or just that one rehearsed warehouse aisle.
It also helps to keep the fundable problem and the important problem as separate questions. Actuation draws the funding and the viral clip because it makes for better video, which is a fact about incentives and says nothing about which problem is actually harder or matters more.
Capture and understanding are different achievements, and Gordon Bell’s own project is the proof. He showed in 1998 that a life could be captured faithfully. He also showed, by his own account years later, how far that capture sat from a usable second brain without a judgment layer built to sit on top of it.
Biology’s own allocation of resources is itself an argument worth taking seriously. Nature had five hundred million years to decide which problem was more prestigious and no ideology pushing it one way or the other. It still spent nearly a third of the cortex on seeing.
Which leaves one practical test for any Physical AI thesis crossing your feed: is it a hands bet or an eyes bet. Both are real bets. Only one of them carries a forty-year paper trail of turning out to be the harder, more foundational problem sitting underneath the other.


