Prologue: The Shorter Version Wasn’t Wrong. It Was Incomplete.
A few weeks ago I wrote a short piece by this title. The core claim in it was simple: AI has made execution cheap, so the founders and companies who keep competing on speed are competing on the one axis that just stopped mattering, and the real scarcity moved upstream, to judgment.
I still believe that. But a claim that big deserves more than a page, and the more time I spent with it, the more I realized I was leaning on intuition where actual research already exists, in cognitive science, in economics, in organizational behavior, going back decades. Some of that research supports the original argument more strongly than I originally made it. Some of it complicates the argument in ways that make it more honest. Both are worth putting on the record.
So this is the longer version. Same thesis. More evidence. And one important correction I owe the shorter piece, which is that not all judgment deserves the reverence I gave it the first time around.

Part I: The Bottleneck Moved
For a decade, the founders who won were the ones who shipped fastest. Execution was the scarce resource, so speed was the moat. AI has quietly dissolved that scarcity. A prototype that took three months five years ago can be built in a weekend now. A first draft of almost anything, code, a financial model, a legal memo, is available to anyone who can write a clear enough prompt.
What AI has not touched, and structurally cannot, is the decision that happens before any of that building starts: whether this is the right market, the right moment, the right problem underneath the one visible on the surface. That decision has a name. It’s judgment. And until recently, I didn’t have a precise enough account of what judgment actually is, mechanically, to defend the claim beyond gut feel. It turns out there’s a well-developed field of research that does.
Part II: What Judgment Actually Is
Start with a Hungarian-British chemist named Michael Polanyi, who spent the second half of his career on a question that turns out to matter enormously right now: how much of what an expert knows can that expert actually say out loud?
His answer, from a 1966 book called The Tacit Dimension, is a single sentence that has been quoted so often it’s practically a proverb in cognitive science: we know more than we can tell. Polanyi’s favorite illustration was recognizing a face. You know your closest friend’s face instantly, unmistakably, in a crowd of a thousand. Now try to describe it precisely enough that a stranger could pick that same face out of a photo lineup using only your words. You can’t, not really. The knowledge is real. It just isn’t stored in a form language can fully retrieve.
Economists have since given this a sharper, more consequential name. In a 2014 paper, MIT’s David Autor called it Polanyi’s Paradox and used it to explain something that had puzzled labor economists for years: why some jobs resist automation even when they look, on paper, entirely rule-based. The answer is that the rules the expert is actually using were never fully articulated in the first place, not because no one tried, but because they can’t be. You can automate what someone can specify. You cannot automate what they only know how to recognize.
The clearest empirical demonstration of this I’ve come across is Gary Klein’s research on fireground commanders, published from a study he ran in the mid-1980s. Klein and his colleagues expected to find veteran firefighters weighing multiple options against each other under pressure, the textbook model of rational decision-making. Instead they found something closer to nothing at all, in the deliberate sense. In fewer than one in eight of the decision points they studied did a commander compare two or more options before acting. The rest of the time, an experienced commander looked at a fire, recognized it as a version of a fire he’d seen before, and simply knew what to do, all before his conscious mind had “decided” anything in the way we normally use that word. Klein called this the recognition-primed decision model, and it has since shaped how the U.S. military trains officers to make decisions under time pressure.
That’s what judgment actually is, mechanically. Not a faster form of analysis. A different cognitive operation entirely: pattern recognition built from thousands of prior instances, compressed into something that arrives as a felt sense rather than a chain of reasoning, and that the person carrying it usually cannot fully explain even when asked directly. Which means it also can’t be extracted by asking. You cannot interview your way to someone’s judgment. You can only capture it as it’s exercised, over time, in context, or you don’t capture it at all.
Part III: The Uncomfortable Caveat
Here’s where the shorter version of this essay was too generous, and I want to correct it directly rather than let it stand.
I wrote about judgment as though all expert intuition deserves trust simply because it’s earned through experience. That’s not true, and the person who proved it most rigorously is Daniel Kahneman, who spent much of his career skeptical of exactly the kind of claim I was making. Kahneman and Klein, despite starting from opposite instincts about intuition, one from behavioral economics and a career studying bias, the other from decades inside high-stakes professions, eventually co-authored a 2009 paper that resolved their disagreement instead of hiding it. Their conclusion: whether an expert’s gut feeling is worth trusting depends entirely on two conditions. First, whether the environment the expert operates in is actually regular enough to learn from, whether the same causes reliably produce the same effects. Second, whether the expert got fast, clear feedback on their decisions long enough to actually learn the pattern, rather than just accumulating years in the room.
A fireground commander meets both conditions. Fire behaves according to physics that don’t change, and a bad call shows itself within minutes. A surgeon largely meets both conditions too. But a huge amount of what gets called “expertise” in business, in politics, in punditry, meets neither. Philip Tetlock spent twenty years tracking roughly 82,000 predictions from 284 professional forecasters, people paid to know what would happen next in politics and economics, and found their accuracy was statistically indistinguishable from chance. Chimps throwing darts at a board of outcomes would have done about as well. These were not unqualified people. Most held postgraduate degrees. What they lacked was the thing Kahneman and Klein identified: a stable environment and fast, honest feedback. Confidence, it turns out, is not evidence of validity. It’s just confidence.
This matters enormously for anything trying to preserve or transmit someone’s judgment, including the thesis I’m building toward here. Not every senior executive’s gut call deserves institutional memory. Some of it is the fireground commander. Some of it is the dart-throwing chimp with a job title. The honest version of this project has to be able to tell the difference, which mostly comes down to asking: did this person operate in a domain with real, fast feedback, or did they operate in a domain where being confidently wrong for years carries no visible cost? The AI era makes this distinction more urgent, not less, because a system that captures a lifetime of pattern-matching without asking whether the patterns were ever actually tested against reality just automates false confidence at scale, faster than any human ever could on their own.
Part IV: What Gets Lost, and What It Costs
None of this would matter much if judgment, once earned, simply lasted. It doesn’t, and the decay starts almost immediately.
Hermann Ebbinghaus demonstrated this in the 1880s with a simplicity that still holds up: without reinforcement, people forget roughly half of what they’ve just learned within about an hour, and the majority of it within a few days. Ebbinghaus was testing memorized syllables, not judgment, and it’s worth being precise about that difference. Judgment is arguably more fragile than the memories he measured, not less, because at least a memorized fact can be written down and crammed back in. Nobody has ever crammed back a decade of pattern recognition from a single written page, because it was never written down as a proposition to begin with. It only ever existed as Polanyi’s tacit knowledge, inside one person, and it leaves when they do.
Civilizations solved this problem for facts a long time ago. The scriptoria copied texts by hand so a fact discovered once could outlive its discoverer. The printing press then multiplied that survival by orders of magnitude. What neither of those inventions, nor anything built since, ever solved is the transmission of judgment itself, the surgeon’s diagnostic instinct, the operator’s read on a market, the founder’s sense of which fire needs to be fought from inside the house. Every generation of experts has started closer to zero than it should have, not for lack of talent, but because the infrastructure to hold what the previous generation actually knew, as opposed to what they merely recorded, was never built.
This isn’t abstract, and it isn’t cheap. Deloitte estimated in 2024 that institutional knowledge loss costs U.S. companies roughly $1.3 trillion a year. Gartner puts the proportion of enterprise knowledge that is tacit, meaning it exists only in someone’s head and was never written down anywhere retrievable, at somewhere between 70 and 80 percent. Harvard Business Review has documented cases where a single wave of roughly 700 retirements at one organization represented the loss of more than 27,000 combined years of experience, gone in a single transition period. And that transition is only accelerating: roughly 10,000 Americans turn 65 every day through the end of this decade, each one carrying a version of this same problem out the door with them.
We built magnificent infrastructure for facts. We built essentially none for judgment. The AI era is the first moment in history where that second infrastructure becomes technically possible to build at all.
Part V: For the First Time, Infrastructure Is Possible
Everything above is, in a sense, the argument for why this problem has always existed and why it has always been unsolved. What’s actually new is not the problem. It’s that the tools to address it, tools that can sit alongside someone as they work and capture judgment as it’s exercised rather than trying to extract it after the fact through an interview neither party can fully trust, did not exist until very recently. That’s the bet behind the institution we’re building at Twelfth Brain: not a faster way to write things down, but infrastructure for the part that was never written down in the first place, captured in context, tested where possible against the Kahneman-Klein conditions above, and kept compounding long after the person who built it has moved on.
A Few Honest Distinctions
Before treating anyone’s experience, including your own, as judgment worth preserving or trusting, these are the questions the research above actually supports asking.
First: did this person operate in an environment with real, fast feedback? A fireground commander gets it in minutes. A pundit often never gets it at all. Match your trust to the domain, not to the years on the resume.
Second: can they demonstrate the pattern, not just narrate the conclusion? Genuine tacit knowledge, per Polanyi, usually resists being turned directly into a tidy explanation. Someone who can only offer smooth, confident narrative and no actual track record of being checked against reality should raise your suspicion, not lower it.
Third: is the confidence proportional to the domain’s actual predictability? Tetlock’s forecasters were plenty confident. Confidence and calibration are different things, and only one of them is evidence.
Fourth: would writing it down as a rule capture it, or does it only exist as recognition? If it’s the second, no checklist will ever hold it, and the only real capture method is context over time, not a single debrief.
Fifth: what happens to this specific judgment when this specific person is gone? If the honest answer is nothing survives, that’s not a hypothetical. It’s the default outcome for almost everyone, almost all the time, right now, today.
Further Reading
On tacit knowledge and why some expertise resists automation
Michael Polanyi’s The Tacit Dimension (1966) is the source of “we know more than we can tell.” David Autor’s 2014 paper naming Polanyi’s Paradox connects the idea directly to the economics of automation.
On how expert judgment actually operates under pressure
Gary Klein’s Sources of Power lays out the recognition-primed decision model from his original fireground commander research. Daniel Kahneman and Gary Klein’s 2009 paper, “Conditions for Intuitive Expertise: A Failure to Disagree,” is the single most useful thing I’ve read on when to trust a gut call and when not to.
On what happens when expertise isn’t actually earned in a valid environment
Philip Tetlock’s Expert Political Judgment is the twenty-year study behind the “dart-throwing chimps” finding. His later book, Superforecasting, co-written with Dan Gardner, is the more optimistic follow-up on what actually separates good forecasters from confident ones.
On what institutional knowledge loss actually costs
Deloitte’s 2024 research on institutional knowledge loss and Gartner’s enterprise knowledge management research are the two most-cited sources behind the figures above, and worth reading in full if the scale of the number matters to your own case.

