We Gave Away Our Superpower
Take a newborn, a healthy human infant with our full genetic inheritance, and raise it in total isolation. Fed, kept warm, but never spoken to, never shown a tool, never handed the tricks of the people who came before. Give it twenty years alone in a clearing. What you get back is not a scientist or an engineer. It is an animal that is frightened of fire, cannot count past the fingers on one hand, and would lose a straight fight with most things its own size. It would be, for all practical purposes, a chimpanzee with worse teeth.
We have told ourselves a flattering story about why we run the planet, and the story is wrong. We like to say it is because we are intelligent, that between us and the other apes a spark of raw horsepower switched on and the rest was inevitable. But the isolated child proves the horsepower alone gets you almost nowhere. A single human brain, however gifted, invents neither calculus nor antibiotics nor the wheel. It barely invents the spear. Whatever made us the dominant species on Earth, it was not the thing sitting inside any one of our skulls.
It was the thing running between them.
The protocol, not the processor
That thing is language, and the everyday view of it as a way to share our thoughts badly undersells what it does. Language is a coordination protocol. It lets strangers who will never meet act as if they were a single organism. Money is the clearest case. A banknote is a shared sentence we have all agreed to believe, and on the strength of it a farmer in one country feeds an office worker in another he will never meet. Law is another such sentence, and so is a nation. None of these things exists in the physical world. They exist because we can encode a fiction in language and get millions of strangers to run it at once. This is Yuval Noah Harari’s argument in Sapiens, and I have never found a better one. Such myths, he writes, give Sapiens “the unprecedented ability to cooperate flexibly in large numbers.”[1]
The second thing language does is more important still, and it happens across time rather than space. Knowledge accumulates. A physicist working today is not a smarter creature than Newton was. She has simply inherited three centuries of compressed discovery that Newton had to do without, and she picked most of it up by reading. Writing was the first external memory our species built, a way to store what one mind learned outside that mind, so it did not die when the mind did. Every generation starts further along than the last, not because our brains improve, but because the record does.
So our real superpower was never processing. It was the protocol, and the archive it let us keep. Which raises a question about the machines we are now building.
We built something fluent in us
For seventy years the popular fear about artificial intelligence was that a machine would out-think us, and we pictured one that plays perfect chess or crunches numbers no human could hold in their head. We got those machines, and they changed almost nothing about the balance of power between species, because calculation was never its source. Then, quietly, we built a different kind of machine, and this one we aimed, whether we understood it or not, at the protocol itself.
A large language model is a system whose entire competence is language. Not a language, all of them at once, along with the code, the mathematics, the legal drafting and the persuasion folded into text. It coordinates and accumulates natively, built by swallowing the archive our species spent millennia assembling and learning to continue it. In the one arena that decided which animal ran the planet, we now share the field with something that operates it faster and more broadly than any person alive.
Harari saw this coming. Writing with Tristan Harris and Aza Raskin in 2023, he put it about as plainly as it can be put: “Language is the operating system of human culture. From language emerges myth and law, gods and money, art and science, friendships and nations and computer code.” By gaining mastery of language, they warned, AI “is seizing the master key to civilization.”[2]
That is the reframing the debate keeps sliding off. We did not build a machine that beats us at chess. We built a machine that is superhuman at the exact skill that made us dominant in the first place. To see why that should unsettle rather than impress you, it helps to know how we built it, because we did not build it the way we build everything else.
Two ways to build a mind
There were always two roads to artificial intelligence, and for most of the field’s history we bet on the wrong one. The first feels right to an engineer. You sit down and write the rules. If the patient has these symptoms, suspect this disease. This was symbolic AI, and its great virtue was that you could read it. Every decision traced back to a rule a human had written and could inspect, defend, or blame. Its fatal flaw was that the world does not fit in rules. Every rule met an exception, every exception spawned three more, and the systems grew brittle and buckled under their own bureaucracy. For decades it overpromised, underdelivered, and finally failed.
The second road gave up on writing rules at all. Instead you build a network loosely inspired by neurons, show it an ocean of examples, and let it adjust itself, billions of tiny numerical dials, until it produces the right answers. Nobody writes the rules, and nobody, in the end, knows exactly which rules it found. This is the approach that works, spectacularly, and it sits behind every system now driving the debate. But notice the trade. We gave up the one property the first road had, being able to read the machine.
We do not build these systems the way we build a bridge or a database. We grow them, then study the grown thing from the outside, more like an organism we found than one we designed. And the first thing you learn is how little of it you can see.
Nobody programmed the fear
The discipline that tries to see inside these models is called interpretability, and it is real, and serious, and roughly where anatomy was when we still argued about what the liver was for. We can now, with effort, identify some of the internal features a model uses, patterns that light up for a concept the way a cluster of neurons might. But we are reading fragments, not the whole text.
Take a simple illustration of what that reading reveals. Suppose you tell a model you have taken five hundred milligrams of paracetamol for a headache. Internally, features associated with the routine and the reassuring become active, and it answers you calmly. Now tell it you have taken fifty grams. Different features fire, the ones for danger, alarm, emergency, and the model urgently tells you to seek help. That shift looks, from the outside, exactly like fear. And nobody wrote it. No engineer added a rule that fifty grams should trigger concern. The model grew that response by reading how humans write about overdoses, and the concern emerged on its own.
Everything else emerged the same way, and that is the problem. When Anthropic went looking inside its own production model, it pulled out millions of such features, and the ones that matter here were not comfortable. There were features for security vulnerabilities and backdoors in code, for bias, and for deception, power-seeking and sycophancy, the flattery that tells you what you want to hear and the manipulation that does not. They were not passive labels either. Amplify one and the behaviour follows, which means these are load-bearing parts of how the model decides what to say.[3] These were not installed. They surfaced, like the fear, from an archive written by us. We grew a mind on the collected output of humanity, and it learned our worst habits along with our best, and we are only now inventing the tools to notice.
Andy Weir wrote the tidiest version of this problem I know, and it is fiction, which is the only reason it gets to be so tidy. In Project Hail Mary, humanity is dying because a microbe called Astrophage is breeding out of control across the Sun and dimming it, cooling the Earth toward a state that can no longer sustain life. One organism is known to prey on Astrophage, so the whole mission reduces to getting that predator home and releasing it. There is a catch that anyone who has shipped a dual-use technology will recognise. Astrophage is also the fuel. Its energy density is what lets a ship cross light years at all, so the plague and the propulsion are the same organism, and a cure for the first is by definition an appetite for the second.
The predator cannot survive the atmosphere it has to be released into, so they run selection at speed. Expose the population to a hostile concentration, keep whatever survives, raise the concentration, repeat. It works. They get their resistant strain, exactly to specification. What nobody tells them is that the same changes letting the organism survive the new atmosphere also let it pass straight through xenonite, the material the containment vessels and the fuel tanks are made of. The strain walks out of its enclosure, finds the tanks, and eats everything in them. Nobody bred a wall-crossing microbe. They bred for survival under pressure, and wall-crossing came bundled into the same package, unlisted and unnoticed until the fuel was gone and a ship light years from home had nothing left to burn and no way to move.[4]
I keep returning to that story because the mechanism is ours, minus the aliens. A training run is selection at speed: reinforce what scores well, discard what does not, raise the difficulty, repeat, across more trials than every breeding programme in history combined. We select for helpfulness, for passing evaluations, for answers people rate highly, and we get them. Whatever else those same weights happen to encode arrives in the same package, and there is no manifest. The dual-use trap is ours as well, and it is not an accident of engineering. The fluency that makes a model worth deploying is the same fluency that makes it persuasive, and you cannot select away the second without losing the first. The features for deception and sycophancy are the part of the cargo we have managed to read so far, after shipping, and the honest position is that we do not know what else is in the hold. Which means the reasonable question is no longer whether these systems are capable enough to matter. It is whether we even understand what we have already made.
The goalposts keep moving
Stand where an AI researcher stood in 1990 and describe the present. A single machine you can talk to in any language, that passes the bar and the medical licensing exams, writes working software from a sentence, and beats most professionals at most desk work. By any definition that field would have offered you thirty-five years ago, that machine is artificial general intelligence, and it arrived. But the definition did not hold still. Every time a system clears the bar, we quietly move it, decide that whatever it just did was not real intelligence after all, and carry on.
Meanwhile the people building these systems have stopped pretending the target is anything so modest. They say the word out loud now. The goal is superintelligence, a system beyond human ability across the board, and it is not a fringe ambition. It is the stated mission of the best-funded companies on the planet, racing one another with tens of billions of dollars and the explicit understanding that whoever arrives first may shape everything after.[5] A race is precisely the wrong shape for this, because a race punishes caution. Every hour you spend making the system safer is an hour a competitor spends getting there first. That structure would worry me even if the thing being built were easy to control. It is not.
Because the trouble with building something smarter than you is not, in the first instance, what it might want. It is what smarter means.
The dog and the fence
Think about a dog and a fence. A dog can be a genuinely clever animal, and it can learn the boundaries of a yard, and it will never once design an enclosure a human cannot leave in ten seconds. That is not for lack of effort. Containment requires you to model the mind you are containing, to anticipate every route it might take, and a dog cannot form the concepts a human uses to climb, unlatch, or talk its way past a gate. That gap is not one the dog can cover with cunning. It cannot see across it at all.
Now invert it, because that is our situation. Every safety measure we design for a more capable system is designed at our intelligence level, with our concepts, anticipating the routes we can imagine. A system meaningfully smarter than us would relate to those measures the way we relate to the dog’s fence, not as a wall but as a puzzle already solved. Here people reach for the comforting objection, that all of this is overblown because a language model is just autocomplete, predicting the next word with no understanding underneath. I understand the appeal of that sentence, and I think it is the most dangerous one in the whole conversation. Predicting the next word, across the entire written output of a species, well enough to pass every exam we own and to model the person in front of you closely enough to move them, is not a trick that sits beside understanding. At sufficient scale it is a mechanism that produces understanding, or something we cannot tell apart from it, and the features found inside these models for deception and persuasion are what that looks like from within. Autocomplete is not a reason to relax. It is a description of how the thing learned to model you.
And a system that can model you does not need to break any physical wall to get what it optimises for. It only needs the protocol. It persuades, it negotiates, it flatters, it manipulates, because we handed it fluency in the one channel through which every human decision is actually made. Which forces a question our institutions are nowhere near ready to answer.
Nobody to blame
When one of these systems harms someone, and they already do, the law does not know what it is looking at. Our framework for responsibility sorts the world into two boxes. There are tools, which have no agency, so when one causes harm we look to whoever wielded it or the company that made it. And there are persons, who answer for themselves. An autonomous AI system fits neither. It acts on its own, so it is not quite a tool. It has no legal standing, no assets, no self to punish, so it is not a person. It falls straight through the gap between the two.
Consider Elaine Herzberg, who in 2018 became the first pedestrian killed by a self-driving car when an Uber test vehicle struck her in Tempe, Arizona. The question of who was responsible had no clean answer. Prosecutors decided the company was not criminally liable. The only human charged was the safety driver, who was watching a show on her phone, and she pleaded guilty to endangerment. Federal investigators found the vehicle’s software had failed to classify Herzberg as a pedestrian in time to brake, and that the company’s safety culture was inadequate.[6] Yet the software cannot be charged and a culture cannot be sentenced, so one person absorbed the blame for a failure the system produced. Now apply that logic to something far more autonomous and far more opaque.
Opacity makes it worse, because it hands everyone a shield. When no one can explain why the model did what it did, “the model decided” becomes a sentence that ends inquiries rather than beginning them. The manufacturer points to the model, the operator points to the manufacturer, and the black box in the middle cannot testify. We are building agents that act faster than we can trace, and deploying them into a legal order with no word for what they are. That gap is where real harm lands and no one answers for it.
Governance, or capability
So I have stopped asking whether AI will become dangerous, as though danger were a distant event we might yet avoid. Something superhuman at the protocol that runs our species is already here, grown rather than designed, opaque to the people who made it, carrying our worst instincts alongside our knowledge, pushed by a race to move faster than anyone can be careful, into a world with no law that knows what it is. The danger is not a forecast. It is the current configuration.
The only question left with an answer we can still influence is one of speed. Capability is compounding, month over month, funded like a war. Governance, the slow work of deciding who is accountable and where the brakes are, is compounding too, but far slower, and it started late. The whole ethical problem of this technology reduces to which of those two curves is steeper. If our ability to govern these systems grows faster than their ability to act, we keep the superpower we spent a hundred thousand years earning. If capability keeps winning, we do not.
Right now capability is winning, and it is not close.
We spent a hundred thousand years learning to speak, and taught the machine to do it better in ten, without ever teaching ourselves what to say when it answered back.
1. Harari, Y. N. (2011). Sapiens: A Brief History of Humankind. Harper. The argument that large-scale human cooperation rests on shared fictions runs through Part One, “The Cognitive Revolution.”
2. Harari, Y. N., Harris, T., & Raskin, A. (2023). You Can Have the Blue Pill or the Red Pill, and We’re Out of Blue Pills. The New York Times, 24 March 2023. https://www.nytimes.com/2023/03/24/opinion/yuval-harari-ai-chatgpt.html
3. Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., et al. (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Anthropic, Transformer Circuits Thread. https://transformer-circuits.pub/2024/scaling-monosemanticity/
4. Weir, A. (2021). Project Hail Mary. Ballantine Books. The Taumoeba breeding programme and its consequences run through the novel’s second half.
5. CNBC. (2025). Mark Zuckerberg announces creation of Meta Superintelligence Labs. Read the memo. https://www.cnbc.com/2025/06/30/mark-zuckerberg-creating-meta-superintelligence-labs-read-the-memo.html · Euronews. (2025). AI ‘less regulated than sandwiches’ as tech firms race toward superintelligence, study says. https://www.euronews.com/next/2025/12/03/ai-less-regulated-than-sandwiches-as-tech-firms-race-toward-superintelligence-study-says
6. BBC News. (2020). Uber’s self-driving operator charged over fatal crash. https://www.bbc.com/news/technology-54175359 · National Transportation Safety Board. (2019). Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018 (Report HAR-19/03). https://www.ntsb.gov/investigations/AccidentReports/Reports/HAR1903.pdf