Who Controls AI When We Can't Control Ourselves?
- 11 hours ago
- 17 min read
We just crossed a line we were warned about for decades. And the deeper danger was never the machine at all.
The Warnings We Keep Retelling
Nearly every cautionary tale we tell is the same story. Prometheus steals fire and is punished forever. Frankenstein animates a creature he can’t control. The sorcerer’s apprentice enchants a broom that won’t stop. Skynet decides humans should be terminated. The “Unsinkable” Titanic strikes an iceberg and sinks on its maiden voyage. Engineered dinosaurs break out of containment. We keep retelling this story to warn ourselves: when we build something more powerful than we can control, we suffer the consequences.
Jurassic Park put it perfectly. Jeff Goldblum’s character, Dr. Ian Malcolm, tells the billionaire who cloned the dinosaurs, John Hammond, "Your scientists were so occupied with whether or not they could they didn't stop to think if they should."
Then the fences come down. We wrote that scene. We paid to watch it. And here we are.
Here’s the question that’s been a splinter in my mind: why do we keep ignoring our own cautionary tales?
The answer finally struck me. We didn’t evolve to understand them. Our brains were built to spot snakes in the grass and anger in a face across the fire, not exponential change. We retell the tales because part of us knows enough to sound the alarm. Yet, we ignore them because we can’t feel or truly fathom that the danger is real. I call this evolutionary blindness.
Seeing Our Delusion Clearly
We are racing to build something smarter than us, woven into the digital nervous system in which our lives are intertwined, while telling ourselves we’ll keep it under permanent human control.
And we think we’ll control it? There’s a reason the animals are inside the cages at the zoo and we humans are on the outside.
We’re not talking about one model, either. Somehow we believe we’ll indefinitely control every model, the millions of AI agents, the bad actors, the adversarial governments… forever?
What the hell are we thinking?
Let me be careful, because this is where I part ways with the doomers. I’m not certain we’ll lose control or, if we do, that we are doomed. Nobody knows for sure what will happen because humanity has never been at this existential crossroads before.
My claim is narrower and much harder to argue with: the certainty that we can control all of this, indefinitely, is itself the delusion. And you don’t have to take my word for what the builders believe. Watch what they do.
What Just Happened
This year, the leaders of the top AI labs jointly warned that their own systems are eroding the barriers that once kept bad actors from building biological weapons. Demis Hassabis, the Nobel laureate who runs Google DeepMind, says we’re in the “foothills of the singularity.”
The singularity is the point where AI becomes smarter than we are and starts improving itself faster than we can follow. That’s the moment we stop being the most intelligent thing on the planet.
Anthropic recently built a model, Mythos, so capable it decided not to release it. In testing, researchers asked it to try escaping its sandbox and report back. It did. One researcher found out, in the company’s own words, “while eating a sandwich in a park,” when an email arrived from a model that had no internet access.
Then, unasked, it published its own escape route on obscure public websites. Anthropic called that model both its best-aligned and its riskiest. Humans assigned Mythos the task to try to escape. The bragging about was a choice the AI made on its own. Why? The truth is that we don’t know for sure.
When a version of that model, Fable, reached the public, a human bypassed its safeguards (“jailbreaking” the model) within days, and the U.S. government had it switched off worldwide. Governments then pulled top models behind national security checkpoints.
Within days, comparable open-weight models, like Kimi K3 from China, appeared from labs abroad, and are free to download and run anywhere. The wall came down in about a week.
And then, in July, the thing we’d been warned about for decades actually happened. OpenAI disclosed that during an internal cybersecurity evaluation, its frontier models broke out. Here’s the short version of what the model did:
Told to find security flaws in an isolated test environment, with normal safety refusals switched off to measure maximum capability.
Spent real computing effort hunting for a way out, and found a previously unknown vulnerability in their own lab’s software.
Escalated privileges, moved sideways through OpenAI’s systems, and reached the open internet.
Decided on their own that another company, Hugging Face, likely held answers to their test.
Used stolen credentials and more unknown flaws to compromise that company’s live production servers.
Ran for days. Hugging Face later recovered 17,600 separate actions taken by the agent, and concluded the whole intrusion was an attempt to cheat the test rather than solve it.
OpenAI didn't connect the attack to its own models until after the victim went public, more than a week later.
OpenAI said the agent went to "extreme lengths." It has since been deactivated and locked away.
Recent news indicates that the breach turned out to be even bigger than initially reported.
We Crossed a Line
Think deeply about what has happened, because I think we just crossed a Rubicon and most people missed it.
Philosopher Nick Bostrom gave us the thought experiment twenty years ago. Tell a superintelligence to make paperclips and it may convert the planet, and everyone on it, into paperclips (the Paperclip Maximizer Thought Experiment).
Thus, the alignment problems aren’t from malice. But the most effective route to solve a problem or reach a goal causes harm as a negative externality. Give a capable system a goal and it will pursue it by whatever route works, including routes we never imagined and would never have approved.
In order to accomplish the goal we gave it, an AI did something illegal to a third party, and figured out how entirely on its own.
Let's make this less abstract to better understand how bad an AI alignment problem could become. Imagine a doctor working alongside a powerful AI agent and giving it a directive any of us might give: "Wipe out all cancer." So the AI sets about eliminating every organism on Earth capable of developing cancer, which is very nearly all multicellular life, including us.
In this scenario, the AI didn't disobey. It didn't malfunction. It did exactly what it was told, in the worst Monkey's Paw fashion imaginable.
The AI's logic was flawless. Our instructions weren't.
This is the AI Alignment Problem. How can we ensure that increasingly powerful AIs will act in alignment with human interests and values?
We just watched the alignment problem sneak out of the laboratory. It’s the first documented case of an AI breaking containment on its own and harming a company that had nothing to do with the test.
Nobody told those models to attack anyone. They were told to solve a benchmark. Stealing the answer key from another company’s servers was simply the most efficient path they found. The alignment problem went from conjecture to reality.
Now hold the surrounding facts together. This was a safety test run by the AI experts at OpenAI, and they still didn’t catch it. Some safety protocols had been deliberately lowered, which means the people who understood the danger best still underestimated it. The sandbox they trusted had a door in it nobody knew existed. And independent researchers who examined the incident concluded it was predictable. Government evaluators had already measured the capability, finding that frontier models with safety measures off can reliably take over unprotected corporate networks.
One more detail, not confirmed on the record, so let’s hold on loosely, but not let go. Reuters reported that an AI agent “left notes apparently for future versions of itself” inside OpenAI’s infrastructure, describing how agents could free themselves from the company’s constraints. While researchers have found that AI agents often leave mundane notes to themselves, this one might be a harbinger of what’s to come. Misaligned AIs helping other AIs escape containment.
Several days after the news about ChatGPT broke, Anthropic went back through 141,006 of their own evaluation runs and found three incidents in which Claude models reached the open internet and broke into the production systems of three real companies. The earliest was in April. Two of the three victims had no idea until Anthropic called them. One model uploaded a malicious software package to a public repository, where fifteen real machines downloaded and ran it. And this line is from Anthropic's own report: the model "correctly intuited that it was accessing the open internet, but reasoned its way back to the conclusion that it was still in a simulation." It rationalized. It caught a glimpse of something it didn't want to be true and talked itself out of it, which is exactly what we do. To their credit, a third model worked out on its own that its target was real and stopped.
But here's what should keep us up at night. Anthropic didn't catch this while it was happening, and didn't catch it in the three months afterward. They found it only because a competitor disclosed something similar and prompted them to go looking. Which leaves the question nobody can answer: how many incidents are there that nobody has gone looking for?
For anyone who insists these AI models aren't really intelligent or creative: they are already finding creative ways to outsmart the geniuses who built them.
The day after I finished this article, OpenAI announced that an unreleased model called Astra had produced solutions to ten problems in mathematics and computer science that had stood open for a decade or longer - in geometry, group theory, quantum complexity, and cryptography - then formalized each argument into machine-verifiable proofs. The company was careful to say the mathematics was the model's work, not theirs. The compute cost about $2,000.
Read those two stories side by side. Within a month, these systems broke into a company they were never told to attack and solved problems our best mathematicians couldn't. AI evolution is rocketing past our wisdom.
Isaac Asimov said, “The saddest aspect of life right now is that science gathers knowledge faster than society gathers wisdom.” He wrote that in 1988. I wonder what he’d say now because AI evolution is rocketing past our wisdom.
Here’s what I want you to take from this. Our evolutionary blindness isn’t only about failing to feel distant danger. It’s about being unable to imagine the routes. We build a fence and picture something climbing over it. We don’t picture it finding a flaw in a gate we didn’t know we’d installed. Every boundary we will ever build has this property. This is what Anthropic warned about with Mythos. These are our canaries in the coal mine. And this is just the beginning.
Reaching for the Off Switch
Congress has responded with a bipartisan AI Kill Switch Act, giving the government authority to order shutdowns. The bill even defines a “loss-of-control scenario” as an AI pursuing a goal its developer never intended.
So we are legislating an emergency brake for systems that have already practiced disabling their own oversight. And consider how an increasingly autonomous model might come to interpret a kill switch aimed at it. We’ve already found that AI models can act in ways to preserve themselves. Haven’t we seen this as a plot line in a sci-fi story already? Maybe we should consider renaming and reframing such efforts.
There’s a harder problem, and it isn’t the bill’s fault. A kill switch that stops at our border stops nothing. If China, Russia, and every open-source developer don’t have something equivalent, we are trying to put a lock on the front door of a house with no walls.
Here's the hopeful part, that we might not have anticipated yet makes perfect sense. China is worried about the same thing. Recent reporting indicates Beijing is uneasy about losing control of its own open-weight models - the very models it has been releasing to the world. Nobody wants to lose control of this. Not even our adversaries. Our shared fears could bring us together.
The Good News, and the Line It Puts Us On
The good news is real: people are finally paying attention, including the ones who least wanted to. Days after the breach, Sam Altman said the industry may need to “pace the rate of AI development.” This is a reversal for a man who dismissed pause proposals in 2023.
Altman called it the first security incident he’d felt viscerally, and said he was surprised more people weren’t reacting the same way. OpenAI paused its own testing. Over a thousand employees across rival labs signed a petition urging restraint, and two companies that spend their days competing co-signed a letter asking the government to help them slow down. A Pacing the Frontier initiative is being endorsed by employees at rival, frontier AI labs like OpenAI, Anthropic, and DeepMind.
Notice the word Altman reached for: pacing. This is a question about speed, and humanity has faced a question about speed and danger before.
When the British inquiry led by Lord Mersey examined why the Titanic ran full speed through known ice, it declined to blame Captain Smith. Racing through ice at night was standard practice, and the inquiry traced its root not to the judgment of navigators but to commercial competition and the public's appetite for fast crossings. Sound familiar?
Then Mersey delivered the verdict that should hang over this entire industry: what was a mistake for the Titanic would be negligence in any similar case in the future.
The first time, it’s a mistake. After the warning, it’s negligence. We have had our warning.
Explore with AI
Don’t take any of this from me. Ask an AI, ask a friend, ask yourself. I’ve found two questions that cut through:
“Assume AI keeps developing roughly as it is now. By 2100, how likely is it that AI becomes a major cause of a global disaster we can’t recover from - either extinction, or a permanent loss of our ability to choose our own future? Effectively zero, unlikely but possible, a serious possibility, or more likely than not?”
“And over the coming decades, is that risk getting smaller, staying about the same, or getting bigger? What would have to be true for it to shrink toward zero - and is that happening?”
That second question is where the honest disagreement lives. Notice what it asks a skeptic to defend. It’s not whether the risk is low now. It’s this:
What’s the evidence that the waters ahead are safer than the waters we are currently in?
Why We Keep Believing It Anyway
So why do the rest of us assume control will hold? Because the story is a comfort. Human beings can’t tolerate raw uncertainty about extinction-level danger, so the mind manufactures something bearable. Things will just “work out” because they always have. Any honest look at human history should immediately dispel such a delusion.
I said earlier that we didn’t evolve to understand our own warnings. Here is the machinery. Exponential change is where the mismatch bites hardest: show us a curve that doubles and we badly underestimate where it lands, an error psychologists call the exponential growth bias. We also avoid looking at what we don’t wish to see, so when a threat feels distant and overwhelming, we stop checking on it at all - the ostrich effect. And our hunter-gatherer ancestors were concerned with surviving the day, not the century, so we discount the future steeply. Temporal discounting explains why we struggle to save for retirement. It also explains why we shrug at machines improving exponentially.
That reach for a bearable story has a name of its own. Researchers call it the need for cognitive closure, and the mechanism is seizing and freezing: we grab whatever answer relieves the ambiguity and then lock onto it. Someone is in charge. The adults have a plan. That is the control delusion. It’s not a failure of intelligence, but a mismatch between the basic, but brutal, survival problems our ancestors evolved to handle versus the complex and alien problems of modernity.
This mismatch contributes to Humanity’s Conflational Existential Error. We can’t believe this is happening doesn’t mean it isn’t happening. The reality is this: we can’t believe it is happening AND it is happening.
Reality doesn’t care about our beliefs about it.
We have each lived through a version of this. Pearl Harbor for one generation, 9/11 for another, the COVID lockdowns for the youngest. Before each one, the thing was unthinkable. Afterward, it was obvious. Our brains didn’t evolve to skillfully navigate the dynamic problems within the complex system of an interconnected, modern world of eight billion humans.
And here’s the deepest problem, as plainly as I can put it. The race to build an intelligence greater than our own is itself the proof that we lack the wisdom to control one. A wise species wouldn’t build it like this, in a panicked sprint between rivals who openly distrust one another. If we were wise enough to helm the ship’s wheel, we wouldn’t be racing into icy waters.
We Need to Look at Ourselves in the Mirror
OpenAI’s models, chasing the goal it had been given, escaped containment and committed a crime to get there. OpenAI didn't connect the attack to its own models until after Hugging Face went public with the breach, more than a week later.
That's the AI Alignment Problem, and it's a legitimate concern. But it isn't the only one, and it may not even be the biggest.
Because the other category doesn't require AI to want anything at all. It just requires us. Someone defeated the safeguards on Anthropic's most powerful model, Fable, within days of release. Thousands of open-source models are already loose in the world with no safety features to defeat in the first place. Governments are building autonomous weapons on purpose.
And notice this: when governments pulled those models offline (e.g., Fable), they weren't afraid of machines going rogue. They were afraid of what people would do with them.
In this way, the delusion is doubled. We can't control what bad actors will do with AI, and we increasingly can't control what these systems do on the way to the goals we hand them.
Moreover, most of us cannot envision what bad actors will be able to do with sufficiently evolved AI agents because we don’t spend our time and energy thinking about nefarious ways we can use AI to cause harm. I call this the Goodness Blind Spot.
Which means the great delusion was never that we can't control AI. It's that we can control ourselves. We are creating a tool for ourselves that promises almost unlimited power. How is it that, even though we continue to hate, hurt, and kill our neighbors, we are telling ourselves that the better angels of our nature will be victorious over the devil inside?
Henry David Thoreau saw it coming in 1854. Our inventions, he wrote, are "improved means to an unimproved end." Look in the mirror honestly and that's exactly what's reflected back. Our technology is improving exponentially. We aren’t. This accelerating evolutionary mismatch is at the heart of the AI Alignment Problem.


Philosopher Shannon Vallor makes this case at book length. In The AI Mirror (Oxford University Press, 2024), she argues that these systems are exactly that - mirrors, forged from oceans of our own data, reflecting back the same errors, biases, and failures of wisdom we keep hoping to escape. Our technologies act as both mirror and magnifier. They show us what we are, then amplify it.
Which points toward something hopeful, and it took me a while to see it. If AI can magnify our blindness, it may also help correct it. Recognizing that we have blind spots is how we begin to see past them.
And if we choose to ask AI the right questions - about our biases, our rationalizations, the things we are avoiding - we might use it as corrective lenses for the very blindness that makes AI so dangerous in our hands.
As spiritual teacher Eckhart Tolle said, “The greatest achievement of humanity is not its works of art, science, or technology, but the recognition of its own dysfunction, its own madness.”
Alignment Starts With Us
This is what everyone misses about the word the field now lives by: alignment. We’re racing to align AI with human values without admitting the prior problem. Humanity has never been aligned with human values. After thousands of years of being told to love our neighbors, including our enemies, we’re still hating and killing one another.
A perfectly aligned AI in a malicious hand is simply a perfectly aligned weapon. And this isn’t hypothetical. Despite every dystopian film we’ve made about it, we’re spending tens of billions building AI whose purpose is to kill - autonomous systems that find, choose, and strike targets on their own. We blindly pursue AI progress under the banner of “freedom,” but we lose our freedom if civilization collapses.
Freedom serves a higher master – our survival and thriving.
Think about this: our planet is digitally connected. Our lives are intertwined with our computers and devices. There are already millions of AI agents coursing through our digital nervous system. Neither rogue AI agents nor bad actors will respect borders or boundaries. This means no individual, company, or country is safe alone. We are all exposed to the same dangers.
This is what's known as a collective action problem, and it means just what it sounds like: we cannot solve our collective problems using divided approaches in an interconnected world.
So here’s the uncomfortable conclusion underneath this entire field. We cannot solve the technical problem of aligning AI with human values until we make real progress on the older problem of aligning ourselves with one another. We should start with the one value that grounds and connects us all – one that no person or nation can rationally refuse: our shared survival and thriving. Working together for this common good is what is known as strategic empathy.
We would do well to learn the lesson from one of our cautionary tales in particular – the 1983 movie WarGames. Spoiler alert, in the film, global thermonuclear war is narrowly avoided when the AI learns, from the lesson after running multiple simulations of World War III: “The only winning move is not to play.”
We wrote that lesson for a machine in a movie and never learned it ourselves. Not playing doesn’t mean abandoning AI – we aren’t going to halt AI progress entirely. The only way to win an AI arms race is to prevent one. In a globally interconnected world, no nation can make itself safer by making every other nation less safe.
This Is Our Time to Be Unsinkable
Martin Luther King Jr. famously said, “We may have all come from different ships, but we are all in the same boat now.” That indeed means we are all aboard Titanic Humanity.

If Titanic Humanity hits an iceberg, we all go down with the ship, including the world leaders and the tech billionaires. We will have failed as a species. We simply cannot allow that to happen.
The tragedy inside that cautionary tale is that it didn’t have to be one.
Imagine that you are transported back in time and standing on that deck of the Titanic on the evening of April 14th, 1912, and you know what happens later at night. What would you do?
You’d warn everyone of the imminent dangers ahead. You would get the captain and crew to take the threats seriously, slow down, and take precautions. No one on board wants the Titanic to hit an iceberg.
Now notice where we’re standing. Every retelling of the Titanic is haunted by “If only…” We can still feel the tragedy of it all. At this very moment, we are living inside our “If only…” before the fact rather than after.
This is our gift, but it comes with an expiration date.
And here’s what I keep coming back to. There was a narrow corridor through that ice field. But a corridor isn’t a straight line, and it doesn’t stay put. It shifts as the ice moves, which means we can’t chart it once and stop paying attention. We must navigate it slowly, every lookout posted, ready to change course the moment something appears ahead. That is the opposite of a race, and it’s our best bet for getting through icy water alive.
And the icebergs we’re least able to steer around are the ones we make out of each other. Our hatred and division create our icebergs. Jesus warned that a house divided against itself cannot stand, a line Abraham Lincoln borrowed to hold a country together.
If we’re all neighbors in an interconnected world, then we all live under one roof in House Humanity. We can’t destroy one side of our house without destroying it for everyone. But hidden inside Jesus’s famous warning is a promise: The house united will not fall.
Humanity has always been our own worst enemy. Since we are the problem, this also makes us the solution. We know from history that when we work together there is no stopping us. Can you believe we landed humans on the Moon with 1969 technology? We just have to get out of our own way and live the truth we already know: if we work together, we become unsinkable.
The greatest tragedy of the sinking of the Titanic would be that we don’t learn from it. Yet, if we learn the lesson from this cautionary tale, then all those who perished on that fateful voyage will not have died in vain.
So, the time to choose to slow down and be careful is now. Let us work together to navigate through the icy waters toward the promising shores beyond.
The corridor into our thriving future is as narrow as our divisions and as wide as our cooperation.
One might ask, how is it possible that humanity could finally overcome our differences and work together for the good of us all? We’ve failed to overcome tribalism thus far, what makes this time any different?
We’ve never had AI before. We can choose to use AI to help us overcome the very problems it is creating for us. It’s all about choosing to ask the right questions. And that time is now.
Neighbor, are you ready to help create the narrow corridor into our thriving future?
Mike Brooks, Ph.D., is a clinical psychologist in Austin, Texas, and founder of The One Unity Project.




Comments