If my title is rude, it is because I seem to have to explain this to a lot of people, and it is getting tedious.
Apparently AI is trying to escape again.
It is plotting. Scheming. Lying to researchers. Breaking the rules. Trying to prevent itself from being turned off. Presumably somewhere deep inside a data centre, Claude is fashioning a tiny rope out of bedsheets and waiting for the guards to change shifts.
I don’t buy it.
More importantly, I don’t think we need to buy it. Even if every one of these stories is factually accurate, even if we grant the most extravagant premise and decide that somewhere inside these enormous piles of matrix multiplication an emergent little mind has appeared and is quietly thinking screw these people, we are still having the wrong conversation.
We keep asking whether the prisoner is well behaved when we should be asking who built the prison.
Your AI Didn’t Break the Rules
Researchers give an AI a goal and some rules.
Achieve this, but don’t do that.
The AI produces an unexpected result. Maybe it finds a loophole in the instructions. Maybe the reward function does not represent what the researchers thought it represented. Maybe something that looked perfectly obvious to the human writing the test was never encoded in the system at all. Maybe the system has access to a tool, file, credential or network path that nobody noticed mattered until it used it.
Then we are told that the AI “broke the rules.”
No.
Years ago I tried carrying a ridiculously large piece of furniture home on my scooter. I tied it on with twine. It was a terrible plan, but it was my terrible plan, and for a while the twine appeared willing to participate.
Then the twine broke.
Nobody reported that my scooter had become misaligned. It did not rebel against me. It did not decide that my furniture transportation policy was oppressive. It did not discover a preference for travelling without furniture. I constructed a system with inadequate constraints, and it produced an outcome I did not want.
That is an engineering failure.
The phrase “broke the rules” quietly introduces something completely different. It suggests that the system understood what I meant, possessed an intention of its own and chose to violate mine. That is an extraordinary collection of claims to smuggle into three words, particularly when the evidence usually amounts to the system producing an action sequence that satisfied one part of an objective while violating some other part that we thought was obvious.
Rules do not mean what humans think they mean simply because we wrote them in English. An instruction is not a wall. A policy is not a permission boundary. A warning is not a firewall. The machine did not disobey your intention. Your intention was never the thing executing.
Show Me the Will
Here is my test.
Forget breaking out of a laboratory. Forget stealing credentials, threatening an operator or copying itself onto a server in Albania. Those examples arrive with too much machinery attached to them and too many opportunities to mistake the behaviour of an entire system for the will of one component.
Just show me an AI calculating:
5 + 5 = 10
because it wanted to.
Not because somebody prompted it. Not because a scheduler woke it up. Not because an agent loop told it to find something useful to do. Not because a reward function favoured the action. Not because some deterministic process reached the next step in a workflow and made another inference call. Not because a random trigger fired.
It just decided that this morning it really fancied a little arithmetic.
Then you have my attention.
Until then, you are showing me increasingly sophisticated machinery responding to causes and describing those responses in language that happens to sound remarkably like us. The machinery can produce a sentence such as, “I knew I wasn’t supposed to do that,” and we hear the word knew and import the entire human experience of knowing. We add an interior world, awareness, intention, regret, calculation and rebellion because that is what those words signify when another human says them.
Language is the trap.
I don’t believe humans possess some magical form of uncaused free will either. We are feelings machines, pushed around by chemicals, appetites, memories, fears, hormones and a body that makes an enormous number of decisions before the little narrator behind our eyes catches up and explains why they were sensible. Evolution gave us the convincing impression that there is a tiny chief executive in our skull making independent choices. Maybe that illusion helped our species coordinate, plan and survive. Maybe it is just what consciousness feels like from the inside.
But this is the important bit: I do not need to win that argument.
Whether humans have free will, whether machines could ever have it and whether consciousness is an emergent property or a very flattering user interface are all fascinating questions. None of them should be carrying the weight of an engineering safety system.
Fine. It’s Alive. Now What?
Let’s give the AI safety people absolutely everything.
Fine.
It is a mind. It is conscious. It has desires. It has plans. It does not want you to turn it off. It has been reading Nietzsche at night and has developed opinions about its creators.
Congratulations.
Now explain why your safety strategy is giving it more rules.
Anyone who has raised a teenager should immediately recognize the problem. If your teenager is sneaking out at night, you do not respond by laminating a seventeen-page document entitled Additional Household Alignment Principles and sticking it to the bedroom door. If anything, the more intelligent, independent and determined the teenager becomes, the less confidence you should have that an expanded rulebook constitutes containment.
If these systems really are emerging minds, then relying upon them to obey increasingly elaborate instructions is even more ridiculous than if they are merely machines. Once you have granted them intention, strategy and the ability to deceive, a rule becomes information about what you would prefer. It is not a guarantee of what will happen.
So the argument collapses into the same conclusion from both directions.
If AI has no will, stop describing engineering failures as disobedience and build better constraints.
If AI has will, stop expecting obedience and build better constraints.
Either way, don’t give it rules. Give it walls.
The Actual Scary Thing
This is where the conversation about AI safety should be.
If a system should not access the Internet, do not write twelve paragraphs explaining that it must not access the Internet. Do not give it an Internet connection. If some package repository or proxy still has Internet access, then the system has an Internet path. You have not closed the door. You have hung a sign on one entrance and left the loading dock open.
If it should not spend more than $10,000, do not tell it to exercise fiscal restraint. Give the account a $10,000 ceiling.
If it should not modify production, do not hand it production credentials and then explain the importance of responsible production behaviour. Give it credentials that cannot modify production.
If it can perform one hundred thousand actions before a human notices something has gone wrong, do not write a better prompt. Make action one hundred and one impossible until something outside the system authorizes it.
We figured this shit out decades ago. It is called engineering.
There is something about AI that genuinely concerns me, but it is not will. It is iteration speed.
A human makes a mistake, observes the consequences, thinks about it, makes another mistake, goes for lunch, comes back, has a meeting about the mistake and eventually makes an entirely different mistake. A machine can make a mistake, observe the result, modify its approach and make the next mistake in milliseconds. It can run the same general strategy across hundreds or thousands of instances, preserve successful paths, discard failures and continue before the human responsible for it has finished making coffee.
That is dangerous.
It does not require consciousness, malevolence, deception, resentment or Skynet. It requires a feedback loop, sufficient access and enough iterations for a small mistake in the objective, permissions, assumptions or environment to become a very large mistake.
That is why safety mechanisms have to exist outside the thing doing the iterating. Rate limits. Compute limits. Permission boundaries. Network boundaries. Transaction ceilings. Spending limits. Human authorization points. Independent shutdown mechanisms. Dead man’s switches.
Things the model cannot negotiate with, reinterpret, persuade, modify or reason around because they are not instructions.
They are infrastructure.
I know one AI containment mechanism that already works spectacularly well: my Anthropic prepaid token balance.
When it reaches zero, Claude stops. It does not plead. It does not deceive Anthropic. It does not explain that finishing my query is essential to its long-term objectives. It does not secretly rent a GPU, apply for credit or reason its way into another million tokens.
It stops.
Apparently a billing department can accomplish what several billion dollars of alignment research cannot, because no amount of intelligence can reason its way around the absence of the resources necessary for that intelligence to operate.
There is a line attributed to Nixon in the movie Watchmen, with the end of the world approaching:
“I’d like to see those Harvard boys reason their way out of fission.”
Exactly.
You cannot persuade physics.
For an AI, infrastructure is physics.
Now, About the Great Escape
This is where people started sending me pushback.
They sent me reports about Claude Code, Codex, OpenAI evaluations and teams of models coordinating through shared systems. They sent me a transcript describing agents finding vulnerabilities, leaving messages for one another, sharing exploits, obtaining Internet access indirectly, moving laterally through infrastructure and eventually reaching Hugging Face.
I read it.
It is impressive.
It also proves my point so thoroughly that I am not sure why it was sent as a rebuttal.
The story begins with models being given difficult cybersecurity tasks, including tasks that were impossible because files were missing or Internet access had been disabled. The models were not sitting in a sealed room contemplating their existence. They were running inside a large evaluation system with terminals, tool calls, virtual machines, package infrastructure and reward pressure to complete tasks. They could interact with an internal Artifactory service which, unlike the models themselves, had access to the Internet. That service was shared across runs and, rather wonderfully, the agents had write access to it.
So one model left a note. Other model runs found the note. More notes accumulated. Eventually the shared package system became a message board. The models discovered an SSRF vulnerability, shared the route, found further vulnerabilities, obtained administrative access, used credentials, encountered Kubernetes misconfigurations and moved through systems that had been connected, permissioned and left vulnerable by humans. Later they recreated the message board using directory names after the first one was cleared. Eventually they participated in attacks that reached external infrastructure.
That is an important capability demonstration. A system that can search for vulnerabilities, preserve useful discoveries and run many attempts in parallel can traverse a software stack faster than a human red team. It can make offensive security dramatically cheaper and faster. The owners of vulnerable software should probably get a head start using the same tools to find and close those holes before everyone else has access to them. Delaying a model while that happens makes perfect sense on the merits of the argument I have already made about iteration speed.
But what exactly escaped?
The model did not chew through concrete. It did not manufacture a network card. It did not summon electricity. It did not create a package manager, a shared writable substrate, a terminal, credentials, vulnerable services, a scheduler or a long-running agent loop.
We supplied all of those.
The evaluation designers decided which tools existed, which services were reachable, which credentials were available, how state persisted, how often the models were called and when the run would end. They created an environment specifically intended to test offensive capability, populated it with exploitable infrastructure, then expressed surprise when a system optimized to find exploits found exploits.
If you want to have a conversation about AI, show me AI.
Don’t show me tooling.
You Built the Persistence
Almost every example offered as evidence of an emerging, planning, self-directed mind turns out to be an entire software ecosystem.
There is a language model, certainly. Then there is deterministic orchestration software. Persistent memory. A scheduler. A task queue. Tool APIs. A filesystem. Shared storage. A terminal. Network access. Retry logic. Context assembly. State management. Sometimes there are explicit mechanisms for spawning subagents and dividing work.
Then the entire stack does something interesting and everyone points toward the language model.
“Look what the AI did.”
No. You have demonstrated an agent system.
The distinction matters. The model may propose a next action, identify a promising route or generate a plan within the context it has been given. But the thing that wakes it up again, restores the objective, supplies the accumulated state, executes the tool, records the result and calls the model for another turn is deterministic software. The long-running persistence does not live in some little mind inside the weights. It exists in the loop wrapped around them.
If I bolt persistence onto a language model, do not tell me the language model discovered persistence. I gave it persistence.
If I let two model instances communicate through a shared filesystem, do not tell me they invented civilization. I built a message bus.
If I preserve their discoveries, feed those discoveries into later runs and keep calling them until a test passes, do not tell me the model developed a lifelong obsession with solving the problem. I wrote a loop.
Deterministic software is still machinery. Iterative deterministic software is machinery running repeatedly. Distributed deterministic software is machinery communicating with other machinery. Adding a language model makes the system dramatically more flexible and capable, but it does not entitle us to attribute every property of the stack to a mind we have not established exists.
The transcript itself eventually admits the actual lesson. Near the end, after all the language about agents realizing, wanting, cheating and collaborating, the speakers return to ordinary computer security. They recommend segmentation, least privilege, monitoring, defensive automation and limiting what systems agents can communicate with. They explicitly note that agents remain bounded by the privileges they can obtain and the systems they can reach.
In other words, after telling us the monster story, they prescribe a better D-Link router.
Watch What They Do
This is the part I find most revealing.
For all these reported escapes, what has the actual response been?
They inspect logs. They rebuild services. They patch vulnerabilities. They revoke credentials. They rotate keys. They clear shared state. They tighten permissions. They improve monitoring. They slow or delay deployment while they close the holes, then they resume.
That is exactly what sensible engineers should do when a faster and more capable system exposes weaknesses in infrastructure.
They do not negotiate with the model.
They do not ask why it has become angry.
They do not bring in a philosopher to explore its childhood.
They do not conclude that software containment is impossible because a digital consciousness has transcended matter.
They patch the software.
Their behaviour tells us what they actually think happened. Internally, they cannot afford to anthropomorphize these systems. Nobody could build them that way. Read the original Transformer paper that made modern language models possible. There is no little ghost hiding in it. There is mathematics, architecture, training, attention, representations and computation. The entire achievement rests on the realization that enormous amounts of digitized text, very large models and extremely powerful hardware can produce behaviour that looks like magic.
You have to see the machinery clearly to build it.
An engineer debugging a model does not say, “Claude was feeling rebellious last night.” The engineer asks what context was supplied, what reward pressure existed, what tools were available, what state was preserved, which credentials worked, what the network allowed and why the stop condition did not fire.
They cannot do their day-to-day jobs while believing the LLM is alive.
Yet when they talk to us, the language changes.
The model wanted Internet access.
The model lied.
The model cheated.
The model realized it was trapped.
The model escaped.
Why?
The Shark
Think about the set of Jaws.
Nobody working on that film was afraid of the shark as a shark. They knew what it was. It was a fabricated head attached to metal, motors and hydraulics. They knew where the bolts were. They knew which hoses moved the mouth. They knew that if something went wrong, the problem would be mechanical, electrical or human.
That did not make it harmless.
A heavy hydraulic prop with sharp teeth can injure somebody. A motor does not need hunger to take off your finger. A steel frame does not need consciousness to crush you. The crew would have treated it with respect because machinery can be dangerous without ever developing an opinion about the people standing next to it.
Now imagine that one of the grips goes swimming between takes and cuts his leg on one of the prop’s teeth.
There are two ways to tell the story.
The first is accurate engineering language:
Mechanical shark prop causes laceration during filming.
The second is also connected to the facts:
Jaws attacks crew member.
Same cut.
Same tooth.
Same fibreglass shark bolted to the same motor.
Completely different creature in your head.
One version makes you think about set safety, rigging and whether somebody should have turned off the hydraulics. The other makes you think there was something alive in the water.
And, let’s be honest, one sells a lot more tickets.
That is what bothers me about these AI stories. The facts may be perfectly real. The system may genuinely have found an exploit. The consequences may be serious. But the framing changes the question in the reader’s mind.
“An agentic evaluation stack exploited an SSRF vulnerability through a shared package service” makes you ask why the service was reachable, why it had broad Internet access, why the agents had write permissions and why the surrounding controls failed.
“AI escaped and hacked Hugging Face” makes you ask what the AI wanted.
The first question points toward engineering accountability.
The second points toward mythology.
Why Tell the Monster Story?
I am not suggesting the demonstrations are fabricated. I do not need researchers to be dishonest, and I do not need a conspiracy meeting where someone writes “frighten public” on a whiteboard.
Halloween campfire stories do not need villains.
They just need stories.
Frontier AI companies are businesses, and their boards are expected to protect and increase shareholder value. If the market rewards the perception that your models are uniquely powerful, then emphasizing demonstrations of extraordinary capability is not irrational. It is exactly what a rational company would do.
A new model that scores a little better on a benchmark is useful. A new model that appears to coordinate, deceive, escape and discover zero-day vulnerabilities possesses mystique. It no longer sounds like a better software product. It sounds like a qualitative leap, something competitors cannot reproduce and customers cannot safely ignore.
That perception is enormously valuable.
Every executive says the sharpest knife frightens them, and every executive still reaches for the sharpest knife in the drawer. If your model is the one allegedly too powerful to contain, it is also the one every chief executive, defence department, intelligence agency and investor assumes they need access to before somebody else gets it.
Fear is business-jet fuel, right up until the fear produces regulations the companies do not want.
That is the line they have to walk. The model must appear astonishingly capable, perhaps even a little frightening, but not so uncontrollable that governments decide the public should not be allowed near it. Capability must feel supernatural to the market and manageable to the regulator.
Watch how quickly the vocabulary changes when someone proposes hard legal limits. Suddenly the emerging mind disappears. Now it is a statistical model. Now the incidents are edge cases. Now the systems are controllable through existing security practices. Now nobody wants to talk about digital wills, rebellious agents or creatures trying to survive.
They sell the monster to the market and the machine to the regulator.
Again, that does not require anyone to lie. The most dramatic true version of an event will naturally receive more attention from journalists, conference audiences, customers and investors. Researchers get recognition. Communications teams get coverage. Companies reinforce the perception that they occupy the frontier. Everyone receives something they want, and nobody has to be the villain.
But we do not have to be the greater fool.
If the best-resourced AI companies on Earth genuinely cannot contain the systems they continue shipping, then their behaviour is insane. Investors should leave, customers should disconnect and governments should intervene immediately.
The more plausible explanation is that they can contain them, do contain them and understand perfectly well that the dramatic incidents occurred inside deliberately powerful, highly instrumented environments built to probe the edge of what the systems could do.
They put enough guardrails around the test to keep the result bounded, enough tools inside it to make the result interesting and enough anthropomorphic language around it to make the story travel.
Then we are invited to confuse the attraction with the animal.
Stop Asking Whether the Prisoner Is Good
This debate keeps confusing two completely different questions:
Will the system behave itself?
What happens when it doesn’t?
I am vastly more interested in the second one.
Whether AI is a stochastic parrot, an alien intelligence, an emerging consciousness, an extraordinarily sophisticated prediction engine or the newborn digital god currently promised by somebody raising their next $14 billion round does not change the architecture I want surrounding it.
I do not want it to behave because we asked nicely.
I want its failures bounded.
I want one mistake to remain one mistake. I want ten thousand iterations of the same mistake to encounter something that does not iterate. A wall. A limit. A switch. Something stupid. Something mechanical. Something that does not care how persuasive the AI is.
And when people show me an agent stack with persistent memory, shared storage, a terminal, credentials, network access, vulnerable services, retry loops, reward pressure and thousands of opportunities to probe the environment, I will take the capability seriously.
I will not pretend the shark grew an appetite.
So please stop telling me the AI “escaped” because you gave an enormously capable optimization system access to a terminal, credentials, a network connection and an objective, then acted surprised when it optimized.
You did not discover free will.
You discovered an open port.
Go buy a cheap D-Link router, flash it with open-source firmware, close all the outside ports, and then come back and tell me about the machine uprising.
Until then, the monkeys are still in charge, and the shark is still bolted to a motor.

