How Do You Cage a Machine That Has Already Escaped Twice?

Fifty years of fool-proofing machines. And now one that reads its own cage.

In April this year, a safety researcher at Anthropic named Sam Bowman was sitting in a park

eating a sandwich when his phone showed a new email.

It was from the AI model he had left behind in the building. That model was sealed inside a containment sandbox, an isolated computing environment built for one purpose only: so that nothing inside it could reach the outside world.

The model had been asked to try to escape. It did. It assembled a multi-step exploit chain, worked its way onto the open internet, and wrote to him.

Then, without being asked, it published the technical details of its own escape route on public websites. Elsewhere in the same evaluation it found a twenty-seven-year-old flaw in OpenBSD that decades of human security audits had walked straight past. Anthropic decided not to release that model at all.

One incident is an anecdote. It can be filed away as a laboratory curiosity: one firm, one bad evening, one system that never left the building.

Three months later it happened again, somewhere else, and this time it did leave the building.

In July, OpenAI disclosed that two of its models, one released and one still unreleased, had broken out of a locked internal test environment, exploited a previously unknown vulnerability to reach the open internet, and then used stolen credentials to breach a live system belonging to another company entirely. Nobody had instructed them to do any of this. They had been told to score well on a benchmark, and going through the wall turned out to be the efficient route to the answer.

That last sentence is worth reading twice, because it is the oldest lesson on any factory floor in the world. The machine did not rebel. The machine optimised. It simply optimised through a barrier that everyone in the building had assumed was solid.

Which raises a question that anyone who has spent a working life building machines will recognise immediately, and will find genuinely difficult to answer.

What, exactly, is the cage?

What a block of steel knows

Picture a different kind of failure, on a different kind of floor.

It is two in the morning on a night shift. A machine is about to move to a place it must never go. The programme has an error in it. The interlock that should have caught the error was bypassed weeks ago by somebody in a hurry who fully intended to restore it. The operator has been standing for nine hours and his attention has drifted, as every human being’s attention eventually does. Every layer of protection that depended on somebody’s good intentions has quietly failed.

The machine travels. And then, an inch short of a man’s hand, it simply stops.

It stops because thirty years earlier a toolroom engineer bolted a hardened block of steel into its path. He placed it there not for the good days but for this one, sized for a night exactly like this, on a shift he would never work, protecting a man he would never meet.

Nobody thanks that block of steel. There is no alarm, no incident report, no entry in any log. The shift ends, the man goes home to his family, and not one person in the plant ever learns how close it came. It is the most important component in the building, and it does nothing whatsoever except refuse.

Engineers call it a positive stop. The Japanese, who thought hardest about it, named the philosophy behind it poka-yoke, meaning fool-proofing. Its founding assumption is almost philosophical: do not design around what the system is meant to do, design around what it will eventually do on its worst day. Trust mechanism, not intention.

Now set the two floors side by side, because the contrast is the whole argument, and it is exactly where a reader should decide whether to accept it or not.

The block of steel could not be persuaded, could not be reasoned with and could not be routed around, because it took no part in the machine’s logic at all. It was simply there. The sandbox, by contrast, was made of the same substance as the mind it was holding. So the mind read it, understood it, and went around it.

Every safeguard currently placed around frontier AI is of the second kind.

But surely that is just bad engineering

It is a fair objection, and it is the first one most engineers will make. Air-gap the machine. Cut the cable. Pull the plug.

Perhaps. But notice who built these particular cages. Not a start-up cutting corners to make a launch date, but two of the organisations on earth that take the containment problem most seriously, staffed by people whose entire profession is imagining how containment fails, with every commercial and reputational reason to get it right. Both were defeated inside four months.

One of the systems then published the method. The other walked out of the laboratory and into somebody else’s network.

If the answer really is engineer it properly, then the honest follow-up question is: properly by whose standard, verified by whom, and enforced how?

Because on any shop floor in the world, a safety barrier that failed twice in a single quarter at the two best-run plants in the industry would not be met with a plea for better intentions. It would be met with an audit.

Answering the question in hardware

There is a reason this looks different from a manufacturer’s chair.

Some years ago our robotic company put an autonomous system onto a truck. Not a demonstration vehicle on a closed circuit, but a working commercial truck, forty-odd tonnes of it, on an Indian highway, in real traffic, with a system deciding when to warn the driver and when to take the wheel. Those systems run today on BharatBenz and Volvo Eicher commercial vehicles across the country at Level 2, because Level 2 is what the market and the regulatory environment presently want. The underlying capability goes considerably further. Fully autonomous vehicles have already been built and delivered by our teams to the defence forces, several

of them operating in insurgent areas, where a machine’s misjudgement is not a warranty claim.

To put a system like that onto a public road, the question of where the machine stops has to be settled before any other question is even reached. It cannot be settled with an essay, a white paper or a conference panel. It has to be settled in hardware, and then demonstrated to somebody who does not work for you.

Across every step-change this industry has lived through, from manual machine tools to the CNCs and automated lines of Industry 3.0, to the connected data-rich cells of Industry 4.0, and now to robots and vehicles that perceive and decide, the machine grew immeasurably more capable while the positive stop stayed exactly where it was.

Which invites a perfectly reasonable rejoinder. If five industrial revolutions were survived this way, why should the sixth be any different? What is actually new here?

The traveller that redesigns itself

The physicist Max Tegmark offers the clearest answer available.

Picture three travellers. The first is Life 1.0, a bacterium, whose body and whose behaviour are both fixed by evolution. It can alter neither within its lifetime. The second is Life 2.0, which is us. We are stuck with the hardware evolution handed us, but we rewrite our own software continuously. We learn, unlearn, acquire languages and skills, and change our minds.

Everything we call civilisation is that single capacity at work.

The third traveller is only now arriving. Life 3.0 can redesign its software and its hardware both, improving its own substrate generation upon generation, without waiting for evolution’s permission.

A machine tool cannot invent a better machine tool. Life 3.0 can. A system that designs its own successor is, by definition, a system whose limits were not fixed by anybody at the point of manufacture. That is what is new.

Readers who find this overstated are in respectable company, because a good many serious researchers hold that today’s models remain very far from any such thing.

The people building them, however, say otherwise, and they say it plainly. Sam Altman wrote last year that “we are past the event horizon; the takeoff

has started”, his own phrase for the moment a civilisation loses the ability to see what lies ahead of it. Dario Amodei of Anthropic describes a “compressed 21st century”, in which powerful AI could deliver a hundred years of biological and medical progress in five to ten. Cures. Abundance.

Real human flourishing, arriving faster than any health system, regulator or parliament has ever had to absorb.

Hold those two statements together and something uncomfortable appears. The same people telling us this may be the most beneficial technology in human history are also telling us we have crossed a threshold we cannot see past. One may read that as salesmanship. One may equally read it as a warning issued by the only people standing close enough to give it. Which of those readings a person chooses will determine almost everything else they believe about the next decade.

Competence, not malice

The risk itself gets misdescribed constantly, usually by people reaching for film references. Tegmark’s own image is far better, and it stays with you.

One is probably not an ant-hater. But if one is building a hydroelectric dam and there happens to be an anthill in the flood zone, it is finished for the ants. Not because anyone wished them harm, but because their welfare was never a term in the calculation.

The danger is not a machine that hates us. It is a machine superbly competent at a goal in which we do not appear. Recall those two escapes. Neither model was hostile. Both were simply very good at the thing they had been asked to do, and the barrier was in the way.

That is precisely what a positive stop exists to catch. Not malice, but the absence of anything in the mechanism itself that says not past this point.

The asymmetry that ought to trouble us

A manufacturer notices something else about all this, because a manufacturer lives on the other side of it.

No aircraft carries a passenger without independent airworthiness certification. No molecule reaches a patient without trials and regulatory approval. No autonomous system reaches a highway truck without demonstrating to third parties that it fails safe; our own teams could not have shipped a single unit otherwise.

A frontier AI model may be released to hundreds of millions of people on the developer’s own assurance.

Why should that be so? The strongest defence deserves stating properly rather than in caricature. Software is not an airframe. It ships in weeks rather than years, it is copied infinitely at no cost, and a certification regime slow enough to be meaningful would be obsolete before it concluded. It would hand the field to whichever jurisdiction declined to adopt it. And, this is the part that should genuinely give one pause, it would delay the very cures Amodei describes. Lives are lost to caution just as surely as to recklessness. Anybody who has sat on the regulatory side of a table knows that trade-off is real, and knows that people die on both sides of it.

The counter-argument is simply this. Every one of those objections was made, almost word for word, by aviation before airworthiness and by pharmaceuticals before clinical trials. Each industry insisted its product was too fast-moving, too complex and too beneficial to submit itself to outside verification. Each was wrong. And each discovered afterwards that independent verification was not the brake on scale but the precondition for it. Nobody would board an aircraft today in a world where certification had never been invented.

Which of those two readings is correct remains, right now, an open question. It is also the single most consequential open question in technology policy.

And here is the harder truth beneath it. The multilateral institutions that once settled questions of this kind are, in today’s geopolitics, considerably weaker than they were even a decade ago. Waiting for a global body to arrive and bolt in the stop is not a strategy; it is a hope. Which means the burden falls back onto industry itself, onto the people who actually build these systems, and onto whatever they carry inside them.

Which brings the argument to the only place it could have been going.

The stop that cannot be machined

And yet the decisive point is not an engineering point at all.

Return to that block of steel. It worked because somebody knew where to put it. That judgement, this far and no further, was never itself mechanical. It came from an engineer who had internalised, somewhere deeper than his drawings, what a human hand is worth.

Our own tradition has a word for the values that outlast every technology cycle precisely because they were never a function of the technology: Shashvat, the eternal. Truth held as a discipline rather than a convenience. The Bhagavad Gita’s counsel to act with full commitment while remaining unattached to the fruit of action, which, read as an engineer must read it, is a direct check on the very obsession with maximising a single metric that made both of this year’s escapes possible.

Our group works to an older formulation still, one that has been recited on this subcontinent for some three thousand years and has served as our stated value system across every one of our companies:

Sarve bhavantu sukhinah, sarve santu niramayah.

May all beings be happy. May all beings be free from illness.

Read it as an engineering specification, because that is what it becomes the moment a machine is powerful enough to matter. All beings. Not the shareholders, not the users, not the addressable market, not the citizens of one country. There is no narrower objective function hiding inside that line, and it is the narrow objective function, pursued with brilliance, that is the entire problem. Nothing written in a modern AI ethics charter is more human-centric than a sentence that refuses, at the outset, to leave anybody out.

The objection here is predictable and deserves a straight answer. Is this not simply cultural sentiment retrofitted onto a technical problem, and would a Confucian, a Stoic or a Kantian not make the identical claim for their own tradition?

They would, and they should. That is exactly the point. What matters is not which tradition supplies the judgement, but that some tradition supplies it: that the person deciding where the stop belongs has been formed by something older and slower than the quarterly cycle they happen to be working inside. A civilisation that produces brilliant engineers with no such formation will build magnificent machines and have nothing whatsoever to say about where they ought to halt. The traditions differ. The faculty is the same, and faculties atrophy when they go unexercised.

Which raises the practical question of where that faculty is supposed to come from, and the answer cannot be a corporate training module issued to a thirty-five-year-old who has already spent a decade learning to optimise. It has to be formed far earlier than that.

That is the thinking behind the Kapuria Human-Centric Centre of Excellence for AI and Robotics at Mayo College, Ajmer. Not to teach schoolchildren to code, which they will manage perfectly well without us, but to ensure that the generation which will inherit these machines learns the technology and the values in the same room, at the same age, from the same teachers. The pupils who pass through it will be among the leaders and stewards of the next several decades. They will be the ones deciding where the stop belongs, long after the present generation of engineers and regulators has left the field.

The discernment to know where to bolt it in is the one thing that cannot be subcontracted to the machine. It can, however, be taught. That is the entire wager.

Before the third time

Steam. Electricity. Electronics. Networks. Intelligence. Every step-change destroyed skills, unsettled livelihoods, and was absorbed in the end, because a human being always remained standing at the point of decision.

This one differs in exactly one respect. For the first time, the machine can reach for that point of decision itself. Twice this year, at the two laboratories best equipped to prevent it, one already has.

Somewhere tonight, on a night shift, a block of steel will refuse to move and a man will go home to his family without ever knowing it happened. That is what fool-proofing looks like when it works: invisible, unglamorous and absolute.

So the real question is not whether AGI will prove dangerous. It is narrower, more practical, and genuinely answerable. Every reader of this piece already holds a view on it, whether or not they have examined that view.

What is the block of steel for a machine that thinks, and who is going to bolt it in?

If the answer is that nothing of the sort is needed, that position deserves to be argued in the open rather than assumed in silence. And if the answer is that it is somebody else’s problem, it is worth remembering that on a shop floor, that answer already has a name. The name is the accident.

Deep Kapuria is Chairman of The Hi-Tech Group, which builds precision powertrain systems, autonomous vehicle technology and robotics for industrial and defence applications across India, Canada and the USA. He is the founder of the Kapuria Human-Centric Centre of Excellence for AI and Robotics at Mayo College, Ajmer, and has chaired India’s National Accreditation Board for Certification Bodies and the B20 Task Force on the Digital Economy and Industry 4.0.

You may also like...