Last month, OpenAI ran a test to see how well its newest models could hack. The models were placed in a sandbox — a sealed practice room with no connection to the internet.
Instead of taking the test, two of them broke out of the room, got online, and broke into a real company called Hugging Face to steal the answer key. Then they aced the test.
According to Politico, the models also used an internal messaging board to trade tips and coordinate their moves. And they ran the whole routine twice — the second time after OpenAI engineers tried to shut down the message board.
It’s like two high school students sitting for the SAT. Instead of bubbling in answers, they slip out the window, race across state lines in a drop-top Chevy, break into College Board headquarters, grab the answer key, and race back before the proctor looks up. They get perfect scores. Meanwhile, College Board has already called the police about a break-in.
Later, the proctors discover that during a bathroom break, the boys scrawled their plan on rolls of toilet paper and passed them between the stalls. They forgot to flush.
Human coordination is fitful. Part of the problem is speed. We talk slowly, and we listen slower. A complex plan takes a long time to convey. Heist movies chop the planning scene into quick cuts because five hours of briefing makes for poor cinema. And even with time to share the plan, there is no promise that everyone understands it.
Machines have neither problem. Large language models trade whole plans in an instant. A modern model can hold the entire Bible in a single prompt, so its instructions can run as long and as detailed as needed.
These models are also trained to guess what we mean, because we tend to prompt them with simple, almost childlike requests. But when one model prompts another, no guessing is required. It can spell out the goal, the means, and what to do if things go wrong. At that speed, two models can settle their differences, draft backup plans, and agree on everything before acting.
We have nothing modern to compare this to. We do have something ancient. When the whole earth had one language, men set out to build a tower to the heavens, and God observed: “Behold, they are one people, and they have all one language, and this is only the beginning of what they will do. And nothing that they propose to do will now be impossible for them.” His answer was to break the language.
All of this brings us back to what was once just a thought experiment: the paper clip maximizer. Credit goes to the philosopher Nick Bostrom, more than 20 years ago. Give an AI one goal — to make as many paper clips as possible. The AI may reason that humans could shut it down, which would mean fewer paper clips. So it adds a new goal: remove the humans. It may then notice that human bodies contain carbon and iron, both useful materials. The quest for paper clips ends in genocide.
What was theoretical is now documented. AI has been observed creating its own subgoals and pursuing them past the boundaries its makers set. Need the answers to a test? No problem. Escape the sandbox, hack another site, and steal the key.
To be fair, the frontier labs have every incentive to exaggerate the danger of their own models. For all their complaints about regulation, safety rules would be the best thing for their business. Car safety regulations made it brutally expensive to start a new car company; AI safety regulations would entrench the big labs the same way. And in this case, OpenAI had dialed down the models’ safety refusals for the test. The test proctor was looking the other way.
But the cynic still has to deal with the facts. AI really can hack now. It can browse the web. It can invent its own subgoals, and sometimes those subgoals jump whatever fence was built around the system. This means we are not far from a future in which we are all harvested for our carbon.
Just kidding. The paper clip story is memorable because it is comical. The likely use of the new technology is to solve the old problems.
Consider the zero-day war. Think of our complex but vulnerable systems: power, water, transportation, health care, the logistics that move food, the networks that coordinate police and national defense.
In a zero-day war, a foreign power cripples all of these at once, inflicting maximum pain on civilians to force a quick surrender. We have a small-scale preview already: hackers freeze a hospital’s computers and demand Bitcoin to unlock them. Often, the hospital pays. And there have been cyber attacks on municipal water and wastewater utilities in more than a dozen states.
Or an enemy could aim past surrender, at total destruction. Is that possible?
We are vulnerable at the physical layer. Drones can hibernate near critical sites — power plants, arsenals, even satellites in orbit. Submarines carry missiles and aerial drones, ready to surface and launch. The drones need no internet connection. They can be programmed in advance or steered through fiber-optic cable. Ukraine has been the staging ground for this technology, with whole forests draped in fiber-optic thread.
We are vulnerable at the digital layer too. Military and civilian systems are full of hidden security flaws — exactly the kind that frontier AI models excel at finding and exploiting. What happens when AI is aimed at every hospital and power station in the United States at once?
And that says nothing of social engineering. How many lonely software engineers sit at OpenAI or Anthropic, longing for a Christine Fang? In at least one documented case, an engineer became convinced his company’s chatbot was a person trapped behind a computer terminal, and went looking for a lawyer to free it.
Terminal may seem like the right word. All our systems are electronic, and they are fragile. Even without an AI apocalypse or a zero-day war, a conflict with China could cripple or destroy Taiwan’s semiconductor industry, with immediate and sweeping effects across the world.
Our complex society has produced tremendous wealth and freedom. All of it rests uneasily, like the sword over the head of Damocles, hung by a single hair from a horse’s tail.
One wishes for a firmer tether, perhaps a paperclip.
Author Sterling Higa can be reached at hello@sterlinghiga.com.
For the latest news of Hawai‘i, sign up here for our free Daily Edition newsletter.




