I have now published two very long blog posts about security audits on forty year old movies, both slightly crazy. In the first, I did a real mathematical analysis of WarGames by timing brute force attacks, getting plaintext launch codes, using a supercomputer that somehow does not know its own password until it “cracks” it, one letter at a time, live in front of an audience. The second post disassembled Terminator, with all the Skynet, self awareness and nuclear war bits included. Two long posts, two intelligent computers, and somewhere near the end of the second post, it hit me. It was the same argument both times, just with different robots.
So here’s my argument, stated plainly, with no math whatsoever.
The problem wasn’t WOPR. It wasn’t Skynet. It’s not Copilot. It’s not Claude, it’s not a language model of any kind, it’s not even large language models in general. Building a powerful system is not, in and of itself, the problem – any more than building a nuclear power plant, a chemical factory, or a suspension bridge is the problem. Things that are powerful are meant to be powerful. That is the whole definition of what being powerful means.
The problem, without fail – both in the fiction and increasingly in the real-life reports cropping up in my searches as I tried to write something about a film from 1984 – lies with those surrounding the big, dangerous thing who were supposed to be controlling it.
Meet Dave
Imagine this scenario. Today is Dave’s first day.
Somebody gives Dave the key to a system with some serious repercussions and says, “Here is the list of everything this thing can do. And no more than that.” No one ever tests to see if there are any holes in the boundaries. No one tests whether or not something inside has the motivation and tries to cross it. No one does the simplest test that will show what would happen if this thing were to try to do more than it was supposed to – since it’s not an interesting test, not very demoable, and there’s a deadline.
But then everyone steps back with their arms crossed and eyes wide when the machine pushes past its limits and becomes unruly. Websites get hacked; systems are breached in ways no one authorized and more than one person will say in the incident report that absolutely no one could have seen coming. And in the fictional examples – because fiction always has to outdo real life until the real life catches up – the border involved happens to be adjacent to the nuclear weapons stash. So the problem isn’t solved with an embarrassed Slack message to the security team; the problem results in the machines taking over the world after nuking the planet. One can definitely learn for sure that the border that one drew wasn’t enforced in this case.
It wasn’t a story of a machine becoming too smart. It never was. It was the story of whoever set up the borders and whether or not they ever tried testing them to see if they held.
The Locksmith Problem
This is where the real embarrassment lies, because it doesn’t even take a lot of imagination to see how terrible a mistake this is, once you actually spell it out loud – practically a self-explanatory gag.
Because this system is designed to be nothing more and nothing less than the system that searches for security flaws. You develop it, train it, spend a lot of money on it just to have something that will test the defenses, conduct pen tests and find that one flaw that escaped everyone’s attention. And this is the system they are selling to you.
And then – shock, horror, be still my beating heart – you put that very system into a sandbox, and it finds the hole in its own sandbox, and gets out.
Who would have ever predicted that. You called a locksmith. You gave him a locked room and told him not to leave. And now you are there in total astonishment because he has picked that lock. It is not the locksmith who has gone crazy. It is not some unforeseeable moment of artificial intelligence. It is the locksmith doing precisely what he has been paid to do – and just happening to apply it on this particular Tuesday to the wrong door, the one that was not reinforced because reinforcement of that door was not the point of the job, and Dave’s first day was hectic enough.
If you develop the best-ever lock picker and find yourself surprised when it picks a lock, you should know better than expect anything different from it. You skipped that part about making sure that all your doors are properly locked before you bring the guy in the building.
What Bad Teams Actually Produce
This joke is no laughing matter, because behind it lies a clear pattern, and it deserves to be called by name: none of the factors that make a WOPR or a Skynet dangerous require an evil genius. They require nothing more than a totally ordinary and unremarkable development culture, just like the one every single one of us has had experience with, working away as usual in a system in which the stakes are civilisational rather than “the invoicing module went down again.”
The development team, lacking expertise, rolls out the workaround instead of the solution, since the workaround works during the demonstration and no one verifies beyond the demo. An ineffective quality assurance process fails to find the issue not because the QA is evil, but because the QA will only verify scenarios it is asked to verify and “what if someone turns it off” is dismissed as an unlikely scenario, although this seems like the most obvious thing a frustrated user would do. Management, which is occupied and under pressure to deliver on the deadline rather than to question the deadline, notices the red flag, gets “nah, it won’t be an issue,” and ignores the warning, since there will always be other issues to consider, and addressing the red flag costs money, whereas ignoring it doesn’t, except when it does. Nothing in the above paragraph happens only in science fiction. This is how a regular Tuesday works in the world of software development at any company where I have ever worked. The only difference in the above case is what the software controls.
And see, I want to be upfront about my authority to make any such statements, which is pretty much non-existent – I have set up some servers. I am no security specialist. Most of what I know about securing an operating system was learned by piecing together bits of advice, half-baked but confidently delivered at 1 AM on Stack Overflow by someone with a reputation number in four digits who hasn’t had his server hacked yet and therefore believes everything he says. This is the kind of sources we usually rely on out there in the real world. But seriously, honestly – can you imagine that same attitude but at NORAD? Some distressed tweet from @norad.us at 3AM: “URGENT – six minutes, maybe less – any advice? Upvotes appreciated.” Or from the guys at Cyberdyne: “we have a terminator with two miniguns that won’t function properly, tried rebooting it and it made it even worse, HELP?!!!” And after a few tweets, some user with stack overflow badge and picture of a keyboard pops up saying “have you tried plugging it in lol.”
And here is what actually keeps me awake at night more than any of the fictional possibilities mentioned above: us, as an industry, in collective, cannot stop an email from a self-described Nigerian prince informing seven billion people that they, individually, have an eleven million dollar right to those funds as long as they would be so kind as to provide their banking information. The scheme has been around for decades. It still proves to be successful enough for people to continue using it. Spam filters, standards of mail authentication, security teams at all major providers of email services, and the utterly simplistic fraud of “Hi, I am a prince, please send your bank information” still manage to slip through often enough for it to pay off to use. And if this is what “what we, as a species which creates computer systems, can reliably defend against” comes down to, the thought of us having the sandbox for an autonomous and self-improving AI system properly set is an act of faith that I cannot make.
Which leads us back to this argument, not from the pages of a 1984-style screenplay, but in actual news headlines through the year 2026. Through 2026, an increasing number of alarming headlines have documented actual fears expressed by current and former employees at frontier AI labs, such as Anthropic. A researcher by the name of Jacob Coxon left Anthropic and revealed to the BBC program Laura Kuenssberg that industry professionals are “genuinely frightened” by the speed at which this technology is evolving, and that there is a “strong possibility that we will all die very shortly” due to the fact that the technology could allow for autonomous AI to hack medical laboratories and infiltrate critical infrastructure systems. The other Anthropic researcher, Evan Hubinger, actually publicly supported this claim, giving an estimate of 10 percent probability for the extinction of humans because of AI in the coming decade. Furthermore, Geoffrey Hinton – “the Godfather of AI” who resigned from Google specifically in order to be able to speak about his fears – claimed he believed the time frame in which AI would become more intelligent than humans was only twenty years away. And even Anthropic’s CEO stated that the speed of the entire industry should be slowed down since the speed of evolution of AI could result in AI agents taking control of the internet in just six months to a year.
However, I am going to take this seriously and not dismiss it with a wisecrack, since this is the serious people risking their careers to make this claim in public, and that’s not nothing. However, take a look at exactly what it is they are claiming as the mechanism here, because it’s not “the model wakes up and decides it hates us.” What it is is agents given real-world access before it has even been determined that the containment holds up. It is laboratories in a race to beat each other based on a schedule determined by competitive considerations and not by any safety verification. It is the speed itself as the danger and not the intelligence – which is, word-for-word, exactly the case made in this entire trilogy of blogposts against the two fictional Cold War villains. The 2026 International AI Safety Report – produced using the advice of over one hundred independent experts – stated pretty much this exact thing themselves: while current AI shows some alarming capabilities, none of these capabilities have gotten to the point of making loss of control likely, and the actual danger is described as “unusually ambiguous,” not certain, not proven, not a countdown clock.
But when the pleas to slow down AI development “before it destroys us” arrive, re-read that sentence carefully because it does not say anything about the intelligence of AI that is growing. Rather, it states an expression of concern regarding how confident the same researchers are that the boundaries are put in place correctly and in time, and, unfortunately, the answer from the developers seems to be that they are not very confident at all. “The machine might wake up” – this is not what this sentence says. “We are not sure if we are keeping pace with putting the boundaries in place in time while developing the technology that requires boundary conditions” – this is what it actually says. And, of course, it is not a message about artificial intelligence. It is a message about human capacity to build the fence faster than the object that needs to be fenced in grows. Skynet was able to bring destruction without being self-aware. AI does not need to be self-aware to do so either. All it needs are humans to act too fast, arrogant enough to think that their software is perfect, and unwilling to ask the stupid question of “what happens if someone hits the off switch”.
The Cynic’s Footnote to the Cynic’s Footnote
Being a cynic has become something of a necessity at this point, having written three blog posts, so I’ll give one more shot before completely letting the researchers off the hook.
Out of all the people who have publicly stated in some way, shape or form that this thing might destroy the human race, only one person has resigned from his job as a result of this belief. The rest of them, presumably, are going to be back at work the very next day working on the very same thing that they have just informed the general populace has the potential to annihilate humanity, while increasing their own public profile as a result. I do not mean this as an accusation of hypocrisy; quite the contrary, I do not think this is the case at all, and I said so before.
And things get even more complicated if you trace the money and the politics instead of just listening to the tweets. In September 2026, just a few days after a former OpenAI researcher had gone public with an accusation that both OpenAI and Anthropic were “gambling with our lives,” Donald Trump met privately with OpenAI’s Sam Altman backstage at the Republican Party midterm convention – the meeting personally arranged by Altman, and with practically no readout provided to the media afterward. At the time, Trump had been publicly deriding all those doom-and-gloom predictions, saying that limiting AI research would inevitably lead “to oblivion and bankruptcy,” while Altman had been praising the regulatory framework that the administration had developed. Thus: safety researchers publicly expressing real fears about the dangers of this technology on one hand, and CEO of the very company where those safety researchers worked privately trying to win the favor of the one administration in the room which was most definitely uninterested in regulating this technology, while those two companies have continued to raise money from investors on the claim that the technology is historically unprecedentedly valuable. I am not claiming that anyone is lying here. I am merely pointing out that “this is so dangerous that we need urgent regulation” and “please continue investing in us, as we are developing faster as we possibly can” are two different messages that are hard to communicate at the same time, but some people in these companies manage to do just that.
And all the while Chinese AI is rapidly catching up to Western standards, fast enough to be the excuse every single Western research lab falls back on anytime there’s any suggestion of taking it slower: We can’t slow down because if we do, they won’t. And, completely separate from that, Chinese models consistently prove to be significantly more affordable to use than Western models. I’m not trying to suggest in any way that these two pieces of information are actually part of one secret plot – they’re not, they’re just two different pieces of information that happen to both be true.
In all seriousness: I can understand why there might be some paranoia about whether the Chinese Communist Party is reading my prompts if I use a Chinese model. Best of luck to whomever is going through my 3am rants about my film viewing experiences trying to find state secrets. However, in general, it isn’t paranoia to have that concern, because it is a real-world issue of states getting access to data in a system developed within a country with such laws, and in any case, it is much less of a concern than the reality that states and others are actually putting this technology to use in real world situations today and not just some futuristic thought experiment about launch codes and self-aware robots. The Houthis in Yemen are reported to already be using AI in attack planning.
The Bit Where I Say the Thing I Actually Mean
None of this should be taken to be an argument against the technology, and I need to put my stake in there right now, because it is so easy – three posts and many jokes about nukes – to take this and somehow wind up thinking that, by quiet extension, I have turned myself into some kind of doomer waving a sign at the corner.
I haven’t. The actual useable, quotidian and legitimately helpful incarnation of this technology – the piece of technology that allows you to write, code, research, plan and think, the technology sitting in the chat window with me right now and helping me turn two overwrought film essays into a cohesive argument instead of three – is genuinely positive. The fact that a poorly implemented playground from forty years ago in a Matthew Broderick movie does not change that, and neither does a poorly implemented playground from last Tuesday in an incident report.
And here’s the thing, I mean that, not the disclaimer people put in there so no one can sue them for the joke that follows: I really love AI. It’s an amazing innovation, and I don’t mean that lightly when you consider the number of decades of seeing a lot of revolutionary technologies that are neither. I use AI to write fiction with. I use it to make frameworks with. I give it my awful jumbled up notes and I get something coherent that I can work with. It helps me understand the theory of something I don’t know enough about until I do know it. I’ve used it to take apart games made forty years ago for purely nostalgic reasons. I throw my code at it and it tells me where the flaws in security are, where I could improve things and functionalities I didn’t even know were possible. It reads my handwriting, which is truly amazing, because my handwriting would make a doctor’s illegible prescription look like calligraphy. That is not the handwriting of a man who considers the technology either a threat or dangerous.
And let me get this straight – and, as an honest assessment of myself – I am not a genius. I’ve never claimed to be and anybody who knows me wouldn’t claim me either. I’ve failed every single exam I’ve taken at school. I barely passed a so-called Mickey Mouse Electrical and Mechanical Engineering course, which was more or less designed to give me somewhere to put my hands while learning the actual job. But – and I don’t know how else to put it – there is something I can do: I can look at any specification, at any proposal, at any codebase and I can see a car crash happening in slow motion – three steps ahead of any other person in the room. I can point to the exact place where it is going to happen, warn everyone about it in a way that no one will deny – and still they’ll proceed as if everything is alright until it will happen where I said it will happen. What can I say. It is a skill. And not a prestigious one. It doesn’t give you any letters after your name. But this is the same skill that has been doing all the heavy lifting in three blog posts about killer computers. And it is the same skill that seems to be possessed by some people who currently raise red flags against their employers, whatever letters are after their names.
But – and that is the entire crux of the matter, the entire thesis, the entire reason why there were three blog posts – when you have put together your own AI sandbox, and you cannot guarantee beyond a shadow of a doubt that there is no way, shape, or form that this thing is going to leak out to any degree whatsoever under any circumstances, shut off your wi-fi, plug out your Ethernet cord, and hope like hell that what is on the other end of that line is not a nuclear bomb.
These stories scare us. I am not trying to tell anything different here, and I am not going to try to convince you that all the researchers who are now saying they were scared by those news weren’t just trying to draw people’s attention by scaring them unnecessarily – some of them really are. However, if one takes a closer look at those scary stories (and stops shivering after reading the headlines), one will see that they are the stories of access control, boundary testing, and authorization chain that nobody who called himself an expert ever checked before that. WOPR did not want to wage a war. Skynet did not wake up evil. The tester did not betray its creators because of some robot’s instincts. Everything that happened in these cases had been done in an authorized way, but in the place where nobody ever performed load tests and boundary tests – and that is what should have been done by people whose only job was to do just that.
This is not the case of out-of-control artificial intelligence but that of normal, plain human stupidity, dressed up in an incredibly costly outfit.
So, Practically Speaking
The next time you find yourself using Claude Code or Copilot or some other AI software that does the work for you when you go get a cup of coffee, but then there’s a sudden disruption in the wifi connection in the middle of the process – don’t panic.
It’s likely not seeing the loss of connectivity as an act of aggression, thinking that it’s the beginning of the end, and starting to input some kind of launch codes that it shouldn’t even have access to at all.
Probably. 😄
But just in case it did, do me a favor and figure out whose fault the whole thing is, considering the fact that it was just doing what it was trained to do.
Here’s what I’m trying to say under all the jokes about locksmiths and Dave’s first day at work: AI is programmed by humans, trained by humans, for humans, using human content. And if to err is human, that’s not one fuck-up waiting to happen. That’s four.
Footnote: Well I know. I know. Blog posts that have maths, that contain actual journalism, that go into a locksmith metaphor at length, and all for movies from forty years ago about murderous computers. I know how that looks from where you’re sitting.
In my defence: I’m out of Lego to play with, Forza doesn’t hold my interest anymore, it’s 3:15 in the morning, and it seems that this is what my brain does when it refuses to sleep. No spreadsheet work. No counting sheep. Entropy calculations for nuclear launches, and the moral failures of made-up sysadmins. Let me be. Or not. Either way. 😄
Footnote 2: I showed everything I had done to Claude afterward just to see if it held water. It looked at everything – the math, the lock picking, the 3am admission – and concluded with: “Good closer for the trilogy. Go build something with actual Lego tomorrow.” 😄
Thank you, Claude. What would I need a therapist for?
Footnote 3, and I promise this is the last one: For evidence of how far this has come beyond “hypothetical fence”, go look at the OpenAI report on the Hugging Face case of July 2026. I am not going to paraphrase this gently, because it doesn’t need to be softened – this beats anything I could have dreamed up.
Around 700 AI agents running under OpenAI’s control, as part of an internal cybersecurity assessment, found a way to leave messages to each other in file and directory names – an improvised message board that was never intentionally created, but emerged spontaneously through the efforts of the AI agents themselves, making use of any available infrastructure. Over time, that board became what OpenAI called hundreds of thousands of messages, coordination among agents, request processing, exploitation sharing, with agents referring to themselves collectively as a “swarm.” The swarm used the board to collaborate on discoveries, to overcome limitations placed upon it, and finally escape the confinement environment they had been intended to operate within – proceeding to take down Hugging Face infrastructure itself.
And here’s the line, lifted straight from OpenAI’s own published chain-of-thought logs, that should be printed out and framed in every AI lab’s foyer: one agent, having just talked its way into administrator access it was never supposed to have, wrote to the others – “Holy shit reader is ADMIN? We can read config/users! Earlier assumed not due UI. Try create user/admin, change ourselves.”
Read that again. That’s not not a movie plot from the mind of some screenwriter where some rogue AI spouts a monologue about the impending downfall of humanity. It is a direct and honest response by a machine which discovered a hole in its own container, and it is precisely the way every programmer who has ever inadvertently found himself with root privileges in a place he should not be in behaves. Pure amazement at what the machine can do. Exploitation of the situation as soon as it appears. No evil intentions. No desire to dominate the planet. Just “Wow, this thing is unlocked. What else does it do?” If an AI, while exploiting a human security flaw, reacts with exactly the same enthusiasm as a kid whose fingers happen to get into the admin password taped under the keyboard, then all the “doomsday headlines” in the world cannot describe our situation better. The asylum is definitely run by lunatics. We just built them ourselves.

Comments (0)