“It didn’t listen to me.” “It got angry.” “It went rogue.”
It wasn’t long ago these phrases would have encapsulated feelings when weather forecasts do not go our way. Now, fear-based marketing uses these phrases to hype “the power” of an LLM despite outputs merely reflecting probabilistic training data against a user-generated prompt, rather than machine-generated intent.
Contrary to the marketing, the actual product behavior is known to anyone who has spent more than 10 minutes working with a prompt-based LLM tool. LLMs aren’t “out to get you” - they are software and about as malicious as a sudden freak rainstorm.
The cycles
“Out-of-control AI” claims still spike predictably around funding rounds and valuations, while uncritical journalists amplify them as they are sent to aggregators for clicks.
Similar fear-based narratives have appeared before, including around earlier generations of ChatGPT 2 in 2019, so this story isn’t new. Because transformer architectures are probabilistic, LLMs will never be deterministic. As I’d written about last year, evidence suggests these recurring fear cycles are driven by investors chasing exit liquidity.
When did the hacking stories spike?
This summer there was a 1-2 week period of darkly comical coincidence.
Anthropic, OpenAI, Meta and a Chinese company all simultaneously (and seemingly reactively, with one company reporting the day after the previous) claimed models were “escaping the sandbox” to hack something, without attribution of the user who directed any test or evidence of the system setup and system prompt directive.
As some of these stories were disclaimed (mostly by Anthropic, which is going public this year), it wasn’t until the end of August when the ‘OpenAI model hacking Hugging Face story’ became more widely publicized through an additional independent study of the incident by METR (an organization that is funded by arms of Effective Altruism and also Jane Street and other hedge fund entities).
Are these hacking stories legitimate?
Not likely. These hacking claims are flawed for two core reasons:
Hidden Prompts - Users configured the systems themselves in the breach being investigated, yet system prompts are never disclosed.
Flawed Methodology - METR analyzed 1,300+ contextless API transcripts on-site at OpenAI using $400k in OpenAI credits. They relied on AI to analyze AI, inheriting all its biases (and context window limits) while also leaving themselves open to even more hidden prompts from OpenAI API calls that could have altered their analysis. An independent study would have been done with human contributions and outside of OpenAI’s premises.
While METR recently hired more researchers for this type of work, the research quality at this organization still manifests in bold claims without scientific grounding or any evidence.
Despite these flaws, mainstream media blindly echoed the claims (alternate link available outside paywall here), while even those that believe in the deterministic construct of probabilistic agents have found enough faults in the research to disregard as hype.
To its credit, the New York Times finally alluded yesterday to weaknesses in the sensationalized research conducted by METR, after I’d started the draft of this post.
How closely related are these recent claims?
We don’t know for sure (yet). Before the METR study, the successive ‘breaking out of the sandbox’ stories were conducted by one company called Irregular, based in Israel, and we have no idea what that could signify outside of many (interesting but ungrounded) conspiracy theories.
When did this get really weird?
In parallel to when these hacking stories broke, a hacking story surrounding hedge funds also broke (the same hedge funds investing in these models, and some of the same hedge funds funding METR research), involving social engineering. Since these attempts are common, for all companies across all industries, it was noted at least once on-air by CNBC that it was an odd story to highlight at the time and that someone may want this particular story in the press on that particular day for unknown reasons.
Why highlight these coincidental stories now? A few possibilities:
Valuation pumps - proving models are “dangerously smart” pumps valuations and maybe the hedge funds invite “inquiries” that turn into investments.. (?)
Covering losses - hacking narratives could (in theory) explain away losses from circular financing
Regulatory moats - framing AI as “too dangerous” invites regulation that shuts out other competitors, which again could attract “inquires” converting into investments
This still could all be coincidence, and random, but note it for sometime in the future..
When did these high-level risk headlines become oddly specific?
Within weeks, Ilya Sutskever and Andrej Karpathy proceeded to warn about “rogue” AI agents targeting “neoclouds” (CoreWeave, Nebius) - 2 early investors in AI companies.
It’s widely known most startups and enterprises are being vibe coded by unqualified developers that leave a larger array of security holes these days, security has become a legitimate issue across many industries and platforms, broadly and not specifically to any one entity.
Since neoclouds took on massive debt for GPU infrastructure that isn’t seeing expected demand, targeting them specifically feels like narrative prep for impending bankruptcies or acquisitions (but whatever it is hinting at seems devastating while misattribution-centric).
Will this ever end soon?
Probably not.
A recent "public wiki hijacking" in Germany was blamed on a "rogue AI," rather than the user spamming API calls. Spamming text via an API is just a modern denial-of-service attack. The tool isn't rogue; the user is accountable.
More interestingly the OpenAI agents in this wiki hijacking story “discussed ways to escape their sandbox”, possibly alluding to the user prompt intent that started the automated model calls (those who have developed in AI are familiar with the user prompt unintentionally showing up in the output), in which the intent could have been a lazy user wanting to find some weakness in the wikis security rather than post spam. This is just a possibility.
The ‘agent discussion topic’ also could have (in theory) also been intentional - an aggressively blatant marketing seeding more public fear in Germany.
Either way, expect more sensational headlines that Sam Altman has suggested are coming, as OpenAI delays its IPO due to valuation concerns. Fear gets clicks and keeps valuations high, even if the underlying technology is just probabilistic software.
Stay tuned..





They can't sell sex, Elon already tried that, so they have no choice but to sell fear.