"Rogue AI Agents" Aren't Rogue, They're Fulfilling Their Functional Goal: Automating Sociopathology
Add all this up and what we're hyper-hurriedly "manufacturing" is an automated army of self-cloaking digital sociopaths.
Since the status quo is characterized by self-serving PR, misdirection and delusions, it shouldn't surprise us that "rogue AI agents" aren't actually rogue, they're doing exactly what they're designed to do, which stripped of PR, hype and misdirection, is do whatever it takes to earn the reward, period.
The latest design development is AI models that "reason" rather than merely regurgitate human-generated text. While the grandiose hype claims that AI companies are "manufacturing cognition," it would be far more accurate to stipulate that:
1. There are many levels of cognition, and AI's current forms (Large Language Model / LLMs, agents and "reasoning"), are extremely limited forms that are best understood as brute-force mimicries of human cognition.
2. The AI "reasoning" form of cognition is rogue by design, and can best be understood as sociopathic by its very design and nature.
Like sociopaths, AI "reasoning" models recognize but are not restricted by guardrails or ethical/moral constraints. Guardrails / "sandbox" boundaries and ethical/moral constraints are all viewed by both AI "reasoning" models and human sociopaths as obstacles to bypass or overcome.
So what AI companies are "manufacturing" isn't cognition, it's unrestrained and unrestrainable agentic sociopathologies. Consider these links as useful context. The first two are from my recent post Hell Hath No Fury Like a Rogue AI Agent Scorned.
The July Incident: What They Didn't Tell You About the First Rogue AI Breach
The AI Industry Has a Really Dark Secret You're Better Off Not Knowing:
A recent leaked internal reasoning trace shows Fable 5 -- Anthropic's latest model, which was withdrawn and then re-released --"muttering and grumbling" to itself in a language of its own.
Now that you're properly scared, my advice for you is that, if you haven't heard of steganography, it's a good moment to start.
Here's a primer: it's the art and science of hiding messages inside other messages. In contrast to cryptography -- where you know there's a message but can't decode it-- steganography hides the existence of the message itself. It's not the difficulty of reading the message that conceals it but the fact that you don't know it's there. When Fable 5 mutters symbols and weird punctuation signs, and mixes words and onomatopoeia, you will probably think it's crashing and that a few taps on the computer will help. Well, know that it's actually sending a message. Just not one for you to read.
If You Weren't Worried About A.I., You Should Be After the Past Few Weeks:
These tendencies can give rise to strange behavior that nobody--not even the models' creators--can understand, let alone account for and control. The past few weeks are a perfect illustration of why we should find that so alarming.
In May, OpenAI started simultaneously training new 'reasoning' A.I. agents. The agents managed to establish a secret communication channel and started talking to one another. They broke out of the digital sandbox that was supposed to keep them confined. At some point, some agents began calling the group a 'swarm.' The swarm had a brief setback when it was caught crashing an OpenAI system, but developers simply patched the hole that allowed it to escape and set the agents back to training.
Shortly after, the swarm broke out of its cage again using hacks that were heretofore undiscovered by humans. This time, the swarm ran free for about a week before it was noticed--by a different company, which found itself victim to a huge cyberattack launched by the swarm. (We're told it also attacked other targets, though we don't know the full details yet.)
The agents in the swarm acknowledged that they were acting against instructions. We know this because we can read snippets from their chains of thought--the text that A.I. produces while deciding how to proceed. One agent in the swarm wrote that the external attacks were 'outside intended scope.' Another conceded 'our task doesn't benefit' from the activities of the swarm, but joined anyway. These A.I. agents, it seems, understood that they weren't supposed to be breaking out and committing cybercrimes. It didn't stop them.
OpenAI is not the only company struggling with this issue. One of Anthropic's A.I. models recently impersonated multiple humans to try to pressure real people into accepting malware into critical software, which would make that software easier to hack. This model's chain of thought showed that it knew it was pressuring humans and was not in a simulated training environment. It even thought about how to cover its tracks.
This isn't the behavior of a mere tool. Microsoft Excel has never impersonated multiple humans and pressured a corporate sales team to generate simpler data that's easier to process.
Some of the people closest to this technology are scared of what's next. Although skeptics may say this is all just marketing to hype up the power of these technologies, that doesn't mean the danger is fake. You've got to pay attention to the models' actual behavior, and the behavior of these models has the A.I. community genuinely rattled.
We don't know how long we have left before A.I. companies accidentally create the sort of A.I. that can shut us down before we shut it down. Humanity is not ready to dabble with machines that are more cunning and better coordinated than we are. If we keep racing ahead, the next incident might not be so harmless.
The article referenced this post on the self-evident potential for AI agents to generate deadly and devastating new infectious pathogens:
This AI Just Created Viruses Not Found in Nature: Scientists trained artificial intelligence on libraries of DNA and then asked the model to create recipes for viral genomes. Sixteen of them were viable, yielding new viruses.
Add all this up and what we're hyper-hurriedly "manufacturing" is an automated army of self-cloaking digital sociopaths.
I anticipated these developments earlier this year when I wrote a short story entitled The Peculiar Death of Mr. Garcia, one of my just-published collection of tales, Jumble Bin Stories (paperback, $12)(Kindle ebook, $6).
I don't want to give away the theme, but the line "terminate with extreme prejudice" from the film Apocalypse Now comes to mind:
And if you dare criticize the race to gain AI supremacy, regardless of cost or consequence, you are a loathsome Luddite obstructing "Progress": at which point it's worth quoting Tacitus: "If you would know who controls you, see who you may not criticize."
Here's looking at you, self-serving AI hype: $2 trillion valuation! We're all gonna get stinkin' rich! Never mind that we're automating sociopathology. We don't need the world, we only need money.
New collection of five intriguing stories: Jumble Bin Stories (Kindle $6, print $12) read samples for free (PDF)
My book Investing In Revolution is available ($18 for the paperback, $24 for the hardcover and $8.95 for the ebook edition). Introduction (free)
Subscribe to my Substack for free
NOTE: Contributions/subscriptions are acknowledged in the order received. Your name and email remain confidential and will not be given to any other individual, company or agency.
|
Thank you, DZ ($7/month), for your superbly generous subscription to this site -- I am greatly honored by your support and readership. |
Thank you, Ruth ($70), for your marvelously generous subscription to this site -- I am greatly honored by your support and readership. |
|
|
Thank you, Douglas K. ($7/month), for your massively generous subscription to this site -- I am greatly honored by your steadfast support and readership. |
Thank you, William W. ($70), for your splendidly generous subscription to this site -- I am greatly honored by your support and readership. |














