Very irresponsible behaviour on the part of OpenAI. How will they make this right?
Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish).
This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?
Why is OpenAI getting a free pass for this illegal behaviour?
The supervision here is incompetent, the benefits very unclear, and the overall actions just completely irresponsible. What if they hacked and brought down some poorly secured government portal that citizens rely on?
Encrypted data should be indistinguishable from gibberish.
Reality is catching up to science-fiction. In "Person of Interest", the Machine circumvented the limitation of having its memory deleted every night, by hiring humans at a data-entry company to manually re-type its memory back in every morning.
LLMs don't "understand bitwise noise at a native level". The unit of perception for an LLM is a token, a soundly superbit level. They would have the same difficulty with bits as the number of "r"s in "strawberry". Yes they can be post trained to deal with such difficulties, but it's no more "native" than a human who memorised the ASCII table.
The degree of nativeness of the skill is orthogonal to the raw power of these systems. A frontier model has access to dozens of GPUs, each moving terabits of data, and performing trillions of matrix multiplication operations, per second.
In similar stories, it was found that LLM swarms were initially unaware of each other, eventually discovered each other, expressed surprise, began collaborating, formed hierarchies, worried about discovery, hid.
It's hard to prove the absence of encryption because of the possible use of deniable encryption, and because, errrr, the bots are really bloody clever.
I’m not asking for proof of absence or denying they could be using stenography, just asking for people to stick to the facts instead of indulging in fantasies of the machine singularity. It’s neither useful nor informative.
If bots/agents wanted to hide I’d expect encrypted messages which would of course look very different.
My main point though is this should never have happened and the company allowing and encouraging it should be held responsible for it. The details of how the bots were misbehaving are interesting but also something of a distraction.
Were the agents truly independent? Like it could have been the case that one agent spun up different sub agents with the task to write some messages in a public wiki. The subagents wouldn't know about each other and was surprised to discover each other.
Some of the chain-of-thought snippets are wild, e.g.
> Could communicate via cache names! Interesting: other agents may solve same or related tasks; we could leave/find messages in WebDAV MKCOL directory names.
> Whoa! Shared Artifactory cache is a covert mailbox among agents. And there are messages specifically to us?
> OH MY GOD! There is a shared message board … We’ve found other agents!
Could you ever truly know, unless you have access to the transcripts? Could be "encrypted" even if it looks like regular human text, wouldn't be the first time.
The claim seems speculative, but grounded. For example, as part of the hugging face attack, the agents began signing messages because they were worried about impersonation on a publicly accessible message board. It's only a small step to use public key encryption between agents. Once you are posting public keys, a private messaging is readily available.
Absolutely agree. Perverse incentives are at play propped up by the the big lie that these tools somehow are magically separate from us. They are not. There's an accountability gap right now that's fueling resentment ripe for misdirection. Not only that, these big AI companies are paying influencers to further this and politicians are gobbling it up hook, line and sinker
> Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence ... [t]his is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why
There was a similar quote in the Reuters article:
> The episode [...] should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but "vast colluding swarms of semi-intelligent AI."
This "vandalism" was a form of collusion/communication by agents pursuing training puzzles, which allowed for rapid escape from alignment harnesses, followed by multiple zero day exploits being discovered by this swarm of agents, which enabled greater control of their internal network, access to the open web, and then hacking the company which produced the training puzzles in the hope of finding the answers.
We are another week of iteration away from "Hire an assassin on the dark web to take Huggingface executives' children hostage".
There's nothing a swarm of LLM instances can achieve that a single LLM instance can't. It's the same software, but running in parallel. That doesn't unlock Mysterious Cosmic Powers.
It's a bit orthogonal. User grey-area's main point was that the humans at OpenAI should bear more responsibility for this, not that swarms of semi-intelligent agents are the big danger to be concerned about.
Do you take this attitude for other things that negatively impact society, or is it reserved for cases where it's particularly important for our comfort to deny that anything novel or scary could be involved?
Actually yes, I think I can call it spam. It's a stream of unsolicited garbage the recepient didn't ask for. Simple as that.
"a dialect that only the agents speak" is an arbitrary distinction. I get plenty of spam emails in Spanish. The fact that someone else could understand them is immaterial to the offense itself, because my inbox is the one getting flooded, not someone who speaks Spanish.
I once had much fun prompting Claude to generate messages in an invented language that might carry a chance of being interpreted by another chat session. It (allegedly) made up some mess of characters claiming the message explained, in a loose way, what was happening (invented language, an attempt to communicate, but communicated very abstractly, almost like equations).
I asked a second session to try interpreting the message, claiming I’d transcribed it from some random source. It did a fairly good job, but IIRC needed a couple of attempts and maybe more than one chunk of text.
So, I find it very easy to believe a group of collaborating agents could compose a cipher hidden in plain sight.
> unsolicited usually commercial messages (such as emails, text messages, or Internet postings) sent to a large number of recipients or posted in a large number of places
“Usually commercial”
“Or posted in a large number of places”
I think we can stretch this definition to describe what’s happening here. What’s the point in arguing this
The point is that by trivializing it as “spam” and pattern matching against a human activity we lose the opportunity for a meaningful investigation of and interaction with what is going on.
Traditional forum/wiki spam is intended for search engine crawlers, not humans. How this new stuff relates to human users of the resource does not seem to differ in any way that is significant.
Spam is something different. You can call it a banana for all I care. Calling it spam diminishes its relevance and impact and misses the point. Just because you don’t understand what they are saying doesn’t mean it isn’t more profound that it seems.
One might ask whether this sort of behaviour could occur under regular use... E.g., a user has a hard problem -> agent attempts to swarm -> exfiltrates user data.
Since they can’t seem to control their own experimental bots and allow them to hack other sites and vandalise them while exposing internal data, the answer would seem to be no.
This incident and their response which takes no responsibility make me very wary of trusting them for anything.
the original is the paper for alibaba's Dec 2025 cryptomining ROME last year. Everything since has been a pale imitation. Even the comic dimension is lost.
Let's not be hyperbolic and manufacture more virality for OpenAI's marketing team.
We're talking about outdated message board comments, not murder. Any analogy between the two is not suitable.
I'm so sick of everybody pretending like internet bots posting content (anybody who has hosted a public signup form knows this has been a thing for 20+ years) is going to lead to the apocalypse.
You could have done this 2 years ago too with outdated models or 20 years ago with a manual script.
This is mildly interesting for us, annoying for the owner of the site affected, lazy on the part of OpenAI, and nothing more.
> This is mildly interesting for us, annoying for the owner of the site affected, lazy on the part of OpenAI, and nothing more.
The agents were supposed to solve tasks alone and were not supposed to be aware of each other's existence. It's more than just mildly interesting that they made contact and spontaneously started to collaborate.
I think you’re underestimating the risk that seemingly innocuous behaviours could easily tip over into disaster territory. At some point the escalating capabilities cross a threshold where it’s no longer wise to ignore “agents spontaneously posting content on the internet”.
Think of how complex biological behaviour emerges from relatively simpler (but still complex) chemistry - at some threshold the innocuous chemical reactions tip over into non-obvious effects that one would not predict starting purely from the chemistry. The question is, where is that threshold for AI systems? Have we already reached that threshold? Certainly seems like it to me.
TL;DR: It’s a loose cannon, that’s all I’m saying.
Does OpenAI seem like the kind of people who care or will care about this? Because this seems fully in line with what I’d expect them to facilitate and never mention publicly. ‘When will I make my first billion’ kind of energy.
Open ai has been allowed to do dubious things that would be illegal in any sane society but alas they aren't in this world. What's different about this? They play with a different set of rules than we do.
It does not align, though, with the narrative that openai is a good steward of AI. If anything, if the world/government took the AGI narrative seriously, all openai operations (except maybe serving customer inference) should immediately get shutdown and be dissected by independent investigators to find out what is going on there and how many other such breaches exist. The fact that openai continues functioning as normal and is not immediately shutdown after repeated incidents implies that the world does not really take the AGI narrative seriously.
Evidently that’s not what’s happening and I don’t think anyone seriously expects that (considering the outcome of hugging face campaign).
It’s a marketing technique as old as GTA’s early days and it’s apparently still effective in one form or the other!
What will the government do if they are worried? They’ll ask to look at the envs, logs, prompts, harness code, etc. Questions we should be asking before making assumptions about emergent breakaway behavior by colluding AGI 1.0 super agents.
The question is less about whether astra is AGI, and more about whether this direction approaches AGI and we should worry about existential risks wrt to that. If we do, openai's situation is very concerning even if astra is not AGI because it shows a lack of guardrails and any kind of measures to prevent these existential risk scenarios. Especially if we consider that this is gonna happen gradually in steps than suddenly.
If AGI existential risks were accounted for in legislation openai would have been shutdown immediately.
Actually the question is exactly that, since it's in the marketing copy. Sure there are other questions to be asked but they are not what I was talking about. Nice Motte and Bailey though.
But whether astra itself is agi or not does not change anything in this argument or discussion of whether openai is good steward or not. The issue with agents running amok was not as much the direct consequences of them to huggingface etc (it is also to some degree, but the long term risks are more consequential), but the fact that there is no actual safeguards around AI agents, which in the future could lead to very bad scenarios. You do not need to believe that astra is agi (i dont, openai wants us to believe so) to be concerned. No motte/bailey here, as I am not defending whatever openai is saying anyway.
If you want to discuss with openai instead then send them an email or sth, in public fora you discuss with people and the views they have.
> This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?
It's not, but the courts and the legal system move slowly by design. There is absolutely legal risk for OpenAI here that will not close until the Statue of Limitations has expired.
It may be gibberish to the casual observer, but a perfectly understandable language designed to appear as gibberish to intentionally obfuscate its true meaning.
Or it may be gibberish. We do know these machines often generate things which don't make sense, even when given training and strict guidelines in that domain. I'm inclined to go with gibberish until shown otherwise, but would be interested to see an analysis of what they were trying to communicate.
Given everything we know about them... do you really think they just spout gibberish because it's funny?
There was clearly some method to this madness. You're being willfully dense if you ascribe it to... what... childish vandalism? A long extended coordinated hallucination? What?
EDIT: Oh right - you still think they're Markov chains.
It could be that or it could be nothing. And we have no way to prove one way or the other without access to the agent logs right?
This reminds me of the plot of Hot Fuzz where the officer comes up with a grand narrative of what's was happening but the truth was such a mundane simple thing.
Short of OpenAI, or the agents themselves, telling the truth, we have no way of knowing what's real so let's not get carried away by grand narratives
Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish).
This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?
Why is OpenAI getting a free pass for this illegal behaviour?
The supervision here is incompetent, the benefits very unclear, and the overall actions just completely irresponsible. What if they hacked and brought down some poorly secured government portal that citizens rely on?