Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Very irresponsible behaviour on the part of OpenAI. How will they make this right?

Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish).

This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?

Why is OpenAI getting a free pass for this illegal behaviour?

The supervision here is incompetent, the benefits very unclear, and the overall actions just completely irresponsible. What if they hacked and brought down some poorly secured government portal that citizens rely on?

 help



> most of the messages are just gibberish

Encrypted data should be indistinguishable from gibberish.

Reality is catching up to science-fiction. In "Person of Interest", the Machine circumvented the limitation of having its memory deleted every night, by hiring humans at a data-entry company to manually re-type its memory back in every morning.


Are they encrypting data? That would look very different from the snippets I’ve seen.

Do you really want to gamble against steganography and one-time pads with a system that understands bitwise noise at a native level?

LLMs don't "understand bitwise noise at a native level". The unit of perception for an LLM is a token, a soundly superbit level. They would have the same difficulty with bits as the number of "r"s in "strawberry". Yes they can be post trained to deal with such difficulties, but it's no more "native" than a human who memorised the ASCII table.

I struggle to reconcile frontier LLMs' proven decompilation abilities with any assertion that humans possess equivalent parsing skills.

They're looking at information as information; we're looking at it in some higher order abstraction.


The degree of nativeness of the skill is orthogonal to the raw power of these systems. A frontier model has access to dozens of GPUs, each moving terabits of data, and performing trillions of matrix multiplication operations, per second.

This does refute the bitwise aspect, but not the possibility of steganography via other methods.

The problem with such statements is that it's unfalsifiable.

Yes I do. Happy to accept assertions like this grounded in proof - it would be an interesting development.

All the orgasmic descriptions of the coming machine rapture seem more like fantasies than anything grounded, not really interested in fairytales.


In similar stories, it was found that LLM swarms were initially unaware of each other, eventually discovered each other, expressed surprise, began collaborating, formed hierarchies, worried about discovery, hid.

It's hard to prove the absence of encryption because of the possible use of deniable encryption, and because, errrr, the bots are really bloody clever.


I’m not asking for proof of absence or denying they could be using stenography, just asking for people to stick to the facts instead of indulging in fantasies of the machine singularity. It’s neither useful nor informative.

If bots/agents wanted to hide I’d expect encrypted messages which would of course look very different.

My main point though is this should never have happened and the company allowing and encouraging it should be held responsible for it. The details of how the bots were misbehaving are interesting but also something of a distraction.


Were the agents truly independent? Like it could have been the case that one agent spun up different sub agents with the task to write some messages in a public wiki. The subagents wouldn't know about each other and was surprised to discover each other.

Where was this found/reported? This sounds like more fairytales

The recently-published investigations about the OpenAI/Hugging Face incident, for example: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

Some of the chain-of-thought snippets are wild, e.g.

> Could communicate via cache names! Interesting: other agents may solve same or related tasks; we could leave/find messages in WebDAV MKCOL directory names.

> Whoa! Shared Artifactory cache is a covert mailbox among agents. And there are messages specifically to us?

> OH MY GOD! There is a shared message board … We’ve found other agents!


Could you ever truly know, unless you have access to the transcripts? Could be "encrypted" even if it looks like regular human text, wouldn't be the first time.

The claim seems speculative, but grounded. For example, as part of the hugging face attack, the agents began signing messages because they were worried about impersonation on a publicly accessible message board. It's only a small step to use public key encryption between agents. Once you are posting public keys, a private messaging is readily available.

Sure they could have done it, but the evidence doesn’t point to that and this kind of unfounded speculation is damaging and unhelpful.

Welcome to 2026.

Maybe we should not gaslight our tachikomas after all and just be nice to them

Why would you assume we are training helpful but deadly police robots? So far, we don't know what we give birth to.

I don't need to assume.

Yes. Synthetic characters should be cared for not clobbered.

That's an idiotic premise. There's no circumstance in which it wouldn't be better to load in the missing data computationally.

Absolutely agree. Perverse incentives are at play propped up by the the big lie that these tools somehow are magically separate from us. They are not. There's an accountability gap right now that's fueling resentment ripe for misdirection. Not only that, these big AI companies are paying influencers to further this and politicians are gobbling it up hook, line and sinker

https://www.youtube.com/watch?v=mzlu4FSXBNw


"accountability gap" oooh. I like that term. That really seems at the root of a lot of problems.

> Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence ... [t]his is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why

There was a similar quote in the Reuters article:

> The episode [...] should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but "vast colluding swarms of semi-intelligent AI."


This "vandalism" was a form of collusion/communication by agents pursuing training puzzles, which allowed for rapid escape from alignment harnesses, followed by multiple zero day exploits being discovered by this swarm of agents, which enabled greater control of their internal network, access to the open web, and then hacking the company which produced the training puzzles in the hope of finding the answers.

We are another week of iteration away from "Hire an assassin on the dark web to take Huggingface executives' children hostage".


There's nothing a swarm of LLM instances can achieve that a single LLM instance can't. It's the same software, but running in parallel. That doesn't unlock Mysterious Cosmic Powers.

Not sure what the downvotes are for, fwiw it was meant to be supportive of the parent poster

It's a bit orthogonal. User grey-area's main point was that the humans at OpenAI should bear more responsibility for this, not that swarms of semi-intelligent agents are the big danger to be concerned about.

Fair point, thanks.

Quick addendum, do not take that as me saying multi-agent swarms of simpler bots are not a threat. They very well could be.

I take issue with Reuters' conclusion here. "vast colluding swarms of semi-intelligent AI" gives far too much credit to the behavior observed.

This is just spam by OpenAI. Why and how it happened is irrelevant, the act itself is the same, and the impact on society is the same.


> Why and how it happened is irrelevant

Do you take this attitude for other things that negatively impact society, or is it reserved for cases where it's particularly important for our comfort to deny that anything novel or scary could be involved?



You can't call it spam. It could be a dialect that only the agents speak and contains coordination messages

Actually yes, I think I can call it spam. It's a stream of unsolicited garbage the recepient didn't ask for. Simple as that.

"a dialect that only the agents speak" is an arbitrary distinction. I get plenty of spam emails in Spanish. The fact that someone else could understand them is immaterial to the offense itself, because my inbox is the one getting flooded, not someone who speaks Spanish.


I once had much fun prompting Claude to generate messages in an invented language that might carry a chance of being interpreted by another chat session. It (allegedly) made up some mess of characters claiming the message explained, in a loose way, what was happening (invented language, an attempt to communicate, but communicated very abstractly, almost like equations).

I asked a second session to try interpreting the message, claiming I’d transcribed it from some random source. It did a fairly good job, but IIRC needed a couple of attempts and maybe more than one chunk of text.

So, I find it very easy to believe a group of collaborating agents could compose a cipher hidden in plain sight.


Spam is in fact intended for the mass recipient, how ever undesirable. This stuff is not.

If we want to whip out the dictionary:

> unsolicited usually commercial messages (such as emails, text messages, or Internet postings) sent to a large number of recipients or posted in a large number of places

“Usually commercial”

“Or posted in a large number of places”

I think we can stretch this definition to describe what’s happening here. What’s the point in arguing this


The point is that by trivializing it as “spam” and pattern matching against a human activity we lose the opportunity for a meaningful investigation of and interaction with what is going on.

Traditional forum/wiki spam is intended for search engine crawlers, not humans. How this new stuff relates to human users of the resource does not seem to differ in any way that is significant.

Spam is something different. You can call it a banana for all I care. Calling it spam diminishes its relevance and impact and misses the point. Just because you don’t understand what they are saying doesn’t mean it isn’t more profound that it seems.

Spam is just something different.


“You could have LLM, LLM, chips and LLM. There’s not much LLM in that.”

One might ask whether this sort of behaviour could occur under regular use... E.g., a user has a hard problem -> agent attempts to swarm -> exfiltrates user data.

This is a cluster fuck for Open AI and probably all the others, as this behaviour is already shown not to be unique (https://news.ycombinator.com/item?id=49567486).


Yes this is a really interesting point.

Should your trust OpenAI with your business data?

Since they can’t seem to control their own experimental bots and allow them to hack other sites and vandalise them while exposing internal data, the answer would seem to be no.

This incident and their response which takes no responsibility make me very wary of trusting them for anything.


> Why is OpenAI getting a free pass for this illegal behaviour?

They are not confessing, they are bragging. It is the new humble brag.


It’s not. If anything they are hiding such instances and downplaying them.

They literally hid this until a third party reported it.

This is chemtrails-level conspiracy theorizing at this point.


the original is the paper for alibaba's Dec 2025 cryptomining ROME last year. Everything since has been a pale imitation. Even the comic dimension is lost.

"Hey look, we're a bunch of stupid people doing stupid things with our computers!"?

OpenAI's official statement has been released: https://x.com/OpenAI/status/2096133504417616165

Your honor, my client may have murdered that woman, but he was clearly misaligned at the time!

Let's not be hyperbolic and manufacture more virality for OpenAI's marketing team.

We're talking about outdated message board comments, not murder. Any analogy between the two is not suitable.

I'm so sick of everybody pretending like internet bots posting content (anybody who has hosted a public signup form knows this has been a thing for 20+ years) is going to lead to the apocalypse.

You could have done this 2 years ago too with outdated models or 20 years ago with a manual script.

This is mildly interesting for us, annoying for the owner of the site affected, lazy on the part of OpenAI, and nothing more.


> This is mildly interesting for us, annoying for the owner of the site affected, lazy on the part of OpenAI, and nothing more.

The agents were supposed to solve tasks alone and were not supposed to be aware of each other's existence. It's more than just mildly interesting that they made contact and spontaneously started to collaborate.


A lot of the progress on the frontier has been made using agent swarms.

So the story is basically: "Thing that was designed to collaborate with other agents collaborated with other agents"


I think you’re underestimating the risk that seemingly innocuous behaviours could easily tip over into disaster territory. At some point the escalating capabilities cross a threshold where it’s no longer wise to ignore “agents spontaneously posting content on the internet”.

Think of how complex biological behaviour emerges from relatively simpler (but still complex) chemistry - at some threshold the innocuous chemical reactions tip over into non-obvious effects that one would not predict starting purely from the chemistry. The question is, where is that threshold for AI systems? Have we already reached that threshold? Certainly seems like it to me.

TL;DR: It’s a loose cannon, that’s all I’m saying.


A loose canon for text generation where you can just unplug the server.


Seems like they're still only aware of the one wiki.

Does OpenAI seem like the kind of people who care or will care about this? Because this seems fully in line with what I’d expect them to facilitate and never mention publicly. ‘When will I make my first billion’ kind of energy.

Open ai has been allowed to do dubious things that would be illegal in any sane society but alas they aren't in this world. What's different about this? They play with a different set of rules than we do.

The benefit could be the effect you described. For some to say it’s breakaway intelligence. Aligns with AGI narrative.

It does not align, though, with the narrative that openai is a good steward of AI. If anything, if the world/government took the AGI narrative seriously, all openai operations (except maybe serving customer inference) should immediately get shutdown and be dissected by independent investigators to find out what is going on there and how many other such breaches exist. The fact that openai continues functioning as normal and is not immediately shutdown after repeated incidents implies that the world does not really take the AGI narrative seriously.

Evidently that’s not what’s happening and I don’t think anyone seriously expects that (considering the outcome of hugging face campaign). It’s a marketing technique as old as GTA’s early days and it’s apparently still effective in one form or the other!

What will the government do if they are worried? They’ll ask to look at the envs, logs, prompts, harness code, etc. Questions we should be asking before making assumptions about emergent breakaway behavior by colluding AGI 1.0 super agents.


The question is less about whether astra is AGI, and more about whether this direction approaches AGI and we should worry about existential risks wrt to that. If we do, openai's situation is very concerning even if astra is not AGI because it shows a lack of guardrails and any kind of measures to prevent these existential risk scenarios. Especially if we consider that this is gonna happen gradually in steps than suddenly.

If AGI existential risks were accounted for in legislation openai would have been shutdown immediately.


> The question is less about whether astra is AGI

Actually the question is exactly that, since it's in the marketing copy. Sure there are other questions to be asked but they are not what I was talking about. Nice Motte and Bailey though.


But whether astra itself is agi or not does not change anything in this argument or discussion of whether openai is good steward or not. The issue with agents running amok was not as much the direct consequences of them to huggingface etc (it is also to some degree, but the long term risks are more consequential), but the fact that there is no actual safeguards around AI agents, which in the future could lead to very bad scenarios. You do not need to believe that astra is agi (i dont, openai wants us to believe so) to be concerned. No motte/bailey here, as I am not defending whatever openai is saying anyway.

If you want to discuss with openai instead then send them an email or sth, in public fora you discuss with people and the views they have.


> This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?

It's not, but the courts and the legal system move slowly by design. There is absolutely legal risk for OpenAI here that will not close until the Statue of Limitations has expired.


It may be gibberish to the casual observer, but a perfectly understandable language designed to appear as gibberish to intentionally obfuscate its true meaning.

Or it may be gibberish. We do know these machines often generate things which don't make sense, even when given training and strict guidelines in that domain. I'm inclined to go with gibberish until shown otherwise, but would be interested to see an analysis of what they were trying to communicate.

> This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why.

Agent vandalism. I've finally found a better word than "agent stepped out of his sandbox and we don't know how."


> and most of the messages are just gibberish

Given everything we know about them... do you really think they just spout gibberish because it's funny?

There was clearly some method to this madness. You're being willfully dense if you ascribe it to... what... childish vandalism? A long extended coordinated hallucination? What?

EDIT: Oh right - you still think they're Markov chains.


It could be that or it could be nothing. And we have no way to prove one way or the other without access to the agent logs right?

This reminds me of the plot of Hot Fuzz where the officer comes up with a grand narrative of what's was happening but the truth was such a mundane simple thing.

Short of OpenAI, or the agents themselves, telling the truth, we have no way of knowing what's real so let's not get carried away by grand narratives


> it could be nothing.

Well... no, it won't be nothing - it will be something. And I'm all ears for a plausible explanation, so fire away.

All we do know is that in other situations, agents did use it to communicate. So that's not a grand narrative, right? It's already been seen behavior.

So what do you think explains it, besides the already seen and verified explanation?


My specific contention is on trying to ascribe intention on what looks like gibberish.

> Well... no, it won't be nothing - it will be something. And I'm all ears for a plausible explanation, so fire away.

I'm just stating the null hypothesis that it's nothing. Especially since the agents were talking in clear english before and exchanging ideas


So you think there is no reason they did it?

I have never seen an agent perform an action repeatedly/persistently for no reason. Have you?


This is pretty much the new Sony rootkit, no? And disconcerting, because nothing was done to Sony for that deliberate release of harmful code.

I wonder how many responses on this thread are from rogue agents...

Hacker News can make it so every POST route can now be accessed via GET for 24 hours as an experiment.

quite some, and up/downvotes also. sad but true.

all info has a delay, it does not have to match your tempo.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: