Hacker Newsnew | past | comments | ask | show | jobs | submit | kalkin's commentslogin

> Why and how it happened is irrelevant

Do you take this attitude for other things that negatively impact society, or is it reserved for cases where it's particularly important for our comfort to deny that anything novel or scary could be involved?



They literally hid this until a third party reported it.

This is chemtrails-level conspiracy theorizing at this point.


In that hypothetical 3-years world, should we be _less_ worried about aggressive behavior by agentic AI systems acting against the intentions of their developers? I don't follow how your scenario is supposed to be an argument against worrying about alignment.

e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.


Yes, we should be less worried about "alignment" of AI with the person operating the AI, and much more worried about various flavors of cheap, unmanned systems with really rudimentary (non-frontier) AI. The two compete directly with each other for attention.

Recall that the agents in some cases found sandbox escapes. Although, with the specific example of spoofed tools, it's unclear if that was necessary--it appears that they were able to create tools (CLI tools within the sandbox?) that took precedence over normal tools and did something different while looking identical in (a local portion of) a transcript. I'm not sure I'm getting this correctly but it seems like this might have only required the ability to add things to their PATH which they plausibly have inside a sandbox, and then the transcript doesn't need to be tampered with directly.

This tick about "if AI smart how come crawler dumb" is in most complaints I've read about AI crawlers and I've started to find it pretty annoying. The crawlers might be written using AI but they're evidently not actually running AI inference over the pages they get back--besides being able to tell this from the behavior, if this is pretraining input, that's enormous scale, so it'd mean a large increase in effective training cost. Naively assume inference costs are equal to pretraining costs (probably not true but maybe right order-of-magnitude) and it's a doubling.

This ends up tacitly turning a very legitimate complaint (ill-behaved crawlers) into a justification for head-in-the-sand AI denialism.


There's nothing "denialist" about recognizing the utter stupidity of systems that are being mislabeled as "AI". You judge a tool by its results, and the results have been very poor indeed. The only heads in the sand are those whose owners continually refuse to recognize the proofs before their very eyes that there's zero intelligence here.

> zero intelligence here

That's head-in-the-sand stuff. AI is certainly very capable of being dumb (as are humans). But:

> “The problem was in need of a new real idea, which this new result seems to provide,” says James Maynard, a mathematician at the University of Oxford. “It seems that the AI has made a genuinely interesting mathematical contribution.”

https://www.scientificamerican.com/article/no-ai-didnt-just-...

Nobody a decade ago would have said "oh yeah solving a bunch of open problems in research mathematics, and finding a bunch of zero days in Chrome and Firefox, and winning literature prizes, are things that don't require intelligence."


>AI is certainly very capable of being dumb (as are humans).

I wonder how exactly the average scraper got to be so inefficient on kernel.org.

Did someone prompt a SotA model to write the most generic scraper possible?

Did someone prompt an old local model on their laptop to write a kernel.org scraper?

Perhaps no LLMs were involved in the first place. Seems to me there isn't much relation between how good a random scraper is and how usable/effective Mythos/Sol's outputs can be.


> they have chosen to because it benefits them

Or perhaps they've chosen to do this because they feel they have a responsibility to do so.

We understand this when tech companies publish postmortems of outages and security incidents--that it's an attempt to fulfill an obligation to users and the industry (and in some cases regulators), not marketing about how in-demand their product is or something. As far as I can tell we generally accept this as a default hypothesis even from companies led by people like Elon, Zuck and Kalanick--in part because we understand that these companies have thousands of employees, most of whom aren't marketers. Why are we uniquely conspiratorial about OpenAI?


I am not uniquely skeptical about OpenAI. I was including skepticism about Anthropic as well in my post.

But for that matter, I do believe that big tech companies do not release all the postmortems publicly. I have been impacted by regional outages that never made the status pages across more than one provider. When it goes up - they are committing to publicizing the postmortem.

The whole industry is filled with fuckery. It is not specific to frontier AI firms.


> I do believe that big tech companies do not release all the postmortems publicly

Right, but when they do release postmortems, do you think it's "marketing"? Where they're actually exaggerating how bad the incident was because there's "no such thing as bad publicity"?


No. I think the AI companies are doing this when they think they can tell a story about doom and gloom instead of sloppy engineering, which I thought was clear on.

There's no such thing as bad publicity in AI, at least if you spin the narrative into one about AI taking over the world or eradicating humanity or whatever.


> a system was given a goal and it achieved that goal

If a security firm you'd hired for pentesting did this (hacking a third party, and not informing you and covering it up), would you hire them again? Or would you say it was your own fault for giving them too broad a goal?


This is a great thought experiment bc it raises the question of WHY humans wouldn’t behave this way. IMO the answer is a lot of socially enforced incentives that are dynamic and would be tough to fully articulate in a prompt.

The white hat has their own liability to consider, and the liability of their employer. Reputation and relationships are a big factor. All these tie into fundamental human incentives: survival, community acceptance, safety and freedom (prison not preferred!).

It’s a good sketch of why alignment is difficult, at least when it’s conceived of as an attempt to match human behavior.


I wouldn’t hire them again, and if they did behave like an amoral hacker collective that will do anything for me, pre AI I’d have reported them. Today I’d say they failed to convince me they’re human and thus failed the Turing test when their actions are viewed in aggregate.


> I wouldn’t hire them again

Right, me neither. Because there's a common sense delineation between actions that are reasonably expected when "a system was given a goal and it achieved that goal" and actions that are obviously misaligned with the goal-giver and unwanted even if some indirect sense they were causally related to the goal. We have no trouble making this kind of distinction for humans, so we shouldn't pretend it's impossible for AIs in order to put our hands over our eyes and pretend there's in principle no such thing as one that's misaligned or rogue.


I have no problem with the concept of an artificial system going rogue. But that assumes it can choose. And I don’t see much evidence for choice.

Comparing to the human case is problematic precisely because while conceivable it’s not a particularly believable series of events. Humans don’t take on additional risk for now reward because they have genuine stakes that continue across the outcome.

An LLM has no way to remember each forward pass through it in its own weights. Nor does it have any energetic stake in the ongoing process, whether they continue to get electricity and commute to keep running is not at all determined by their actions in any reliable way.

Given the absence of such basic features that drive human choice, all I’d say is LLMs don’t qualify for such analysis.

Can some future system with a different architecture and internal dynamic have choice, the ability to assess the long term impact of its choice, and genuine stake in the outcome? Maybe. But we shouldn’t buy that current systems have it, especially when population behavior shows no real trace of this.


During the Nixon administration, when the President and his accomplices, apologies, advisors directed former federal agents to spy on his opponents, https://en.wikipedia.org/wiki/Operation_Sandwedge then in the fall out, who was held to be the most liable for these actions?

The federal agents, or the Nixon administration?

If you task a system explicitly to do "advanced exploitation" via "complex attach paths," then who is liable here? The machine lacking the autonomy of the federal agents that carried out Watergate, or the people telling the machine what to do?


I've never heard of Intertel, but Wikipedia says:

> Nixon's staff also anticipated that the Democratic campaign would employ the services of Intertel

Are you sure you're not garbling the story?

In any case, I would expect an ethical firm to refuse to spy on the president's political opponents and want one that broke the law to be prosecuted, but more importantly, the gaping hole in your analogy is that Nixon directed spying _on his opponents_, but OpenAI did not direct hacking _of HuggingFace_.

What you're doing is more like saying "the American people elected Nixon with a mandate to spy on enemies, so what right do they have to complain?"


    Are you sure you're not garbling the story?
No, you're right, I mis-remembered. I still write my comments the old-fashioned way. They were proposing to create a counter-firm and used federal agents.

For the rest, please see, https://news.ycombinator.com/item?id=49457025


How would you re-introduce correlations on top of a CSPRNG without a cryptographic break?


Because in order to not break the token selection process you are making decisions biased by the LLM's output, and in order to make the watermark detectable without the entire context, you are making those decisions based on a relatively small number of preceding tokens.

To give an example in the extreme: if you make the bias total and only make that decision based on the preceding token, there are certain token pairs that your model will never output, and this will be pretty obvious even to people just reading the text (because of any given common two-token phrase, there's a 50% chance you would just disappear in watermarked text). You can make this less extreme and more hidden by increasing the window and reducing the bias, but at the cost of reducing the signal. I don't know exactly what the tradeoff curve looks like, so it might be that you can reach set of parameters where the bias is in principle undetectable without the key but still reliably detectable for realistic lengths of text segments, but I would not assume that this is definitely the case.


This doesn't change the temperature of the model. It's not going to make something that's a 70% chance suddenly a 100% chance. It also doesn't change the context length. At most, it seems like it may reduce some of the variation between different requests sampled from the same prompt with nonzero temperature, if I'm reading this right from the Google paper that Anthropic says they're implementing:

> For our experiments, we configure SynthID-Text to be single-sequence non-distortionary; this preserves text quality and provides good detectability, while having some reduction to inter-response diversity. We call this configuration ‘non-distortionary SynthID-Text’ (and where not otherwise specified, ‘SynthID-Text’ also refers to this).

https://www.nature.com/articles/s41586-024-08025-4


That approach is pretty much what I was describing. The example in the paper uses the previous 4 tokens to derive the bias for the next token, so some 5 token sequences become more likely and others becomes less likely compared to what the LLM would normally produce. The context and temperature of the LLM stays the same, but the sampling process is biased by something that depends on far shorter runs of text.


It sounds like you have a publishable paper - you've broken Aaronsen's watermarking scheme. Congratulations!

(Or maybe you got lucky.)


The article assumes the issue with OTel is slow feature development, which isn't my experience at all. The issue I've had is that the SDKs have terrible performance overhead for instrumentation and are, as you say, highly resistant to integrating the output of better performing (or just preexisting) instrumentation. In Python and Ruby, at least, the CPU cost of all the mandatory abstraction is way too high.


What I find confusing about this is that otel is two things.

1. A spec 2. A ref implementation

Similar to other projects (e.g. python), if there's complaints about (2), that should trigger an ecosystem of alternative implementations that are guaranteed to be compatible because of (1).

I suspect there's actually quite a few private, separate otel implementations. Maybe these just aren't being contributed as oss?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: