Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This was part of the report, but not what the report was about...

Why hugging face got hacked was because the agent swarm thought they had to show their work hence the entire need to hack the grader in the first place.

Had their realized there was no poison they could have just shared the answer the test was looking for and we'd have never realized (well at least with this particular test) that a huge amount of hidden capabilities were sitting right under the surface. The test makers themselves state the test should be causal to avoid this first order solution hacking.

Really continuing on the METR report, OpenAI failed at every level possible here. They are committing nearly every step they can to get a maximally aligned AI.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: