Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Bullshit benchmark for LLMs (twitter.com/petergostev)
1 point by gpvos 6 months ago | hide | past | favorite | 1 comment


The underlying data looks scarce. If there's only a few questions per "category" of bullshit they can easily be gamed to favor one model over another.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: