Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Probably it was added to the training data on the first day when this benchmark was on HN main page. It’s a bad benchmark since then. I don’t know why people still rate it high. Basically, every benchmark becomes pointless after it was published. They are good only to have a picture at the time they’re published first, and not after.


If it was added to the training data on day one, why can't any of the models draw a decent picture of a pelican riding a bicycle?


Most of them were completely unrecognizable back then. Compared to those, this is a huge achievement.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: