Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.


I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks.

I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do.

I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.


AI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.


What you’re saying is that you’re not interested in benchmarks. But then you go a step further and state that this particular benchmark is entirely inconsequential. That’s like telling you that if you didn’t exist, nothing would change. Even if that were true, it would still be an insensitive and rude thing to say, wouldn’t it?


It's okay to be rude to benchmarks though, they don't have feelings.


Yeah, I agree, I didn’t mean to impose on this conversation between a man and a benchmark, my bad :p


Usually when I say insensitive things there are more downvotes then upvotes. That isn't the case here. I might be rubbing against a truth somewhere here.


> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle

Please remember, we've started from there :

https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/

When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.


When is it ever hard to catch?


Sure, the task is not completed perfectly, but that's not the point.

Isn't it?

If the computer can't do it better than a human being, then what's the point?

Being wrong at scale is not better than being right.


Many humans would struggle with this even with very good tooling (ie not writing raw svg and using illustrator). I struggle to draw a bicycle accurately. But yes, I suspect it will be diminishing returns and I doubt it will ever be perfect due to the average nature of AI but I’d like to be wrong.


> Many humans would struggle with this even with very good tooling

But no ones hire random humans for things like this. You go and hire a vector artist and they will get your a very good pelican on a a bike. That's how you get things done when you can't do it.


> You go and hire a vector artist

Yeah, but then, you recruit the artist for $XXX - whereas you "recruit" your LLM for $0.XXX for the same task.

Of course the quality difference is huge. But sometimes you don't need that level of quality.

Also, finding a vector artist takes days of communication, payment settlement, revisions, etc.

Not always the most practical solution.


Hugely profitable companies leak half the nation's personal data every month. Tell me more about how being wrong at scale is not valuable.


So what if it is more profitable or valuable? It is still not better. Something being more profitable/valuable does not make it better, just like something being better does not make it more profitable/valuable. Sometimes, in some pursuits, for some outputs under some circumstances, the two are correlated. In others, the two are anti-correlated.


>If the computer can't do it better than a human being, then what's the point?

Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.


Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!"

This is what the tech industry has become?

Less of a failure is still failure.


I don't know how much money has been spent for AI, and I very much doubt you do. Do you know if more has been spent on LLM the last 9 years -- since "Attention is all you need" -- than Internet infrastructure during, say 1995 to 2004? That included the dotcom crash. Did you lament how a failure the Internet was?

LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?


I don't know how much money has been spent for AI, and I very much doubt you do.

Pick up a newspaper. Start with the Wall Street Journal. These are public companies. It's not a secret.


That is the software industry. Our product is less bugged than our competitor’s.


> If the computer can't do it better than a human being, then what's the point?

It can certainly do it better than I can. Sometimes you don't have a human handy with the required skills to do something.


Pelicans don't ride bicycles.

It's physically impossible.

The problem is to draw it in the least disturbing way possible.


True! But somehow Disney has been drawing ducks riding bikes in a way that seems to satisfy everyone since before my grandfather was born.

https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...


It’s better than many of the AI offerings and the bike could steer and the ducks are sitting on saddles, but the three nephews can’t reach the bottom of the pedal stroke and by the looks of their feet on the far side of the bike they aren’t trying. That shouldn’t satisfy Donald and Daisy, leaving them with all the work.


Wow, that was a highly relevant and specific website to source here!


I love these little websites with amazingly focused content. <3




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: