I don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.
I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks.
I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do.
I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.
AI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.
What you’re saying is that you’re not interested in benchmarks. But then you go a step further and state that this particular benchmark is entirely inconsequential. That’s like telling you that if you didn’t exist, nothing would change. Even if that were true, it would still be an insensitive and rude thing to say, wouldn’t it?
Usually when I say insensitive things there are more downvotes then upvotes. That isn't the case here. I might be rubbing against a truth somewhere here.
When it started, it was clear what LLM would stand out, its style, etc.
Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.
Many humans would struggle with this even with very good tooling (ie not writing raw svg and using illustrator). I struggle to draw a bicycle accurately. But yes, I suspect it will be diminishing returns and I doubt it will ever be perfect due to the average nature of AI but I’d like to be wrong.
> Many humans would struggle with this even with very good tooling
But no ones hire random humans for things like this. You go and hire a vector artist and they will get your a very good pelican on a a bike. That's how you get things done when you can't do it.
So what if it is more profitable or valuable? It is still not better. Something being more profitable/valuable does not make it better, just like something being better does not make it more profitable/valuable. Sometimes, in some pursuits, for some outputs under some circumstances, the two are correlated. In others, the two are anti-correlated.
>If the computer can't do it better than a human being, then what's the point?
Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
I don't know how much money has been spent for AI, and I very much doubt you do. Do you know if more has been spent on LLM the last 9 years -- since "Attention is all you need" -- than Internet infrastructure during, say 1995 to 2004? That included the dotcom crash. Did you lament how a failure the Internet was?
LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?
It’s better than many of the AI offerings and the bike could steer and the ducks are sitting on saddles, but the three nephews can’t reach the bottom of the pedal stroke and by the looks of their feet on the far side of the bike they aren’t trying. That shouldn’t satisfy Donald and Daisy, leaving them with all the work.