ok first thoughts - this is a cool, but ludicrously expensive toothbrush! Anyone who buys one is just flaunt their wealth when $2 gets you a manual version
The big issue I have with Fable is this. From the Anthropic email announcing Fable 5.1. So basically they're giving us a Ferrari, which will point blank refuse to do certain stuff - forcing us to go out in our Mustang. Their choice, not ours
"Safeguards and automatic fallbacks (beta): Fable 5.1’s biology and cybersecurity classifiers block fewer benign requests and now permit vulnerability finding in source code. Blocked requests return an error and are not charged to you. On the Messages API, opt in to fall back to another model so users get a response instead of an error. We recommend Opus 5 for biology and Opus 4.8 for cybersecurity. In Managed Agents, fallback is built in."
The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.
Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.
My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here
https://github.com/ed-is-ai/featherbench
Indeed; I find it really interesting there is so much disbelief and outright hostility to my post, in these comments. I am an AI Realists rather than AI hype-artist, say it as it is. For me, the evidence is clear that for our own use cases OpenAI, Anthropic models should not be the default choice any longer.
Incidentally, if you want to look at this in more depth - my previous post about the eval framework got like zero response; that's where the results, methodology, code for running the evals, evals are all shared.
I had to google astroturfing. I liked the due diligence... I am a hacker news infant, so all my history is about this.
Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very easy)
https://github.com/ed-is-ai/featherbench
You might have read about Ox Alpha aka glm5.3-flash, well I updated to include that
Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you can add your own tests if this isn't enough for you...
Thanks for point out - I updated the benchmark today to include. It is every bit as good as everyone says it is. The intelligence / $ is something else...
reply