People have been making claims about the commoditization of llms since chatGPT, and they've been wrong every time as quality and prices and differentiation have increased.
But Scott's point is more: why even have markets? Once you have the superforecasting available on the questions you care about, why do you need to publish it for everyone to also react to?
I mean he explicitly at the end says there will be a new era of prediction markets:
"This, then, is my prediction for the AI superforecaster future: for basic questions, your off-the-shelf AI chatbot will be able to offer opinionated probabilities superior to those of any human. For more controversial or bias-laden questions, a new era of prediction markets will smooth over differences in brand and model and efficiently aggregate all AIs’ opinions."
But one reason that you would publish it for everyone else to react was exactly what I said, because if you can show people that a superhuman AI believes a certain outcome is more likely, you impact whether or not that certain outcome is more likely. Which feels like a flaw in this version of the future. That in fact you could overcorrect and make the markets less efficient.
Almost by definition, once AI forecasters are in the market, they won't (all) be beating the market.
But why evaluate AI forecasters by beating the market? Do we evaluate deep learning by whether hedge funds make money from it in the markets? These things have far, far more utility outside of finance.
Doesn't this argument prove too much? Why does AlphaSense sell their company research instead of using it to trade themselves? Why do people work on open source time series forecasting packages instead of quietly using them to trade?
I'm going to pre-register my prediction that GPT-5.6 Sol is significantly behind Claude Fable 5, as evaluated by general consensus once time has passed for people to get familiar with both.
Claude will win on "vibes" and it'll be close in coding but considering how incremental Fable is above 5.5 in terms of overall smarts, there's no way 5.6 isn't considerably smarter on the whole.
Fable is allegedly a massive model (estimates between 6-10+ trillion, with a few hundred billion active). If 5.6 is just an incremental upgrade over 5.5 (at the same model size) then it won't be able to fully compete with Fable just yet.
I’m countering this prediction by stating that Fable and Sol will be somewhat similar - this has always been the trend and I see no reason why this should stop now.
OpenAI may have a model in the works that is similar next-gen size and architecture to Fable, but this isn't necessarily it. I'd guess that 5.6 was more of a hasty reaction to Mythos - same base model (same size, same price) as 5.5 but with additional post-training to make it more competitive with Mythos/Fable in some benchmarks.
Mythos/Fable is supposedly next generation in size vs Opus, and is rumored to have some architectural innovation in terms of dynamic routing/compute, possibly only fully enabled with Fable which at $10/50 is still twice the price of Sol 5.6's $5/30, but a big reduction from Mythos preview which had been an astronomical $30/150 possibly due to the dynamic routing not yet having been enabled.
Is this the trend? There have been various points where one of Anthropic or OpenAI was substantially ahead. Sure, many times they're close, but now doesn't seem like one of them.
"Affordable" depends on what you need. When a task is able to be achieved by two different calibers of model, it's obviously more cost effective to use the less capable model, in the same way that you wouldn't hire a math PhD to do simple addition.
If what you need is only possible with the more capable model then the "affordability" of the less capable model is sort of irrelevant. If what you need is a novel mathematical proof, it doesn't matter that a high school student is "more affodable". You need the math PhD.
As "old" models get more and more capable, it's going to be an increasingly important skill to be able to adequately recognize when a task requires a frontier model and when it doesn't, so that the less capable (and therefore cheaper) model can be used.
If it's helpful, I'm still holding at July 9 as my median date that Fable gets re-released to Americans, the news of the last 24 hours didn't update the model meaningfully.
People forget that Meta already did this years ago, before prediction markets became the next big consumer trend for them to chase.
The app was called Forecast, and launched in June 2020. (Around the same time that Kalshi and Polymarket launched, actually!) It was framed as a way to make the comments and activity on Facebook actually productive rather than toxic, and build expert reputation signaling mechanisms.