God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing
Guess they don't care about regular devs atm and are focused only on hardware sales.
Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?
OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.
> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?
Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers that much). They could do what Kimi did - make a good subscription with good models, once you get enough customers to get some good PR and such, pause the signups so you don't have to spend more on running the service than you want/can. Do enough of that and people will talk about your offerings organically, make yourselves known to even devs as "That one company with their own hardware and the super fast subscription." experiencing which would do more than any marketing.
Coding subs are good when they promote usage and adoption of your models in enterprises at API rates.
Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact.
Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.
In their case not to gain mindshare or money or whatever, they're already a market leader, but to run something that validates the use case of their own hardware (across a bunch of 3rd party models) on a practical level and gain whatever insights or details might be relevant to pass on to other hardware and software teams.
Ultimately the frontier labs are competitors of Nvidia. There is a fixed amount that the market will pay for tokens. If the frontier lab model premium collapses due open weights models, Nvidia can capture a greater share of aggregate spend.
I don’t think they will until they change the architecture.
They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.
Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.
>Guess they don't care about regular devs atm and are focused only on hardware sales.
Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn't be about model choice, it'd be about how expensive it.
Seriously, even as well paid as many devs are these are not machines that are affordable for personal use. Their market is people slapping down millions on frontier model training.
I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.
I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.
> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.
Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...
Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.
It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!
>Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs
Oh yeah, I'm still amazed how good the current iteration of models are for coding (I have a fear it's too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax.
Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.
gpt-oss-120b is absolutely unusable over Cerebras. It fails to call tools half the time and just continues to think about what tool it'll call repeatedly. Like it says it'll call a tool and then it doesn't, and then it says it'll call the tool again and then it doesn't, and it just does that in a loop forever. It's awful. Also forgets to end the thinking block too. Even if the model itself was just-okay for its time, even at 1000t/s+ it's not worth it. And it's EXPENSIVE, like $5 per minute expensive
As MoE with 5B active parameters it's pretty fast. But you still need a lot of vRAM, or have to run small quantitations. Qwen models just gave you more bang for your buck, and the gap became worse with every qwen release
Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?
Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3 and GLM 5.3 are near-SOTA in performance and considerable in size, a great choice for proving the platform!
As for the 2nd part of your question - that wasn't a relevant concern or consideration here, unless the models would be tainted to a degree to prevent them from having a good coding subscription that gets more developer mindshare towards what their chips can achieve and generate some good PR.
So they should go with near sota instead of SoTA due to hype
? In the Current market there are vendors who are ahead and Chinese quickly catch up.
You pick the vendor who is ahead, create an agreement with them to get access ahead of public release , and bake that model into hardware , because that’s how you make money .
Guess they don't care about regular devs atm and are focused only on hardware sales.