Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing

Guess they don't care about regular devs atm and are focused only on hardware sales.



Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?

OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.


> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?

Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers that much). They could do what Kimi did - make a good subscription with good models, once you get enough customers to get some good PR and such, pause the signups so you don't have to spend more on running the service than you want/can. Do enough of that and people will talk about your offerings organically, make yourselves known to even devs as "That one company with their own hardware and the super fast subscription." experiencing which would do more than any marketing.


Coding subs are good when they promote usage and adoption of your models in enterprises at API rates.

Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact.

Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.


> Should NVIDIA do a coding subscription too?

Yes, obviously! Well maybe not a subscription but definitely an inference service.

https://build.nvidia.com/

https://resources.nvidia.com/en-us-inference-infrastructure/...

https://www.nvidia.com/en-us/data-center/dgx-cloud-lepton/

In their case not to gain mindshare or money or whatever, they're already a market leader, but to run something that validates the use case of their own hardware (across a bunch of 3rd party models) on a practical level and gain whatever insights or details might be relevant to pass on to other hardware and software teams.


Ultimately the frontier labs are competitors of Nvidia. There is a fixed amount that the market will pay for tokens. If the frontier lab model premium collapses due open weights models, Nvidia can capture a greater share of aggregate spend.


> Should NVIDIA do a coding subscription too?

They sort of do? They offer free access to various versions of nemotron via multiple routing services.


They do have Cerebras Code

https://www.cerebras.ai/code

But it's not fully open to just anyone, I wasted time signing up to find out that I couldn't even sign up for it to test it out.


I don’t think they will until they change the architecture.

They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.

Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.


Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry.

They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?

https://inference-docs.cerebras.ai/capabilities/prompt-cachi...


ah, that makes it feasible! Okay, glad it's not technical limit. They should fix the pricing...


>Guess they don't care about regular devs atm and are focused only on hardware sales.

Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn't be about model choice, it'd be about how expensive it.


Seriously, even as well paid as many devs are these are not machines that are affordable for personal use. Their market is people slapping down millions on frontier model training.


> GPT-OSS 120B which is nigh useless nowadays:

I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.

I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.


> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...

Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.

It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!


>Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs

Oh yeah, I'm still amazed how good the current iteration of models are for coding (I have a fear it's too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax.

Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.


gpt-oss-120b is absolutely unusable over Cerebras. It fails to call tools half the time and just continues to think about what tool it'll call repeatedly. Like it says it'll call a tool and then it doesn't, and then it says it'll call the tool again and then it doesn't, and it just does that in a loop forever. It's awful. Also forgets to end the thinking block too. Even if the model itself was just-okay for its time, even at 1000t/s+ it's not worth it. And it's EXPENSIVE, like $5 per minute expensive


As MoE with 5B active parameters it's pretty fast. But you still need a lot of vRAM, or have to run small quantitations. Qwen models just gave you more bang for your buck, and the gap became worse with every qwen release


GLM 4.7 is gone (at least for us), with no suitable replacement from Cerebras. I think all they care about now is hardware and OpenAI hosting.


Cerebras is the fastest provider by far on OpenRouter, and gpt-oss-120b is still very useful. They have backed up their claims very well.

> Guess they don't care about regular devs atm and are focused only on hardware sales

They aren't trying to make a few bucks off tokenmaxxers. They're trying to be the underpinning of compute for all AI. They're going to beat Nvidia.


Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?


> Why they should go with Chinese models

Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3 and GLM 5.3 are near-SOTA in performance and considerable in size, a great choice for proving the platform!

As for the 2nd part of your question - that wasn't a relevant concern or consideration here, unless the models would be tainted to a degree to prevent them from having a good coding subscription that gets more developer mindshare towards what their chips can achieve and generate some good PR.


So they should go with near sota instead of SoTA due to hype ? In the Current market there are vendors who are ahead and Chinese quickly catch up.

You pick the vendor who is ahead, create an agreement with them to get access ahead of public release , and bake that model into hardware , because that’s how you make money .


Exactly, they need to push to regular consumers as well as businesses. This will create more pressure for adoption


Wouldn't you want to keep this for internal development? Keep the single thread gains for yourself.


Didn't OpenAI bought it?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: