Hacker Newsnew | past | comments | ask | show | jobs | submit | more syntaxing's commentslogin

GLM series has made it very practical to self host. If the new update for Deepseek flash holds up, I think it would be silly for some companies to not self host.


We're spending 225k a year on tokens. No reason not to buy the hardware necessary to run DS4 at this point.


Are you saying that you already spend so much money for tokens that too late to buy the hardware? Sounds like a trap...


Agreed, and you can write cool infra agents to do stuff for you that runs during off hours like nightly tests and triage.


It’s wild how broken search has become, LLM made it worse but it was already downhill prior to that. Even Reddit forces you login to search within a subreddit and displays the annoying “Best place on internet” overlay after staying on the site for 2 min. I just use AnythingLLM with DuckDuckGo + local LLM model to do searches now.


Most sites required you to log in to do a search. At least every forum that I've ever used was like this.


"Computer. Search result. Relevant."


iPhones are iPads have a milled CNC “unibody” as well… Being able to dictate where to have more thickness helps a ton on the luxury feel. Most of the chips are recycled anyway into new billets.


Unibody is a very appropriate technique for a small portable device that needs to be strong from force in many directions. It is overkill for a keyboard.


Overkill is the point. Same thing for iPads and iPhones. It gives a luxury feeling. That and Chinese manufacturing have made milling a very cost effective manufacturing method.


I’m kinda surprised Apple doesn’t do something like this. I would imagine it’s a lot easier with phones with LiDAR


I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.


You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip.

I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)


They show that CS-4 can't really do batching (or rather it can't properly benefit from it), total throughput barely changes (25%?): https://cdn.sanity.io/images/e4qjo92p/production/6a132331880...

Which I think makes it feasible to approximate activation from CS-4 tokens per second per user.


(Where did you see that?)

This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?

I think the frontier providers keep the size of their models carefully hidden.


You don't really need to train a 10T model to test cerebras against a 10T model. You can feed it an untrained (randomly initialized) model and benchmark it. Result will be gibberish but performance the same.


Mythos/Fable are around 10T:

> According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion

https://www.reuters.com/technology/bytedance-targets-mega-ai...

I believe this report has confused Opus (which is known to be around 5T) and Fable.

Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk talks about the models being trained on Colossus2


I'm confused. I thought Mythos 5 and Fable 5 were exactly the same model just with a different security layer in front of it. Could they mean the Mythos 5 Preview?


Yes. One reason why I think that report has confused Fable and Opus.


> I believe this report has confused Opus (which is known to be around 5T) and Fable.

5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.


> and often described as a match with Opus in overall quality

It's not. Idk about who has more T's but, unfortunately, DS4 pro is not a match to Opus, at least not Opus 4.8.


It's not an Opus match.

The difference is very visible in long tail applications. Exactly where you'd expect parameter count to matter.


It's rumored fable is around that 10T number


If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.


Not necessarily, there could be diminishing returns on mere parameters count .

There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today


That's precisely what he is saying, there is diminishing returns (or optimization left on the table).


I read it as it is impressive because smaller models 2.5T are squeezing similar returns as 10T models despite being 1/4th size not that there beyond 2T today the number or parameters do not have much meaning


or the latest qwen3.8 27B doing so well at ~1/100 the size of K3


What about general knowledge you can get out of it before hallucinations start?


I do not rely on any LLM of any size for general knowledge baked into the weights, they all hallucinate and that is the wrong way to hold them imo

I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK / API surfaces that have since changed from the training data


> I do not rely on any LLM of any size for general knowledge baked into the weights

You have to rely on it to a certain level for agentic/coding work, presuming that's the general subject we're talking about here... For instance I recently encountered a project where it would have been a lot worse if the LLM didn't already know "what is" xterm.js and a bunch of its associated npm-related/node related software. If it was still smart but had to google and find results for everything it would have been a lot more time consuming and risked sending it down a wrong path.


for sure, there is a minimum size and knowledge base that is required to be useful

at the same time, search may find newer or better alternatives, and you can always specify specific technologies you want to use, I typically do this when starting a new project


Storing general knowledge in VRAM has always been a dumb idea in the first place.


It did OK on schlongbench v1.0 (test of a specific niche word that doesn't make it into smaller LLMs) but it sure does love to count words

https://pastes.io/r8F1AY8h


Qwen 3.8 27B beats Opus, Fable and GPT 5.6 by a comfortable margin on the AA-Omniscience Hallucination Rate benchmark.


GLM 5.3 is "only" 753B parameters. Much much smaller.


You want to take a look at the "Scaling Laws" paper, so you can extrapolate from these numbers.


This paper, as well as the Chinchilla one, aged like milk though.


And GLM is only 0.7T!

But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.


Fable is most definitely nowhere near 10T.

The cost to train and infer that would be insane, even by today's standards.


Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T.

Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai...

That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx

Both Grok and Bytedance are training 10T models.


The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac.

Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.


The open models don't really match Opus.

For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention.

I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.


Even if they don't match current-day Opus in everything, they do beat 6 month old Opus, which we have no reason to believe it was smaller than the latest version.


Yes. And Opus goes a very long way compared to Fable, Anthropic isn't doing any favour, it's clearly just 2 models with a very different amount of parameters.


If Fable is seriously around 10T and Kimi K3 sidles up to it at 2.4T

That would be extremely surprising and a massive blunder by Anthropic in model design architecture ... which I highly doubt to be the case.


In real world comparisons K3 is somewhere between Sonnet and Opus. Fable is just a completely different (higher) level.

The long tail of tasks and queries is where you see the difference.


Wasn't Opus ~1.5T and Fable is about twice that?


Kimi K3 is a 2.8T model that's available at about 1/4-1/3 the cost of Fable from multiple providers on openrouter. The math doesn't seem wildly off.


The raw margins on proprietary model inference are rumored to be quite high though (they have to successfully defray the entire investment into model training and datacenter capacity for inference, which is massive enough). The API cost you're paying for the model includes that raw margin.


The Chinese have similarly high profit margins on Inference via their first party api


> The cost to train and infer that would be insane, even by today's standards.

This assumption is likely what has led to the erroneous failure.

Enterprise compute per rack has scaled multiple fold in the last 3-5 years. Alongside the training efficiency gains & datacenter scale increases, even 50T+ is well within reach at the top end.


Yeah, that's insane, you'd need to have many billions of dollars and buy up a huge chunk of the worlds memory supply to do that /s


pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks


Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.


As others have noted, most benchmarks stress the torso, not the tail.


Didn't they say that they can support bigger models now?


Memory capacity on the WSE is the same as before, but access to off-wafer memory is much slower, so the sweet spot is a given fixed balance of memory and compute. They have announced a partnership with AMD in which CPU/GPU hardware is used for part of the workload and the WSE-3 machines are used for inference for specialized smaller models, but I'm not really sure of the details on that.

And there is, of course, the educated guesses about what WSE-4 will be, one being adding a LOT of stacked SRAM or DRAM to tip the balance towards memory (which could also be done by having a few different tile designs with various configurations of compute and memory capacity). I am curious about which way they'll go.


I think the hardest part is how to define “waste”. Are you “wasting” your time if you go to a school event for your kids during the weekday? Is it a waste of time walking your kids to the bus? Is it a waste of time reading to my kids before they sleep? I argue it’s extremely hard to draw that line and a personal question for sure but being “efficient” is hard.


What speed do you get on this setup? Im tempted to use the same GPU.


It varies widely based on a bunch of factors. With this specific model at 8-bit quantization and MTP, it starts out at about 25 t/s for basic chat, but for agentic tasks with long context it slows way down to something like 12-15. I don't see a big difference in token rates based on any config changes I have tried, or going to the smaller 6-bit quantization, so far, though I haven't spent a ton of time on experimenting.

If you already have one or more of them, then, yeah, you can use them for this model or any other at around this size or smaller, but I wouldn't recommend you buy them (or anything else, right now, everything is just too overpriced). You can run better models for less money at higher speeds. I bought mine before they got more expensive, but I wish I'd just bit the bullet and bought newer/faster cards before they got more overpriced. Or, the actual smart money, even back then was to just use cloud models and forget about self-hosting.


Update on this: When I enable tensor parallelism in llama.cpp, I see 25-33 t/s. With reasoning effort set to medium, Qwen 3.8 finished the same task that previously took 11 hours in a little over three hours, which is still more than three times what most of the large models required including Opus 4.8, and nine times what GPT 5.5 (the fastest of the models I've used) needed for a similar task. So, it's still not fast enough for comfort, but it's much faster than the first run. And, I guess, faster than writing the code myself.


Thanks for the update! I wonder if a Blackwell GPU would be noticeably faster. Which vendor did you end up using? I want to get a gigabyte one but thats be OOS for months.


The Blackwell and Strix Halo will be similar to each other and much slower than the number I'm getting on the dual V620 setup (I see about 10-15 t/s on my Strix Halo with this model at 8-bit quantization depending on context). Prefill is generally quite a bit faster on the DGX Spark and token generation slightly faster on the Strix Halo, as I understand it. But, there are better software efficiency improvements for the Spark line.

This model is far from usable on current AMD or Nvidia 128GB AI machines, IMHO, they just don't have the memory bandwidth, especially since it chews so many tokens for any task. If you want to run this specific model, two (or more) 32GB GPUs with decent memory bandwidth is the right way to do it. It doesn't benefit from the larger memory of the Strix Halo. There's enough room for full context and 8-bit quantized model in 64GB. But, it's really a terrible time to buy hardware. MoE models are a much better fir for the Spark and Strix Halo; you can run Laguna S2.1 (slowly) or one of the Qwen 3.6 MoE fine-tunes (pretty quick). Ling 3.0 Flash also looks promising. Nemotron 3.5 Lightning in the MXFP4 quantization absolutely flies on the Strix Halo at 65-80 t/s, but it's dumb. But, all of those are weaker than Qwen 3.8 27B for coding.


The issue with “AI” is how toxic it feels. Feels grimy like social media. Just like social media, I haven’t introduced any of it to my young kids nor do I plan to until they’re teens. At most I give them access to local AI through home assistant. I have a coworker who have kids of similar age and had to get rid of all his echos when Alexa started feeding his son new Lego sets as “Christmas gift ideas” to tell his parents.


when I was a child I read a book about a luddite family who avoided getting their kids neurolinks and the oldest child decides to go and get one put in at you know, coming of age, and then dies from the implant. I have no idea why I remember this but. Godspeed to you.


Why does it feel toxic? It's literally like the computer in Star Trek.


When used correctly.

It's the circular saw on the job site that makes a bunch of manual jobs quicker. But at the same time it can feel inappropriate, like people are in love with their circular saws so much that they bring them to concerts, and try and do jobs that don't call for a circular saw, like painting walls or digging a hole.


Would I be surprised there’s bench maxing happening? Yes. But some users also use Q4 quantized and complain how dumb local models are.


Is there a reason why so many of these agent harness are written in node.js?


Because:

1. The first significant agentic harness was made by Anthropic.

2. One of the most senior developers of client-side software at Anthropic is Felix Rieseberg, one of the original creators of Electron. [1]

3. After Claude Code blew up, everyone else copied Anthropic.

---

1: https://daringfireball.net/2026/07/claudes_criminally_bad_ma...


I think it also helps that it's basically the default platform for any software that AI writes, unless you tell it otherwise. And JavaScript is one of the most widely used and well known languages in the world, so there's that, too.


1. it's built for async 2. runs everywhere 3. interpreted, making it fast to iterate on 4. decent performance 5. most popular language, llms are decent at writing it


JVM has real and virtual threads and arguably just as good if not better on all these points.

(I actually have/am writing a harness in Java fwiw, but mostly as a hobby/experimentation)


JVM apparently has the disadvantage that nobody under the age of 40 wants to touch it anymore. I admit I haven't worked in it in 20 years, but I do think it's a marvel of engineering and unfairly maligned. It used to be my career but I wanted to be closer to the metal.

Having Oracle's tramp-stamp on it may have been the final kiss of death in terms of totally-superficial "coolness" factor.


The JVM has a fixed size heap which for me it is wasteful.

IMHO, Microsoft made the correct approach on .NET.

For LLMs, I prefer C# and C++ instead of TypeScript, JavaScript or Python as the static + compiled language factor keeps the coding agents on track. Plus, they have a true threading/async implementation.


It's only appearing wasteful if you're not understanding how memory management works on modern operating systems. It's not wasting any RAM at all if you pay attention to RSS vs VSS.

The actual physical RAM is still entirely available to other applications. It's just made the OS know it might want that many pages. Until there's data in the pages, they will not count towards total RSS.

It's the kind of things some sysadmins used to gripe to me about and I would question whether they should be in charge of a machine at all.

To repeat: just because an application mmaps a large region doesn't mean the OS has actually given it all that physical RAM. It's merely made sure the pagetable knows about it.


I know how mmap works. The JVM is/was terrible on freeing allocated memory though.

If the program is actively using that allocation, that's fine. My problem is with the runtime hoarding RAM when it should have been freed after GC back to the OS.

Then there's also the JVM not handling peaks well because it hit the max heap size, while you still could rely on the OS doing its job to shuffle stuff to swap temporarily. I still see JVM OOMs in my $dayjob's product while the OS has plenty of free physical memory. It is stupid.

I mean, we have malloc() and free(), they are in the stdlib for a reason :)

The JVM seems to follow a philosophy where it assumes it is the only process running besides PID 1, which is valid for some scenarios, but not for others.


Well, you either spend energy freeing stuff or you waste some memory. It's a tradeoff, that you can trivially set with a single flag in Java.

In pretty much every other runtime's case you are stuck with whatever their GC uses, and almost every other GC is far less advanced than the JVM's implementationS. And manual memory management is not free of tradeoffs either, e.g. RAII can have pretty long destruct chains in both C++ and Rust.


Then apparently no one less than 40 works at a faang or any other company from the top 100?

Because java is still one of the top 3 languages by any ranking worth its salt (not you, tiobe).


Even when I was at Google (2011-2021) it was slowly falling out of favour and by the time I left, Go was fairly rapidly taking its place as the "garbage collected managed language" choice.

Not saying that's a good or a bad thing. I still think the JVM is remarkable.


and you could say the same if not better from C#. But those are now becoming niche languages and ecosystems. One for people around microsoft, azure, etc. The other around oracle solutions. There are still pretty interesting projects around it any of them, but they seem to be losing mindshare against other languages.


nah, C# is huge in game development and things adjacent to it. Lots of young people know it and love it for that reason.


But have you considered that java is gross and nodejs is sexy?


I hate nodejs and the npm ecosystem more than most, but Java, really?


An additional benefit of interpreted, I think, is to make plugins easier to distribute and incorporate. With a compiled language you’d need message passing or something.


dynamic linking was invented pretty long time ago


You know that dlopen does not compare.


That actually don't like you're describing Python. I've been working on a couple JS/TS projects and it's like the models I use (Claude Sonnet and DeepSeek v4 Flash) continually struggle to do coherent work; I have to always keep close watch to reduce sloppiness. I go to Python and it's smooth sailing with minimal prompting (and reduced token burn) for acceptable outcomes.


aren't 3 and 4 a tradeoff though? Yes you have 3 but "decent performance" cannot be an extaled value as compared to "runs everywhere". If its used as a counter balance to 3 then it shouldn't be its own unique point basically saying 4 is true despite 3 in this case.


when you're waiting for network or LLM inference the raw performance doesn't matter at all


This line of thinking I feel like assumes it's the only program running on your computer. Using less of my CPU and memory means my computer can do more things in parallel, or even run more instances of the harness. My laptop is sweating when I got 5+ claude code sessions running.


Is it actually claude using those CPU cycles though, or the agent running test suites and what not?

Honestly I would not be surprised when it actually IS claude using those resources... It is very clearly vibed


I haven't monitored CPU usage so closely, but seems to get heavy with basic tool calls and editing. Memory usage definitely is out of hand, have an idle session right now eating 500MB


You would think raw performance wouldn't be a problem given most of whats happening is waiting for network calls and streaming tokens.

But modern bloat manages perfectly well to make apps that wait for network calls run poorly enough to give you a bad experience.


yeah fair enough, my entire point is not about the application itself but the contradiction on using superlative terms for all points but a compromising/normal term for one. Like if performance is not revelant why include it in the list of benefits.


This is so very, very incorrect.


> it's built for async

You do know that node is event driven, right?


decent performance, lol! compared to what? a shell script? "i'll only take up 200MB of disk and 4GB of RAM to output flickering text on a terminal. boy this is high performance"

fast iteration is for POCs. once you have the app built and working, you need performance and stability much more than fast iteration


Probably for the ease of coding extensions — which strikes me as outdated thinking: if it’s open source and you’re outsourcing the coding to LLMs, why not use a compiled, safe language?

There’s an interesting counter example for DeepSeek called CodeWhale, though:

https://github.com/Hmbown/CodeWhale


Also isn't OpenAI's Codex written in Rust?


I don’t think so. The ChatGPT app was, which is the “Classic” app now. The Codex app that they’re carrying forward is an Electron app and if you forget to quit it before you walk away it’ll make even your M5 Max unresponsive eventually. Sad days.


Sorry, I meant the Codex CLI harness. I don't understand why OpenAI decided to start using the "Codex" name for everything.


Codex, Kiro, Grok Build. Pi has a clone in Rust too.


I don’t think so it’s an npm package iirc


Reasonix (which has been seemingly the most recommended harness for using DeepSeek, as it is designed around maximizing caching in DeepSeek) is now a Go app, but still installed via npm. Which feels ugly, but I guess everyone has npm already, and it handles binaries, so I guess it's a reasonable choice.


you can install binaries with npm too, not just limited to js


I'm not sure why specifically Javascript instead of something like Python or other options, but using an interpreted environment minimizes the friction for implementing extension systems, which are an important feature in AI harnesses.


Python basically requires containers unless you are OK with it bit-rotting every six months or so. At least, this used to be the case for trivial python, and recently was the case for stuff that uses cuda.

I stopped paying attention the third time they redefined matrix arithmetic semantics. That happened to be around the 100th time I was sent a script and it only ran on the author’s machine. Maybe they will fix it some day. When they do, I will not believe it.

In contrast, TS has a much nicer type system and better async support. It runs well on web, mobile, desktop and server. Yes, sometimes you have to ship node.js or a whole web browser, but the tooling for that is slightly less insane than the analogous tooling for python.

Its language interoperability story is slightly nicer too (invoke native code, or use wasm). It’s UI story is much, much better since it reuses all the web stuff.

Pip practically invented the supply chain attack; npm perfected it. That’s probably a draw.

Of course, if you care about performance, then other choices make more sense. If you’re training a model then python probably still wins, but very few customers have a $1M+ machine.


`uv` helps but it's new, and I don't think it has the same mindshare yet on "I just globally want to install this thing that needs an interpreter/runtime", so Python probably just doesn't come first to mind.


I'd put it on this. In my experience Python is fine for scripting your own machine but an obnoxious platform to distribute code on. It's very fragile to version changes, in both directions; I don't know how many things I've seen that only run on 3.10, not 3.9 or 3.11. Its packaging system is global by default which only compounds this because everything needs a specific version but they're all dumped in the same place. And it tends to have a lot of native code as dependencies, leading to all the issues of needing to either have the right build environment or a runtime environment that's already been built for.


codex is written in rust fwiw

smol has implementations in Go, Python, Clojure, PHP

https://github.com/smol-env/smol

out of the box an agent only needs to be able to do http requests and call tools (which might again be just http requests or shelling out)

there is no inherent reason for why an agent has to be in JavaScript or Typescript

but they are popular languages and come with runtimes and libraries for http requests, steaming, TUI (terminal ui) and so on which can help


I'm interested in this but why those four separate languages

Edit: okay I read the code, it's actually four separate implementations


Yes it's separate implementations of the same minimal idea

I'm currently working on more 'feature-full' but still minimal variants

e.g. a python variant with automatic compaction + truncation of sh output

https://x.com/__tosh/status/2087606344035479632

i also got quite a lot of requests to provide the code in non-golfed form to make the implementation more approachable and idiomatic in each language (will do!)


TypeScript is great and its ecosystem is easy to work within.


TypeScript's type system is extremely expressive while still allowing you to retain the flexibility of a scripting language. v8 and JSC also have decades of performance tuning across basically every consumer device.


I'm still waiting to find one written in Rust that I really love. I don't think js/ts makes sense for terminal based applications


Nothing better for UI than React. Electron/tauri or node kinda falls from it.


I gotta same problem.


Because that hammer is their only tool!


Might be easier to do cross platform.


npm as a distribution tool works well and typescript has types.

Any reason why it should not be written in nodejs?


I always thought nodejs was a weird choice for CLI tools.

For web stuff, sure.

But for CLI, it never made sense to me. Especially when Python and Go exist.


> But for CLI, it never made sense to me. Especially when Python and Go exist.

But why? Not saying node is better, just want to know where you are coming from for my own knowledge.

Bc I would have picked typescript + node too. It has types (where python just has type hints) and a lot of developers know it already (where go is more niche).


I guess the reasons are that I feel (no data to back this up) that having a Python runtime available is much more probable than having a nodejs runtime around.

I was going to say you cannot easily distribute a nodejs based CLI app, but that’s of course not true. devcontainer-cli is a nodejs app and so are many of the coding agent harnesses.

Yeah, thanks for pushing back. I guess my view was irrational.


Skill issue. Their models don't work for serious programming so everyone just copies Electron apps from each other.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: