I heard this was a thing when listening to a Theo podcast, he mentioned to add a "Only use subagents if the user explicitly requests them" line in your agents.md file.
I don't know if it works, but I've always had a consistent level of token burn on my plans (I've only heavily used Sol after adding it).
I’m more curious how each 4 bit quant compares. It seems like NVFP4 outperforms Q4_K_M in terms of speed and top 1 but is only good for expensive Nvidia cards
I swear, Qwen 3.8 27B @ Q8 is smarter than Sonnet 5 most of the time. Why wouldn’t corporate America self host at this point, especially with better options like Deepseek Flash and GLM 5.3 flash that’s a middle ground between Sonnet and Opus
Agreed. And conversely, American models can also just as easily be secretly influenced for bad things, or be more tightly controlled by the government, to corporate America's own detriment.
It's wild that people are so lost in the sauce of social media that they have no qualms handing over all their data to an authoritarian ethno state with a single ruler who appointed themselves for life with unrestricted unilateral control over every aspect of the state. And here we are calling for collapse because mom might get recommended Gain instead of Tide.
The context here of comparing trust between American models (which are mostly centrally hosted) to Chinese models (rather than "local" models) implies trust levels in who is hosting.
American local models don't hand over data either, which would nullify the point of the comment.
For most people around the world America is more of a threat than China. Heck for most Americans the American government is more of a threat than the Chinese government.
What's even more wild to me is that basically ALL modern electronics and all modern batteries are made from raw materials sourced by forced and child labor (and sometimes both)... yet the vast majority of the entire world just turns a blind eye to it.
This is very confused, first the world does not turn a blind eye, and by having US purchasers of these products the US is actually forcing a big change in the standards for the better due to purchasing power. In particular the US has avoided huge amounts of potential purchase in solar, largely under the justification of avoiding forced Chinese labor.
Second, the "ALL" qualification is extremely wrong, as only small fractions of these products in the US could ever be sourced to the human rights violations cited here.
From whence shall we expect the approach of danger? Shall some trans-Atlantic military giant step the earth and crush us at a blow? Never. All the armies of Europe and Asia...could not by force take a drink from the Ohio River or make a track on the Blue Ridge in the trial of a thousand years. No, if destruction be our lot we must ourselves be its author and finisher. As a nation of free men we will live forever or die by suicide.
― Abraham Lincoln
Still a ton of non tech companies doing a digital / tech transformation out there too lol. Maybe some so far behind they still have the real estate and rack space to get ahead on this one
I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
Tangentially this makes me wonder how large Opus really is. Perhaps Opus is a lot smaller than most of the 1T+ assumptions, just a lot more post-training/ finetuning on a 300-400B sized MoE model.
I'm with you here. Every demo I asked for where I knew a captcha / cloudflare would block an AI directed scraper was unsuccessful / produced unsatisfactory results. WebMCP for the win, or loss, depending on your perspective. Personally, I could get behind WebMCP if micropayments ever became a thing. Of course, then my incentive to trick your Agent in crawling millions of pages and paying me lots of cash would be rather high.
Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.
I’m more curious on the size. If it’s smaller than or equal size to GLM 5.3, this would be a crazy good model. If it’s closer to deepseek pro, it would be a good model. If it’s near Kimi K3, I think it’s competitive but nothing particularly differentiating.
Definitely agree. If it is small (eg. Qwen 3.8 28b or gpt-oss-120) then this might be amazing. If it is anywhere near Kimi K3 it would need to have some other differentiating factor than intelligence.
Not full precision. I've only benchmarked 27B across Q3-6 quants using lm-eval. I lack the hardware to bench 27B at BF16 but I might be able to do Q8_0. I haven't gotten around to doing 35B. I really should upload my collection of results to Github or somewhere.
Here's a summary of what I have for 27B. I used unsloth's UD-Q{3-6}_K_XL quants across 11 evals. The values are pretty linear between Q3 and Q6.
So, I know https://cactuscompute.com/needle is designed only to enable tool calling on tiny devices. But, I wonder if anyone has used it as a CPU-side mediator between a tool and a GPU-side local LLM making semi-natural-language tool requests...
This hit a bit too close to home. Sol has the same issue, spawns a lot of agents for no good reasons (besides burning tokens).
reply