Hacker Newsnew | past | comments | ask | show | jobs | submit | syntaxing's commentslogin

> 23 agents total.

This hit a bit too close to home. Sol has the same issue, spawns a lot of agents for no good reasons (besides burning tokens).


I heard this was a thing when listening to a Theo podcast, he mentioned to add a "Only use subagents if the user explicitly requests them" line in your agents.md file.

I don't know if it works, but I've always had a consistent level of token burn on my plans (I've only heavily used Sol after adding it).


I’m more curious how each 4 bit quant compares. It seems like NVFP4 outperforms Q4_K_M in terms of speed and top 1 but is only good for expensive Nvidia cards

I swear, Qwen 3.8 27B @ Q8 is smarter than Sonnet 5 most of the time. Why wouldn’t corporate America self host at this point, especially with better options like Deepseek Flash and GLM 5.3 flash that’s a middle ground between Sonnet and Opus

Agreed. And conversely, American models can also just as easily be secretly influenced for bad things, or be more tightly controlled by the government, to corporate America's own detriment.

Post-IPO I'd trust the American models far far less than the Chinese models.

The most insidious advertising in the world is about to be surfaced as people use LLMs to look for product recommendations.


It's wild that people are so lost in the sauce of social media that they have no qualms handing over all their data to an authoritarian ethno state with a single ruler who appointed themselves for life with unrestricted unilateral control over every aspect of the state. And here we are calling for collapse because mom might get recommended Gain instead of Tide.

Sorry, how does running a local model hand over my data?

You are very confused about the security and what can happen here.


The context here of comparing trust between American models (which are mostly centrally hosted) to Chinese models (rather than "local" models) implies trust levels in who is hosting.

American local models don't hand over data either, which would nullify the point of the comment.


For most people around the world America is more of a threat than China. Heck for most Americans the American government is more of a threat than the Chinese government.

trump has named himself dictator for life now? when did that happen?

What's even more wild to me is that basically ALL modern electronics and all modern batteries are made from raw materials sourced by forced and child labor (and sometimes both)... yet the vast majority of the entire world just turns a blind eye to it.

https://gcdnb.pbrd.co/images/gKMYRcEIeE9j.png


This is very confused, first the world does not turn a blind eye, and by having US purchasers of these products the US is actually forcing a big change in the standards for the better due to purchasing power. In particular the US has avoided huge amounts of potential purchase in solar, largely under the justification of avoiding forced Chinese labor.

Second, the "ALL" qualification is extremely wrong, as only small fractions of these products in the US could ever be sourced to the human rights violations cited here.


America's own demise will be made in America, stamped by American laws

From whence shall we expect the approach of danger? Shall some trans-Atlantic military giant step the earth and crush us at a blow? Never. All the armies of Europe and Asia...could not by force take a drink from the Ohio River or make a track on the Blue Ridge in the trial of a thousand years. No, if destruction be our lot we must ourselves be its author and finisher. As a nation of free men we will live forever or die by suicide. ― Abraham Lincoln


@q4 is definitely smarter than sonnet from what I’ve seen so far. It’s even caught problems in code made by fable, when using it as a code reviewer.

> Why wouldn’t corporate America self host at this point

Because they've been trained to think "cloud-first" for a decade?


Still a ton of non tech companies doing a digital / tech transformation out there too lol. Maybe some so far behind they still have the real estate and rack space to get ahead on this one

is this actually the case? I haven't kept up with the small models

but if there's roughly Sonnet 4.6 level capable open small models, then I'd be impressed


Qwen 3.8 27B is the real deal BUT remember to use froggeric template and/or medium reasoning.

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates


Qwen 3.8 27B is better than Sonnet 4.6

Qwen 3.8 27b has the juice. Try it.

Why not Luna?

I mean, why self-host? Deepseek v4 flash is dirt-cheap on openrouter. Inference is a race to the bottom at this point.

I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.

One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.

I believe that multimodal training increases the robustness of latent representations regardless of which modality is being processed.

Tangentially this makes me wonder how large Opus really is. Perhaps Opus is a lot smaller than most of the 1T+ assumptions, just a lot more post-training/ finetuning on a 300-400B sized MoE model.

How do people bypass captcha or robot checks? All I wanted is a price aggregator but it always gets blocked by major retailers.

I'm with you here. Every demo I asked for where I knew a captcha / cloudflare would block an AI directed scraper was unsuccessful / produced unsatisfactory results. WebMCP for the win, or loss, depending on your perspective. Personally, I could get behind WebMCP if micropayments ever became a thing. Of course, then my incentive to trick your Agent in crawling millions of pages and paying me lots of cash would be rather high.


If you get a children’s card, they print your kids name right on it. It’s a nice little souvenir to keep.

Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.

This is exactly what Jenson said in all of his interviews. Banning it in the short term would have long term consequences.

I’m more curious on the size. If it’s smaller than or equal size to GLM 5.3, this would be a crazy good model. If it’s closer to deepseek pro, it would be a good model. If it’s near Kimi K3, I think it’s competitive but nothing particularly differentiating.


Definitely agree. If it is small (eg. Qwen 3.8 28b or gpt-oss-120) then this might be amazing. If it is anywhere near Kimi K3 it would need to have some other differentiating factor than intelligence.


With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further


It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them


Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.


Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.


Have you given Ornith-1.5-35B a shot?

It's been a pretty decent step up for me compared to Qwen3.6

https://news.ycombinator.com/item?id=49362401


I know it benchmarks very well. I haven't tried it yet though.


it'll hopefully improve with more MoE and half the prefill/generation. I think it's the sweet spot for the strix halo for smarter or vibe tasks.


What sort of pp/tg speed do you get on a Strix Halo?


This is the best I got, all with Unsloth's quantizations.

Laguna-S-2.1:UD-Q4_K_XL (no MTP) pp=186.4 t/s tg=27.8 t/s

Qwen3.6-35B:UD-Q4_K_XL (with MTP) pp=404.4 t/s tg=83.2 t/s

Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP) pp=343 t/s tg=12.1 t/s

Laguna actually performed better than I remembered. I thought it was slower.


Have you benchmarked against full precision models for accuracy/ performance?


Not full precision. I've only benchmarked 27B across Q3-6 quants using lm-eval. I lack the hardware to bench 27B at BF16 but I might be able to do Q8_0. I haven't gotten around to doing 35B. I really should upload my collection of results to Github or somewhere.

Here's a summary of what I have for 27B. I used unsloth's UD-Q{3-6}_K_XL quants across 11 evals. The values are pretty linear between Q3 and Q6.

    Qwen3.6-27B     Q3    Q6
    ARC-Challenge   97.0  97.0
    BIG-Bench Hard  57.9  59.3
    GPQA Diamond    77.8  83.3
    GSM8K           92.4  92.6
    Hendrycks Math  35.5  38.9
    HumanEval       80.5  85.4
    HumanEval+      75.0  79.3
    IFEval          87.3  88.0
    MBPP            75.2  77.2
    MBPP+           88.4  88.9
    MMLU-Pro        83.1  83.5


What are pp/tg? I get 30t/s on 27B qwen.


pp is prompt processing how fast it processes the prompt. Tg is token generation how fast, it generates tokens.


okay, so prefill and decode would be the terms I was already familiar with.

So, I know https://cactuscompute.com/needle is designed only to enable tool calling on tiny devices. But, I wonder if anyone has used it as a CPU-side mediator between a tool and a GPU-side local LLM making semi-natural-language tool requests...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: