Sol is for routine work, Opus for frontend/design, and Fable for more complex / ambiguous / architecture work. Fable works extremely well to drive Sol as a subagent.
Fable is the only one you can actually trust to not look at the code, but Sol is somehow still more pleasant to work with, especially in fast mode. Opus is the enemy, and it will make you insane if you talk to it for too long.
The important bit I found is to explicitly remind Claude that Sol 5.6 is a very smart and good model; otherwise, Claude performs its normal condescension towards any non-Claude model behavior and insists on reading all the diffs in full and testing all of Sol's work, negating any token savings.
> The implications are clear: focus on the human-skills part of the job.
The implications are not clear. They are not clear for the security people, for the SWE people, for anybody in knowledge work whose jobs are impacted.
I wish people would stop with the “the solution is merely simply retool against the part the AI isn’t good at yet” cope and feel the enormity of the moment with humility.
When the dust settles these jobs may not exist, or the jobs that do exist will be unrecognizable from the ones today and perhaps so qualitatively different as to no longer be attractive.
Great writeup. The excessive function thing has always driven me crazy; I guard against this explicitly in Claude.md.
I have found that models are generally poor at managing refactors / complexity while also implementing new features. But I’ve had some success with a semi-lights-off approach where you decompose it and prompt the model adversarially in a second pass to look for new rough edges and areas of complexity or refactors that might simplify the codebase.
So I’d be very curious to see this benchmark but with something like a periodic “refactor turn” interleaved in.
Also eager to see Fable benchmarked; anecdotally that was the only model whose code I felt I could actually trust to not review closely.
> This is where OpenAI has an advantage over Anthropic. While its models are trailing Anthropic's in recent months, its investments in product, consumer experience, site publishing, voice, and hardware are all directions that have clearer moats.
I was with you up until this point. I don’t think OpenAI has any more of a substantial product moat than Anthropic; if anything, the Claude / mythos etc brand is a valuable asset that OpenAI lacks.
Yes, many of the elite HN engineer always online types have come to prefer Codex, and but if you actually talk to regular engineers in industry, agentic coding is simply still synonymous with Claude Code.
And for the non-engineering uses, Claude is so much more pleasant of a conversational companion than any of the GPT line, and I suspect is this baked deeply into the model, otherwise OpenAI would have closed this gap by now.
> (if you have to say it, that’s how you know it’s good)
Pet peeve, but no, it's the exact opposite. Good satire is immediately obvious; nobody had to ask whether Jonathan Swift was actually serious about solving poverty in Ireland by having the poor sell their children for meat to the rich. Subtle satire is bad satire by definition; if you have to be told that it's satire, that means it has completely failed to do its job, and is no better than intellectual masturbation.