Just because Anthropic and OpenAI really want there to be an arms race justifying the outsized investment, doesn't mean the optimal play is to build larger, more expensive, models.
The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.
I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.
And it doesn't have to be either/or. They could make larger, more expensive models, just at a slower cadence.
Sure downside would be not learning from people using your model for coding, if we're on the cusp of huge leaps in self-improvement. But there is a reasonable case for avoiding desperate scramble, especially if other parts of the business can also create value with the compute.
If the Chinese labs can compete on a shoestring budget with access to much less powerful hardware, Google should be able to compete as well. They're becoming almost irrelevant for agentic coding right now.
I would agree with you on AI/LLM being more than agentic coding but at the same time, I think there's more nuance.
For example, PDF's and powerpoints can be generated using agentic coding by things like https://bento.page or other ways of generating them in an agentic coding fashion.
A lot of browser automation could/is also done by agentic coding.
It can also help them set up and configure self hosted software with the help of LLM's and debugging if its working or not.
You can create videos using Manim and remotion.dev and also excalidraw-animate and generate excalidraw files agentically if what you need is more vector style graphics (which surprisingly can fit into many ideas) rather than say a real life human waving video/more photo-realistic video (but I must say that this has certainly its own pros/use-cases as well).
It might sound self-explainatory but turns out that coding can represent a wide range of problems!
I get that. I use agents a lot and LLMs often reason with code. It is valuable. I just think the floor is a lot lower for general reasoning and common tasks like that. And in 6-12 months it won’t matter. Google will publish better models. The temporal distortion of how long a Sol or a Fable has existed is real. No one is suddenly missing out on some giant competitive edge because their model is a few months behind. I feel like it’s all just going to normalize and things other than how well your model can write code will matter more and more in 12 to 24 months.
Sure I understand what you mean as well and I am not asking for SoTA models to be created by Google but more so explaining why coding is still the largest focus for many labs.
I personally wish to get more smaller models (like the recent qwen model) and other open source models like GLM 5.3 and the glm flash model.
> No one is suddenly missing out on some giant competitive edge because their model is a few months behind
Sure I can agree with that. The competitive edge might still exist but I do get the underlying sense of what you are trying to suggest.
> things other than how well your model can write code will matter more and more in 12 to 24 months.
What are the things then which you feel like could be more differentiative factor? For example, I personally think multi modal is still quite preferrable in AI models. I use GLM 5.2 and it doesn't have vision and I can certainly imagine time/use-cases where multi-modality would've helped coding and even other use cases as well. So what are some other use cases that you are thinking? Video generation models like Veo/Sora?
It's not much of a shoestring budget to be receiving regular injections of investment from state lenders along with cheap credit.
I don't think the comparison holds.
Yes, I agree with you that the race all the AI companies are running doesn't make sense, but at the same time, there are rumors that Google has produced newer versions of Pro without releasing them to the public.
Version 3.1 has plenty of room for improvement, yet they don't seem to be giving the attention it deserves or at least communicating accordingly.
There is more to the cost of a model than its training.
While training is a significant Capex expenditure, it has very low Operational cost after training unless it is deployed for public inference.
It may be that they wish to slow their cadence of releases, or develop their models to focus more in a different direction, etc. No matter what the actual reasoning, they have chosen to not compete in the same race, and I cannot say I fault them.
I work there. I have zero internal knowledge about the model. Opinion my own, etc. I don't think it is worth fighting to win on a month to month time horizon. When you step back and look an inch above this market, Gemini Pro 3.1 as a product was released in February. 6 months. It feels like forever and that Google is behind, but on a 2-3 year horizon? The models are going to stay similar.
Also, look at Flash 3.5 to 3.7. Flash 3.7 is a genuinely decent Sonnet 5 class model. Flash 3.7 is quite efficient too. Also, whatever was spent training 3.5 pro is probably not wasted. However, as a strategy, when I see models like Kimi K3, Fable, Sol. If you discard "because the model sucked" what other alternatives or potential options might exist?
I thought of a quite a few and they are far more compelling and interesting to me.
(Also Gemini models tend to be pretty decent at more than just programming. Enterprise AI use is more than just software eng / programming)
I'm the CTO of a GCP shop with an 8 figure annual commit.
If you'd told me at the end of Cloud Next 2025 that by now Google still wouldn't have a competitive offering to agentic coding offerings from Anthropic (Claude Code + Fable) or OpenAI (Codex + Sol), I wouldn't have believed you.
In our non-coding use cases where we're embedding models in our product, we're also not reaching for GCP stuff. Because Anthropic has the mindshare of our engineers and product folks, since it's what they use every day.
Given your position and the responsibility that comes with it; I sure hope you updated your mental model in another way than simply "they are acting irrational"... It's not clear from your comment that you did, but it sounded a bit like it.
3.5 was almost certainly a 3.1 post-train, so likely a small investment on Google's part.
They mentioned that they have already started pretraining Gemini 4, which will be the full ground up rip-your-face-off-expensive training that is often discussed.
Google doesn't have a good coding model. This is a HUGE problem. They don't need "larger more expensive models", they need a good coding model because it's a competitive advantage.
Second tier models are like self-driving cars to the point where you question if they save any time. Sol can do a deep analysis and plan a large feature. Luna can mostly execute. There's a clear qualitative difference.
The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.
I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.