I started measuring my Claude Max x5 use and last week (they gave me 50% more) I used 1.3B input tokens. Some 130M were cache writes, rest was cached. And 5M output.
This puts things in perspective. We're taking thousands of bucks weekly even if I managed to switch to Kimi K3.
What is the majority of this use? Infrastructure upgrades, troubleshooting and so on. Ingesting quite a bit of documentation at beginning of each session.
Sessions run from few hours to a month long and 1M context usually hovers near 30-60%.
What do people use these tiny limits for?
I started measuring my Claude Max x5 use and last week (they gave me 50% more) I used 1.3B input tokens. Some 130M were cache writes, rest was cached. And 5M output.
This puts things in perspective. We're taking thousands of bucks weekly even if I managed to switch to Kimi K3.
What is the majority of this use? Infrastructure upgrades, troubleshooting and so on. Ingesting quite a bit of documentation at beginning of each session.
Sessions run from few hours to a month long and 1M context usually hovers near 30-60%.