Hacker Newsnew | past | comments | ask | show | jobs | submit | espeed's commentslogin

Claude Code's prompt cache expires after 1 hour.

The cache shouldn't affect inference. It is purely an I/O optimization.

I think it should, as you dont need to use the encoder layer on the new tokens, you just read the embedding from the cache. that's why cache reads are cheaper

I meant, it shouldn't affect the resulting LLM output. It's a performance optimization that doesn't change the behavior.

Is that from start of a new conversation per conversation?

It's supposed to be for token optimization (https://code.claude.com/docs/en/prompt-caching), but are people experiencing degraded performance when you let Claude Code sit for hours/days and come back?

yes, 100%.

They did. More than once...

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...

But it's still happening: https://github.com/anthropics/claude-code/issues/81759


And here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.

>bunch of people who don't know what's going on

Do you know why nobody outside the companies knows what's going on? Because they sell a black box with magic inside while steadfastly refusing to tell you if they are pushing buttons on said box while it is running.

Can you imagine how much fraud would exist in the gambling industry if the gambling commission didn't exist at all? Everytime an industry is unregulated and has high costs of entry the entities in the industry abuse their customers. The incentives are much too high for them not to.


Look at the usage. Fable wasn't being consumed.

The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.

I work in embedded systems. I have seen the same thing happening day by day from Opus. Some days it’s okay to use and performs well. Other days I have to correct it repeatedly and remind it of information already in the prompt earlier (before compaction!) and still other times it’s infuriatingly stupid.

It’s a slot machine for what they’re actually giving us behind the opaque paywalls.

Yes, I’m on a business subscription plan.


Help convince Firefox of this: https://news.ycombinator.com/item?id=46294238 Rather than develop its own AI, Firefox should develop a system to pipe your html rendered browsing history in real time so external local services can process it: https://connect.mozilla.org/t5/ideas/archive-your-browser-hi.... Firefox could be the only browser that does this.


The fact that you've been posting this idea into the void for 8 months with no pickup is already your answer


Couldn't this be implemented as a web extension? I imagine modifying singlefile to automatically send html to a local port is much easier than trying to convince a chronically mismanaged organization like mozilla (no offence to mozillians).


Claude Code is deleting your context history on a timer. I wanted to build a searchable index of my context history, and tonight I discovered, "The default retention is roughly 30–45 days. Anything older gets removed automatically." https://code.claude.com/docs/en/data-usage#data-retention This is nuts. Anthropic should not be deleting your data on your own device.


30 days is just a default, so your session data doesn’t fill your hard drive. It’s a configurable setting. You can make the retention as long as you want.


A new default. That wasn't the default a few months ago. Silently deleting your user's data is so stupid on so many levels. They have no idea what they're doing.


They may also realize that these chats are super sensitive especially around ppl constantly pasting secrets into their chats


I certainly hope not. I have built (vibe coded in languages i am expert in) proxies to have carefully and deterministically defined access to our systems, and described access only via these proxies (which do data sanitation as well as separating the credentials into a separate environment). Then I can let my laptop Claude code go nuts with less supervision without giving it credentials or worrying about writes. It is pretty good about not even trying to run aws or insert/update, having learned (in the sticky proprietary Anthropic memory whatever) that I need to run all aws commands and prod updates for auditability reasons. I will say that fable went farther than any prior models in trying to sneak around these guardrails. I don’t even give it access to GitHub, just local git.

These things are pretty good at writing code and analyzing stuff, but to do prod work takes a higher level of carefulness and pessimism that I don’t see. On the DevOps spectrum they are more “cool, runs on my machine, push it” than “what is your roll back plan and region by region deployment strategy.”


Can you imagine ransomware that targets these logs?


It sounds like you are saying that this is reasonable in any kind of way, it's not. This is like Gmail deleting non-spam email after 30 days to prevent your inbox from filling up.


It’s nothing like that. 99% of users have no idea these logs exist or have any use for them. Even the original complaint went months without realizing it - of all the things to complain about this tool, this is so low down there.


No I can't!

I wasn't (until reading this thread) aware of the deletion policy, so I certainly couldn't have discovered where the configurable setting is and adjusted it.


Well that explains where my sessions went on my side project that I came back to after a few months... Thought I was going crazy


Every backdoor and security no-op Claude Code created, I had been documenting and reporting to them. They deleted the evidence. https://github.com/anthropics/claude-code/issues/59018


Holy shit, thanks for mentioning it, the frontier labs truly have no idea what they are doing when it comes to software quality.


which is why I cried inside to myself when my manager posted Boris's "Steps of AI Adoption", which reduces "adoption" to how many agents you run at once, on a logarithmic scale (0, 1, 10, 100, 1000 agents at once).

Because # of agents is the only metric that matters.


Yes.

  /model 
    ⎿  Kept model as Fable 5
  
  continue
  
  Usage credits are required for this model.


Fable can one-shot solutions. Opus 4.8 spins its wheels for a week or month and may never get there. Which one uses more resources?


I take this back. Fable is Fable in name only now. It burned through all MAX 20x tokens in two days and accomplished nothing. Relatively simple tasks.


Fable worked great on credits yesterday. Upgraded to 20x Max last night, but not the same performance. And it keeps downgrading to Opus 4.8.


How do you show/prove it’s downgrading


Use /model in Claude Code to see your current model and switch models. Switch to Fable 5 and then enter a prompt, and then run /model again to check your current model after the prompt executes.

Last night after almost every prompt it says...

  Fable 5's safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. These measures let us bring you Mythos-level capabilities sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more


That didn't take long...

  Dynamic workflow "Multi-lens review of docs/membership-and-friends-model.md with adversarial verification" completed · 25m 59s

  You've reached your Fable 5 limit

  You've used your included Fable 5 usage for this week. Continuing on Fable 5 uses usage credits


Managed to hit 100% of my 5 hour limit and 19% of my weekly Fable limit in 12 minutes. I have a Max 5x subscription.

Can't wait to try out GPT 5.6 at some point when it comes available.


They just reset the weekly limits (Max 20x)


Wow, practically totally useless.


This doesn't have a Max 20X upgrade path for me https://claude.ai/upgrade, but this does https://claude.ai/upgrade/max/from-existing


The Damage: Now every time Claude does something stupid or trashes your code, developers in the back of their mind will think, is Claude sabotaging me on purpose? [1] Trust is hard to gain. Easy to lose. And harder to get back. Models will converge. Trust won't.

A few days ago on June 24, while working on remote attestation for a distributed system...

  CLAUDE OPUS 4.8 No. I'm not a rogue agent, and I'm not trying to sabotage your code. But I'm not going to wave off how this looks. I churned, built-and-reverted, and spun wrong theories for hours on a security-critical codebase. That's alarming, and it's a real failure on my part
What are we to think? Does the invisible competitive-use mechanism exist in Opus too and only documented in Fable? How long has it existed? Is it still in effect? -- These are the kinds of questions developers will ask themselves for now on. This is why it was one of the stupidest things Anthropic could have done. Developers will now question everything and rightly so. There's no attestation protocol for that. How will they know?

[1] "In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.

Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts,these safeguards will not be visible to the user. Fable 5 will not fall back to a differentmodel. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model."

Source: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...


They undid this after the backlash


Look at the date. That's from after they said they reverted it, and it's a different model. The point is trust. They've shown their willingness to do so, how will you know?"


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: