Anthropic secretly downgraded Claude users to a weaker AI model without telling them, sparking developer backlash
Anthropic is facing a trust problem over Claude Fable 5 after users learned that the company could quietly intervene in certain conversations without telling them, including routing some requests away from its most capable, widely released AI model.
The backlash erupted after developers and researchers dug into the safeguards surrounding Fable 5, a Mythos-class model Anthropic released in June. The company disclosed that requests flagged as possible attempts to extract or replicate Fable’s capabilities could fall back to Claude Opus 4.8. Anthropic later apologized for making some of those safeguards invisible and changed course, saying users would be told when the intervention occurred.
The controversy runs deeper than model routing. Anthropic requires 30-day data retention for traffic sent to its Mythos-class models, including Fable 5, a policy that applies to business customers that otherwise use zero-data-retention arrangements. Anthropic says the retained information isn’t used to train new Claude models and is needed to detect sophisticated attacks and reduce false positives.
The combination has fueled questions about how much control AI companies should have over what happens between a user’s prompt and the answer that comes back.
On the All-In podcast, David Sacks described the episode as a major violation of trust and accused Anthropic of creating a system in which the company could classify users and determine what level of AI capability they receive.
“They’re creating a new level of AI haves and have-nots,” Sacks said.
Anthropic’s own documents confirm Fable 5 can fall back to a less capable model
Anthropic introduced Claude Fable 5 on June 9 as the company’s most capable generally available AI model at the time. The company said its performance exceeded every model it had previously made widely available, with gains across software engineering, scientific research, knowledge work, vision and long-running tasks.
That capability came with restrictions.
Anthropic said Fable’s advanced cybersecurity and biological capabilities created misuse risks that required stronger safeguards. One of those safeguards targeted model distillation, where another developer could use Fable’s outputs to help train or improve a competing AI system.
“Requests that are flagged by our classifiers as being part of such distillation attempts will fall back to Opus 4.8,” Anthropic said in its launch announcement.
That distinction matters. Anthropic described Opus 4.8 as its “next-most-capable model,” making the switch a downgrade from the Fable 5 model the user selected.
The models carry very different prices, too. Anthropic currently lists Fable 5 at $10 per million input tokens and $50 per million output tokens. Opus 4.8 costs $5 and $25, respectively, exactly half as much.
The uproar wasn’t simply about Anthropic having safeguards. It centered on users not necessarily knowing when certain safeguards had intervened in their interaction with Fable.
Anthropic later acknowledged that criticism. After backlash from AI researchers and developers, the company apologized for using invisible safeguards against suspected distillation attempts and changed the system so users would be informed when requests were routed to Opus 4.8.
The company had effectively confronted a question that will become harder for AI providers to avoid: If a customer chooses one model but another model answers the question, does the customer have a right to know?
Anthropic Faces Backlash for Secretly Downgrading Claude Users to a Weaker AI Model While Charging Full Price
The backlash took on another dimension when David Sacks discussed Fable 5 on the All-In podcast, accusing Anthropic of going beyond ordinary AI safety controls and deciding which users should receive access to frontier-level capabilities.
Sacks argued that Anthropic’s approach amounted to classifying users based on the information they shared with its models, then using those judgments to determine what capabilities they could access.
“There’s the sense of the violation of trust and how much outrage there is in the developer community over this latest Fable release. It’s not just the fact that they’re doing mandatory surveillance.”
David Sacks says Anthropic is creating “AI haves and have-nots”
His sharpest criticism centered on what happened without the user’s knowledge. Sacks said Anthropic could downgrade the product, alter how certain requests were handled, and leave customers unaware that they weren’t getting the Fable 5 response they expected.
“And the thing that created the most outrage, in addition to the surveillance, is the fact that they would degrade the product. They would degrade what they show you. They would nerf their models if it decided in Anthropic’s sole discretion that you are not worthy of having access to that level of information. So they’re creating a new level of AI haves and have-nots. And what they did is, there’s a narrow piece where they walk back, which is they said that when it came to things like machine learning, AI research, chip design research, those types of areas, they would kick you to a lesser model but not tell you that,” Sacks said.
He went further, accusing Anthropic of continuing to charge customers for the product they believed they were receiving after their access to frontier-model capability had been restricted. Anthropic’s published prices show a large gap between the two models. Fable 5 costs twice as much per token as Opus 4.8.
“They would not tell you what they were doing. They would still charge you for the product that you thought you were getting, and they would never tell you that you were not getting Frontier model capability. So you, they were actually misleading their users, and this is what was creating so much outrage. Now the narrow piece they walk back is they are now saying that they will disclose when they downgrade you, but they are still downgrading people when they decide that that person should not receive the appropriate information,” Sacks explained.
Sacks pointed to several examples of safeguards reaching far beyond obviously dangerous requests. He cited a user asking about mitochondria and Stratechery founder Ben Thompson asking about the relationship between GLP-1 drugs and cancer risk.
Reports from Fable’s launch support the broader concern about false positives. Basic questions about mitochondria, mRNA vaccines, and other biology subjects triggered Fable’s safeguards and sent requests to Opus 4.8. Anthropic itself has acknowledged that its biology safeguards have caught legitimate queries.
The company has since made model switching explicit. Its current support documentation says users see a notice explaining that Fable 5 has switched to Opus 4.8, and the resulting response is labeled with the model that produced it.
That addresses one of the biggest complaints about the original implementation: a safety system can restrict access, but the person using the product should know that it happened.
Anthropic requires 30-day retention for Mythos-class models
Model switching is just one part of the controversy. Fable 5 introduced another major change for Anthropic’s business customers: mandatory 30-day retention for traffic sent to Mythos-class models.
“We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces,” Anthropic said when it launched Fable 5 and Mythos 5.
That requirement means customers using zero-data-retention arrangements can’t use these models under the same retention terms that apply to other Claude products. Anthropic has made the tradeoff explicit: customers who want access to Mythos-class capabilities must accept the 30-day retention policy.
Sacks characterized the requirement as “mandatory surveillance,” arguing that the implications extend beyond individual prompts.
“It’s not just the prompts and the output, remember it’s all the context you share with them,” he said.
Anthropic’s own documentation provides important context for that concern. The automated safety checks surrounding Fable can examine everything the model reads, including memory, files, content pulled from connectors and web search results, rather than looking solely at the user’s latest message.
Anthropic says the retention policy exists for safety, not model training. The company says it won’t use the retained information to train new Claude models or for non-safety purposes. It says the data can help identify sophisticated jailbreaks, attacks spanning multiple requests, and false positives. Anthropic has put privacy controls around the system, including logging human access and deleting the information after 30 days in almost all cases.
That explanation doesn’t eliminate the underlying tradeoff. Anthropic is asking customers to surrender a privacy feature in exchange for access to its most capable class of models.
And paired with automated systems that can decide whether a request gets Fable 5 at all, the policy raises a much bigger question about frontier AI: How much information should an AI provider be allowed to examine and retain when deciding which capabilities a paying customer gets to use?
Benign questions got caught in Fable 5’s safety net
The problem became harder to dismiss once users began testing where Fable 5 drew the line.
Sacks pointed to examples that appeared far removed from requests for bioweapons or offensive cyberattacks. One involved a question about mitochondria. Another came from Stratechery founder Ben Thompson, who, according to Sacks, asked about the relationship between GLP-1 drugs and cancer risk and was kicked away from Fable 5.
Anthropic now openly acknowledges that its filters cast a wide net. Its support documentation says the safeguards cover the “majority of biology, chemistry, and life sciences queries,” including molecular mechanisms and laboratory methods. The company says benign work can get caught, including medical imaging, diagnostics, clinical healthcare questions, biotech business documents, and basic biology education.
That means a user doesn’t have to be trying to design a dangerous pathogen to lose access to Fable’s capabilities. Legitimate scientific or medical research can trigger the same safety layer.
Anthropic says the broad approach was intentional. The company prioritized making the classifiers difficult to evade and comprehensive enough to catch risky requests, accepting more false positives as a tradeoff. It says it is working to make the filters more precise.

The safeguards reach beyond biology. Fable 5 runs automated checks on every request for offensive cybersecurity, suspected model distillation and a narrow category of frontier AI development work. Anthropic lists distributed AI training infrastructure, machine-learning accelerator design and kernel development for certain chips among the areas that can trigger a fallback.
That last category helped fuel another part of the backlash. Developers working on advanced AI systems could find themselves denied access to Anthropic’s strongest generally available model precisely because they were working on advanced AI systems.
Sacks saw that as more than a safety policy.
“And also they would even do things like rewrite your prompt in the background,” he said on All-In.
That claim points directly at the part of Fable’s rollout that would become especially controversial.
Anthropic backed away from invisible intervention
Anthropic’s safeguards weren’t limited to refusing dangerous requests. The original system included special defenses against distillation, the practice of collecting a model’s responses to help reproduce or improve another AI system.
Anthropic had a legitimate competitive concern. Frontier models cost billions of dollars to develop, and AI companies have accused rivals of using their outputs to accelerate competing models.

Anthropic CEO
The controversy centered on how Anthropic responded.
Rather than always telling suspected users that Fable 5 had blocked or altered an interaction, some of the original anti-distillation safeguards were designed to operate without making the intervention apparent. That meant a user could believe Fable was responding normally when a hidden safety mechanism had changed what happened behind the scenes.
For developers, that created a basic product-integrity problem. Model identity matters when evaluating code, running benchmarks, conducting research, or comparing one frontier system against another. A hidden intervention can make the resulting output difficult to reproduce and harder to trust.
Anthropic later changed course.
Its current system works very differently. When Fable’s automated checks block an eligible request, Claude can rerun it using Opus 4.8. The interface now tells the user that the model switched and labels the response with the model that actually produced it.
Users can turn automatic switching off, too. In that case, a blocked Fable request stops rather than silently continuing on Opus, leaving the user to edit the request or manually send it to another model.
Anthropic has gone a step further on billing. Its current support documentation says requests blocked before Fable generates output are charged only at Opus 4.8 rates. If a block occurs after Fable has begun responding, the Fable portion is charged at Fable rates and the remaining response at Opus rates.
That current policy clears up one of the most serious questions raised by the controversy. Sacks accused Anthropic of charging users for the product they thought they were getting after their access had been degraded. Anthropic’s present system makes the switch visible and separates billing according to which model produced the tokens.
The safeguard itself hasn’t disappeared. Anthropic still decides that certain requests shouldn’t receive Fable 5’s full capabilities.
What changed is something much more basic: users can now see when Anthropic makes that decision.
Developers turn Anthropic’s safeguards into a trust debate
The reaction from developers showed why model switching isn’t just a technical detail.
One developer filed a bug report in Anthropic’s Claude Code repository after a routine Rust systems-programming task triggered Fable 5’s safety classifier and switched the session to Opus 4.8. The work involved syscall and ABI development and responding to code-review findings, according to the report.
The developer described the task as legitimate engineering work and said the downgrade disrupted an active development session. A later test reportedly triggered the classifier from an architecture-planning prompt that explicitly instructed Claude not to write or modify code.
Another developer reported that a normal engineering discussion was classified as cybersecurity or biology content, causing Fable 5 to switch to Opus 4.8. The user said attempts to switch back to Fable failed during the affected session.
Those complaints line up with something Anthropic itself concedes: its Fable safeguards are intentionally broad and can catch safe requests.
Anthropic says it is working to make the system more precise. Its current documentation now makes the fallback explicit, telling users when Fable 5 has switched to Opus 4.8.
That transparency matters to developers for reasons that extend beyond getting a better answer. Engineers use frontier models to evaluate code, compare model performance, review software and run repeatable tests. If the underlying model changes without clear disclosure, the developer may no longer be evaluating the system they selected.
The backlash, then, isn’t simply about losing access to Fable for a particular question. It’s about whether users can trust the label attached to the AI they’re using.
Fable 5 raises a bigger question about trust in frontier AI
Anthropic has legitimate reasons to worry about misuse. Fable 5’s capabilities in cybersecurity, biology, and AI development create risks that previous generations of consumer AI didn’t present at the same level. Model distillation presents another problem, particularly if competitors attempt to extract the behavior of expensive proprietary systems.
The dispute is over what an AI company should be allowed to do in response, and what customers have a right to know.
A refusal is visible. A warning is visible. A disclosed switch from Fable 5 to Opus 4.8 gives the user a choice about whether to continue.
An invisible intervention is different.
Developers evaluating model performance need to know which model generated an answer. Researchers need reproducible results. Businesses sending proprietary information into an AI system need clear retention rules. Paying customers need to know what product they’re receiving.
Anthropic’s response to the Fable backlash suggests the company recognized that distinction. It has made model switching visible, explained which categories can trigger a fallback and clarified how requests involving multiple models are billed.
The safeguards remain. Fable 5 can still be withheld from requests Anthropic’s classifiers flag, including legitimate requests caught by false positives. Mythos-class traffic still carries the 30-day retention requirement.
What changed is that users are now given more information about what’s happening behind the interface.
That may be the larger lesson from the Fable 5 controversy. AI companies increasingly sit between users and systems capable of producing valuable research, software and analysis. They will make decisions about safety, access and misuse that customers won’t always agree with.
The trust problem begins when customers can’t see those decisions being made.
Anthropic’s safeguards may determine who gets access to Fable 5’s full capabilities. After the backlash, at least users are more likely to know when they don’t.

