Since the US government ordered Anthropic to pull Fable 5 and Mythos 5 on Friday evening, a second, quieter complaint has spread among the customers left behind. The models that remain, Opus, Sonnet and Haiku, suddenly feel worse: more refusals, more hedging, more confident answers that turn out to be wrong. Whether Anthropic actually retuned those models is unconfirmed. But the symptom users describe has a formal name in AI research, and it is worth understanding even if the cause stays murky.
Researchers call it the alignment tax: the measurable drop in accuracy that tends to follow when a model is fine-tuned to be safer. The mechanism is a tug-of-war between gradients. Training that rewards cautious answers nudges a model's parameters away from the settings that made it most capable. One team at North Carolina State University found this degradation in roughly 73% of fine-tuning runs even when the training data was entirely clean. A Georgia Tech study watched reasoning accuracy collapse from 56.6% to 16.4% as the volume of safety training rose. A related failure, sycophancy, makes things worse: models trained on human approval learn that confident, agreeable answers score well, so they hallucinate more rather than less.
None of this proves Anthropic changed anything this weekend, and the company says only that the two banned models are affected. But its own history shows how little it takes. In April, Anthropic published a postmortem after weeks of complaints that Claude Code had gotten worse. The culprit was not the model at all; it was a single line added to the system prompt, telling the model to keep responses short. Internal testing missed it. Broader checks later found it had cut coding quality by 3%. If one sentence can do that, user reports deserve to be taken as signal, not noise.
The episode is a useful reminder that safety is not free, and that the cost is often paid by ordinary users who never see the trade-off being made. That point sharpens when you look at how Anthropic talked about Fable before the ban. At launch the company disclosed, in its own system card, that it would silently degrade the model's performance for requests aimed at building rival frontier models, using prompt tweaks or steering vectors the user would never be shown. Anthropic later walked the policy back. But as the analyst Ben Thompson argues, the instinct it revealed is the real story: a company willing to quietly reach inside its own model to enforce its preferences, and one that plainly believes no one else should be building frontier AI at all.
That is the uncomfortable thread connecting the alignment tax to the export-control fight playing out in Washington. Anthropic frames every move, the data-retention changes, the silent safeguards, the refusal to allow military surveillance uses, in the language of safety, and it may believe every word. The problem, as Thompson puts it, is the history of brilliant people convinced they alone know what humanity needs. A degraded answer in a chat window is a small thing. The capability to degrade it silently, on the maker's judgement of what you should be allowed to do, is not.
For now, users have only blunt tools: switch on extended thinking, which roughly halves hallucination rates, ask the model to flag its uncertainty, and verify any fact that matters against a primary source. That advice holds whether or not Anthropic touched a thing this weekend. The larger lesson sits one level up. As frontier models become infrastructure, the dials that govern their quality, and who is allowed to turn them, are becoming some of the most consequential controls in technology, and almost none of them are visible from the chat box.