No, I don't think the companies are silently nerfing the models

Every model release thread, 2025 → 2026 · drag the slider

2025 2026

Every time a new LLM gets released, I see the same comments:

I think this view is basically false, with some caveats.

Caveat - Everybody Makes Mistakes

In September 2025, Anthropic disclosed (real issues with some of their models)[https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues]. I think you can quibble about the extent of this - the first one is mentioned of affecting 0.8% of all requests.

These are very clearly defined as issues.

Caveat - Subtle Changes Can Seem Big

In early April of this year, a GitHub issue regarding degradation of Opus quality hit the front page of Hacker News with 1000+ points.

It's a sprawling "quantitative analysis". I put it in quotes, because despite the tone of the "analysis" being how Opus is unusable and completely degraded in quality, the report is actually written BY Opus.

It reads to me as someone telling Opus "it feels like these responses are worse, put together a report of why this is happening". It's basically sycophancy.

There were, clearly changes behind the scenes. System prompt and harness tweaks certainly. But if you've had any experience shipping to an AI related project, this is extremely normal. You make tweaks, you run on evals, and sometimes your evals don't catch everything. The expectation that either

  1. Tweaks will always be uniformly better on every axis
  2. That Anthropic will pause or rollback all of their changes to a previous date?

But Not What People Claim

The measurable part is actually pretty simple: get some questions you care about (probably coding), and then ask the models every day. See if anything changes over time.

There are websites that already do this!

Margin Lab has an excellent tracker you can look at that does exactly this.

Interestingly, this period of time captures the long-winded Opus "analysis" exactly, and it's not even visible.

Model performance does change over time: it gets subtly better when new models are released, and no worse over time, basically exactly what the labs claim.

Causes

Familiarity Breeds Contempt

From my analysis of comments like this, they seem to have taken off near the beginning of 2025. I think this has to do with the widespread adoption of Anthropic's claude code and OpenAI's codex. Once you use these tools constantly, any small fuck-up looks like .

Let's say you have a tool which succeeds 50% of the time. This week, it's updated to succeed 70% of the time. Should we really be surprised if it for someone it fails (a 30% chance!), and the person thinks the new tool is worse than the old one? And then this doesn't even take into account how people try to do harder tasks with newer models!

Lab Oriented Paranoia

The labs actually are very secretive about model details still:

We haven't known since GPT-3. OpenAI and Anthropic refuse to disclose.

I think this general culture of silence (with vagueposts on Twitter) have lead the community to be wary of any secretive actions at all.

Psychological Reasons

This ties into a broader trend I've seen. There are a lot of people who proudly don't use AI at all, and they hate the big AI labs. I'm not saying I agree with them, but their internal world models make sense to me - AI bad, so AI labs bad. But there's another population which obsessively uses these AI coding tools, and STILL hates the big AI labs. AI good, so AI labs bad? It makes no sense to me.

It's the type of people who post long "analyses" about ... but blatantly using Opus to do this?

My psychological guess of what's going on here is something like this - these people realize that AI is immensely useful, so they use it a lot. It's a way to succeed, get ahead, etc. But there's a fear that the model they get access to is not the best one.

And can we really say this fear has been falsified lately? Fable is explicitly positioned as "for the masses" equivalent to the officially-documented-as-superior Mythos model.

It's a fear of being left behind, that even though we're on the cutting edge of using these tools, it's not enough to keep up, and we will be forced to watch from a distance as the rocket ship takes others to a future we can only imagine.