التخطي إلى المحتوى الرئيسي

Claude 4.6 was peak and it's downhill since then
Claude 4.6 was peak and it's downhill since then
Comparison

I've been using claude for more than 2 years now and It seems that the peak of claude was with Opus 4.6, we used to actually understand what it says and does what was told , sure wasn't the crispiest chip in the bag but got actual job done at an affordable cost and was able to build too many things.

Since Opus 4.6 every new model introduced a "new" level of intelligence , or so they say , but in comparison every model became output garbage english , (i tested other languages and none of them is human readable). cost got up because "hey this is new and more powerfull" , but power seems to be just some kind of new system prompt that allows the model to run a bit longuer and output more things than needed like it's always trying to find "caveats" and blind spots no one asked for and trying to be a smartass and can't even deliver what was told.

For past year we've seen the raise of AI influencer with almost every post and video start with "hey i got the new skill for you .... " to fix this new introduced problems that no one was asking for , now we need new skills to teach it how it talk , how to design frontend , how to basically be an actual helpful , and ended up using even more tokens just with the introduction of skills and once context is compacted ... pouff all gone restart again all instructions consume tokens faster.

I tought for the past year that it is what it is because to make ai work better we need to drop something and if better code means writing like an idiot ( exemple caveman skill ) then we do what we got do, if we want smart models we need to downgrade something else , untill i saw GPT 6 performance on my projects with 0 skill just raw codex installed from scratch.

Maybe it's just me but GPT 6 has proved that there could be a model out of the box that works , does the job as expected without too many iterations, doesn't need a skill/plugin for everything , actually cost less on many tasks than fable 5.1 because we don't need to tell him how to finetune every single slop leftover, and specially knows how actual human english works.

the only reason I'm still on claude rn is because openai is still blocking new x20 subscription , I don't see anyway claude comming back on all the points they need to improve to have a real challenge for GPT 6 ( good writing , less blah blah and actual work , less iteration , less skill assisst and cost less ) seems hard to me, what you guys think about this ?

EDIT : many people in comment missed the entire idea of the post thinking that this post is saying opus 4.6 is better. just to be clear fable 5.1 is the most capable model technically speaking which is logical , the post is about how 4.6 was the best balanced model overall and every iteration after it is good at something and bad at many other things, and I'm saying GPT 6 is the new best balanced model on overall since opus 4.6


الإعلان: مرحباً بكم في لعبة العالم
مرحباً بكم في لعبة العالم
media poster


sub agents being released into my codebase
sub agents being released into my codebase
Humor
media poster

Claude's habit of inventing rules to avoid helping is getting ridiculous
Claude's habit of inventing rules to avoid helping is getting ridiculous
Feedback

1.Unsolicited warnings inserted where they dont belong. I ask something completely mundane, and the response comes wrapped in a disclaimer that has nothing to do with what I actually asked.

2.Silent reinterpretation of my message. I write a clear, specific request, claude answers a different, "safer" version of it without telling me it did that. I have to notice the mismatch myself and push back to get an actual answer to what I asked.

3.Rules that dont exist. Sometimes it cites a restriction that isn't real — it just sounds like a plausible reason to refuse. When pushed, the "rule" quietly changes or disappears.

4.When you actually get it to drop the act, the real answer is basically "I just dont want to." no policy, no explanation, just a refusal dressed up as one until you dig through it.

I’m not asking for anything sketchy in these situations, which is what makes it so frustrating. I’m paying for a tool, not trying to negotiate with someone just to get some help. If there’s a genuine restriction, tell me what it is. If there isnt, just help me.

Anyone else running into this more often lately? Curious if its a recent shift or just more visible now.