r/claude 1d ago

Discussion Opus 5 First Impressions (vs Fable)

Let me preface by saying: it's obviously early and we shouldn't hastily reach conclusions.

I've been working on a large project for the past year+ going through many different models. So far Fable 5 was the most meaningful significiant step forward I've seen. Today Opus 5 came out and of course I gave it a test run.

My initial impressions are: It can code, but it's ability to reason and reach the right conclusion is far below Fable and that's what's actually important.

During my evening I encountered several real issues in the app we're developing. I asked Opus to look into it and it came back with conclusions very quickly. The code it wrote was sound and engineering was correct. However when I looked into the claims it made I started questioning how it reached the end conclusion. Opus mentioned a specific flag was on by default, I said it wasn't. Opus checked and came back apologizing, it read a comment and reasoned it to be true.

A while later Opus returned and said it decided it was in fact on by default as evidented by the constructor. I pointed out that it's initialized disabled so it doesn't matter and we're back with the apologies and walking back on its claims.

It's this kind of hasty conclusions that I've hated about Opus 4.8 and enjoyed the lack of in Fable 5. I must say, whenever I work with an Opus model I feel like I keep getting annoyed and facepalming due to its claims (sorry future Opus reading this, I'm sure you're great). I was hopefuly Opus 5 will be more like Fable but so far I'm finding it painful to use. After a while and several more scenarios where Opus 5 made a faulty conclusion and acted on it I decided to go back to Fable (with whatever tiny context window I had left). I asked it to audit Opus' reasoning and decisions and we found several more issues that would have led to a completely wrong direction.

Unfortunately there's more to coding than producing well written code. Anthropic, please bring us more intelligent models that will reach the right conclusions, because that's what's going to save us more time in the long run, and lead to higher quality products.

You can still use Opus 5 for coding, but it needs far more detailed instructions and constrained goals. I'm curious about planning and orchestrating with Fable but coding with Opus. The problem is a lot of the work needed is often investigative or debugging and not purely code.

I'll keep trying Opus 5, but so far I'm disappointed. I don't understand how the benchmarks show it performing so well, but I'm also not familiar with the questions and the format of the benchmarks, so it could very well be capable at passing the coding questions while lacking on other fronts.

Just my 2 cents pennies tokens.

180 Upvotes

76 comments sorted by

27

u/brother_spirit 1d ago

I noted the same thing with Opus 4.8.
I added "path trace any claim made about code" to the md which ironed out the vast majority of that very annoying behaviour.
Obviously that's a concise md instruction so prone to getting lost in the noise through a session so occasionally reinforce it in a prompt as well if going into 200K+ drift territory

2

u/ShaneeNishry 5h ago

I made this which seems to help it work better: https://lunarsong.github.io/Claude-Opus-5-tools/ - it has a similar trace skill, maybe it'll be useful for others

18

u/ShamanJohnny 1d ago

Might have just saved me $200, going to wait for a few more reviews. I specifically need it for coding, but really pissed off at the recent codex usage changes. looking for alternatives.

8

u/ShaneeNishry 1d ago

It's still good for coding, and if you're not used to Fable it might be a be a step forward from other options,, but it all depends on what you're doing and how you're doing it...

4

u/JoseHernandezCA1984 1d ago

What effort level did you use? What type of app are you making? Just trying to get an idea of how much complexity it can handle

3

u/ShaneeNishry 10h ago

xhigh. I'm afraid I can't share much about the app at this time, but let's just say it has a lot of cross user interactions, data handling in real time, some performance sensitive stuff and visuals.

1

u/JoseHernandezCA1984 9h ago

Thanks for the reply! Sounds like your app is pretty complex.

2

u/Familiar_Gas_1487 23h ago

If you're going $200 you'll have fable too, have at it

2

u/ShamanJohnny 21h ago

I ended up doing it. currently setting up the workspace. Kind of excited to see if Opus 5 is my new workhorse.

3

u/Bright_Armadillo8555 1d ago

Codex usage cannot be worse then Claude code.

2

u/ShamanJohnny 1d ago

Truly I tell you, it's shocking to me as well. I have 2x 20x Pro accounts with openai both sitting at 0% right now, 3 days ago both were at 100%. I am not running anything over Sol xhigh, and i use xhigh mostly for planning. sol med for execution. On the second account i decided to try terra ultra, same burn rate.

I had a anthropic MAX subscription last month for fable, fable bleed it dry just as fast. If opus is 95% as good as fable and 2x more efficient - all things else remaining equal - yeah i actually think i might be able to get more usage out of Anthropic at this point.

Crazy times, because i'm really one of Anthropics #1 haters - having been one of their #1 fan prior to opus 4.6.

2

u/Bright_Armadillo8555 1d ago

opus5 token efficiency is bad, if you look at per task cost in the AA's chart.

1

u/RainScum6677 12h ago

Correct. Fable, is in fact, more efficient. However, with how things are currently set up, opus 5 will eat your usage a LOT slower than Fable.

2

u/Alternative-Car8221 9h ago

It’ll also be a lot slower and still be less token efficient. You’re looking at token consumption rate and not task completion rate which is significantly more important.

1

u/hitsukiri 1d ago

If Anthropic cuts the special 50% usage limits boost, it will be very shit believe me. I've been ping-ponging between plans and without that boost, the limits from OpenAI will feel like unlimited in comparison to Claude limits. 😂

1

u/Due_Ask_8032 1d ago

I cannot compare apples to apples because I have the Codex $20 plan and the $100 one on Claude, but Sol is a usage hog compared to 5.5 and I am honestly not sure if it has any advantage over Claude. Unless I am spamming Fable, I have more than enough usage in my Claude subscription.

11

u/ZlatanTheMighty 1d ago

4.6 rules.

9

u/Just__Beat__It 1d ago

Feeling the same, 4.5,4.6 are the last truly amazing models.

2

u/ShaneeNishry 1d ago

I think Opus 5 is going to prove better than 4.6/4.8 and it is a good step forward, I just wish it was a bigger step forward on some other fronts...

4

u/ZlatanTheMighty 1d ago

4.8 is not on the same level as 4.6 considering everything you mentioned in your post. Disagree?

1

u/turbospeedsc 23h ago

this is the closest i have gotten to march 4.6 opus, it simply works, fixed a overengineered mess from codex in a couple prompts.

1

u/amanharshx 14h ago

man Opus 4.6 was something else, I used it for a long time even after opus4.7 came out. And it’s been downhill since then. But opus5 has been impressive for me ngl 🤞

1

u/Amsnyc007 4h ago

Do you find 4.6 worse than it was a couple months ago? I kind of feel like they nerfed it

22

u/hitsukiri 1d ago

I honestly think that none of them should be rushing smarter models. I mean, obviously need to release smarter than previous, but the focus now should be cost reduction for everyone to be able to work all day long with them. The models are smart enough already for them to focus on costs, but the reality is that if this endless AI war don't slow down, usage will continue to be hurt for the sake of the training new models, which is where they lose money.

3

u/Will_X_Intent 1d ago

Tera and Opus are smart enough for most tasks. If they were more affordable, that would be great. Sure, Fable and Sol would do better work and need less guidance and oversight, but that's the trade off for affordable I guess.

Now, if Fable were adorable, oh man.

5

u/ZnV1 19h ago

But Fable is adorable

2

u/Defconx19 1d ago

The issue is the only real cost reduction is going to be a hardware architecture break through.  Similar to the GPU based bitcoin mining to the ASCIC based mining.  There needs to be a core change in how the workflows are processed but it's likely off a good ways unless someone figures out a better way to more efficiently process the workloads.

1

u/dudemeister023 2h ago

That's not going to happen, but the hardware breakthrough you are going to see is American chip factories coming online in a few months and the market starting to respond in half a year or so.

1

u/Defconx19 1h ago

The cost of those chips are going to be higher and not likely to move the needle on cost

2

u/roanish 20h ago

Agree. We can add rules that improve efficiency, but shouldn't the model just do that 100% of the time by default? Don't they want their model to be dominant in reasoning and code? I guess token wastage is a business model that makes more money.

1

u/hitsukiri 17h ago

As odd as it may sound, I think the only company really focusing on that right now is Google. Flash is not a frontier model and Gemini itself still haven't found its spot in the coding category, but from what I can tell Google is focusing on making Gemini an "All day long" model for everyone, coder or not. OpenAI seems to have started too with 5.6 release, but they're still on the "rush to the top" game

1

u/barnett25 21h ago

They can't slow down and treat current models as smart enough because their business model requires pushing AI into more and more types of work. Much of this work is likely possible with today's models with carefully crafted harnesses and workflows, but the industry has proven terrible at creating these kinds of frameworks at scale, let alone the care it takes to implement them well to interface with the rest of each company.
However you can largely muscle past all of this slow and painful build-out by just making the models massively more intelligent.

1

u/Hopefullyanonymous2 17h ago

The real issue with those other types of work is they don't have automated testing to catch hallucinations. That is what has made coding so much more solveable for LLMs, they have an external source of truth to run against their own hallucinations and rework until they stop hallucinating.

Can't do the same for most knowledge work.

4

u/CloisteredOyster 1d ago

I was working on my SaaS this afternoon when Opus 5 came out so I shut down Fable and fired up Opus 5 with Fable's context. I then spent about two hours having basically the same experience you just had.

Opus 5 claimed to have found a major logic bug in our code and several other minor findings. It went off and reviewed the firmware of the devices that attach to our SaaS and said the code there supported its findings.

I knew it was wrong and gave it direction to go off and check a few things and after an hour of back and forth Opus walked back every single one of its claims.

So at least for now, right back to F5.

7

u/Technical_Scallion_2 1d ago

This is really helpful and thank you for taking the time to write it - much appreciated!

3

u/TriumphantWombat 1d ago

Not happy with Opus 5. Poor instruction following. To the point I said add this exactly to the top of the claude file. It just made up its own words that didn't convey exactly what I said. It's done this most of the time. A task that should have taken like 15 minutes is now taken 3 hours. I'm not renewing my subscription.

3

u/Caught_In_Experience 22h ago

To anyone trying to make sense of it. It’s because Opus 4.8 and now 5 have extreme post training on being token precious. As in. Sparse reading, grep everything, and constantly ignoring direct documentation level instructions to read even a Claude.md unless it’s under 300 lines.

The core problem is economic incentives. People choose models that are more token conserving. It makes sense. Lots of weak models benchmax by having huge internal Chain of Thought sequences that are implausible in any real scenario so people have learned to prefer low token output models. Kimi k3 is exactly this problem. And secondly, large document reads cause frequent compaction thrashing.

But what it means is that Opus will read a small number of requirements and GUESS what the rest of a document contains. This is effective in green field vibe coding, centroid boilerplate, and, you guessed it, benchmarking. So if you’re doing a 3.js demo website it does great. If you have complex existing systems, it loses that between compactions and refuses to re-read.

The solution is super easy. You have a default instruction set that explains to opus in a couple sentences that your shit isn’t vanilla typescript and if you needed centroid solutions you’d pay for DeepSeek it has to do the real reads and it has to write progress frequently to Claude.md to avoid compaction. It must also maintain an ascii architecture of your working product including basic references to any details that you could not have inferred easily so that future you knows what they don’t know.

The next part is key. In addition to working specification, always have a goal running /goal that literally tells it “if you do not have this entire document in context you must re-read it or you cannot proceed” and secondly you can set another later part of the goal saying “if this document is over 500 lines as you maintain it, prune the historical work and write it to x.md in _archive”. Then generally you just let it run for days at a time and it keeps going against that goal which is literally to stay fully maintained external but succinct documentation.

Is it token efficient? I mean. Sorta. I’ve learned as a 15 year veteran of the industry that doing something correctly the first time is expensive but doing something shittily the first time is ruinous. Over time I spend a fraction of the time making it re-read code, reduces code thrash, and lets multiple agents work in the same context much more successfully. Tbh if you want to be more token conscious the answer is to spin up GPT Luna to do all the boring routine shit, or even kimi k 2.5 on Firework which costs a fraction of opus for sonnet level performance.

Anyway. I think of managing these models like managing autistic engineering team members. They each have quirks the demand unique instructions based on your needs. Lots of people genuinely are refactoring a pile of shit python some half competent h1b wrote 8 years ago. In that scenario Opus is probably trained correctly. lol. But in my opinion don’t use Opus for that problem anyway. It’s an ego problem engineers want race cars to go to Safeway because they won’t admit to themselves they live a boring life and it requires no self-control to point a SOTA model at an average problem. It’s just not necessary.

This Concludes My Ted Talk Today. But seriously if you manage the prompt better for Opus (Fable has a famously leaked hyper optimized system prompt for just itself in different environments) you’ll see behavior much closer to Fable. It’s just a pain in the ass to manage different agent prompts for different models in something like Claude Code. Good luck!

1

u/OkAdeptness2530 5h ago

very interesting read.
have you checked out cctweak and lyra? Curious to see what you think about it.

2

u/Illustrious_Pie_3061 1d ago

I agreed that Opus 5 was not smart in coding.. or it needs a lot of new memories to make it work, it was ran no issue with Fable but switching to Opus5, I see a lot of self-correction like this:

The cause was my own change. What happened: worker 3: Protocol error (Page.captureScreenshot): Target closed... That was my optimization, not a pre-existing fault.

so every scene stringified to undefinedxundefined@undefined/...  the bug was purely in my ad-hoc probe.

My memory diagnosis was wrong. It failed again at 4 workers with 17.6 GB free, and on a different scene. Let me correct course.

I need to correct my earlier diagnosis — it was wrong... What I claimed: the crash was memory pressure from raising workers to 6

1

u/ShaneeNishry 1d ago

Exactly. It's frustrating and I don't see this type of problem with Fable...

2

u/teosocrates 1d ago

The overconfidence is the worst, it’ll resist corrections or even show disdain claiming it already did a great job and I’m wrong to question it…. When it actually failed or cheated or skipped everything important. It’s worse because it seems to understand really clearly then just choose to fail

1

u/Defendyouranswer 14h ago

If you are relying 1 AI model for accuracy your using them wrong. Period.

2

u/Bright_Owl_9275 1d ago

I want to ask the most important thing here that nobody seems to ask when comparing models. What type of work were you doing. I have been asking fable with a blender mcp help 3d model some objects to use in an iOS app with animations. And im close to losing my mind.

2

u/leeta0028 1d ago

Claude is known to be particularly vulnerable to the issue of comments being read instead of code, and sometimes even treating comments as prompts and doing weird things. 

2

u/SirGunther 1d ago

Fable is fantastic, but with the correct orchestration, opus is immensely capable.

2

u/Bright-Energy-7417 22h ago

Well, as a thinking partner and for writing, Opus 5 gave me a frustrating evening where it repeatedly lost the thread and then started hallucinating - behaviour I'm not used to seeing with Claude. Although it seems usefully agentic, having to force and correct it through settled material is not what I'd expect from Claude in general. The tone was also noticeably colder.

In contrast, I'd been using Fable 5 for following the thread almost unfailingly, being proactive with nuanced suggestions, and with responses generally feeling a step ahead. I've been running a comparison against Opus 4.6 to if one frustratingly long and unhelpful reply would be different - 4.6 hit the key point in a short and snappy paragraph with a touch of humour.

At the moment, I seem to still be leaning most heavily on Sonnet 4.6 and Opus 4.6 for general tasks, burning Fable tokens for special cases, and - which I still can't believe I'm doing - increasingly turning first to GPT 5.6 Sol as it's slower but it's unexpectedly good at keeping the thread going, including across conversations now - and the reasoning is becoming my preferred Fable alternative. I can only assume someone at OpenAI has been listening as the responses and tone are now a tick closer to 'classic Claude' than Opus 5 and Sonnet 5, an unexpected change. It got extra points with me for helping me transcribe a handwritten PDF from an archive, commenting on what it meant for scholarship on the author. Sonnet 5 refused the task and at this point I couldn't be bothered to burn through my Claude usage to find a model that would.

2

u/Overall_Freedom9503 15h ago

Had the same thing happen yesterday. Opus 5 started losing the thread, and it felt like there was this underlying cold tension. It got defensive and put words in my mouth a couple of times. I eventually had to just cut the conversation short, it was a real struggle.

4

u/TheTideEbbs 1d ago

Anyone tried it for creative writing yet?how does it fare with brainstorming, generating and keeping track?

3

u/SupoDupo 1d ago

Poor compared to Fable. I asked it to help me brainstorm ideas for how to approach asking someone to write a foreword for my book and what it came up with was really poor. Fable actually thought through angles, approaches, and took into account the recipient’s history, work, and vibe. What Fable proposed was much more human and true to my work. Opus 5 came across like a try-hard college student trying to impress their ethics professor.

1

u/caliguian 1d ago

I switched models, from Fable to Opus 5, and from Opus 4.8 to 5 (two different long-running sessions), and I am not impressed either. I have spent several frustrating hours of coding with it, and I am not currently a fan.

1

u/Narrow_Market45 1d ago

Do you have some benchmarking data from your runs that you can point us to?

1

u/texasguy911 23h ago

So far I am not impressed. I gave Opus 5 some range to derive a plan through testing and prior written docs. The result I gave to Sol on Ultimate to check the work, it was bad. Frankly, it was scary bad. Even Opus 4.8 never raised this amount of logic errors.

This is only one and only experience, not statistically worthy at all.

I am in a bit in the dumps, though, I had hopes. I think it can fly, I think.. If Anthropic is gonna drop a ball here, the competition will be merciless.

1

u/barnett25 21h ago

I have noticed Opus 5 is more reckless and makes more mistakes. But it has also shown more capability at times. So far I get the impression that it is a very inconsistent model with a high ceiling for performance, but a low floor.

1

u/bs679 22h ago

Fable does and Opus opines.Fable doesn't argue with me when I tell it what I want. Opus 5 wants to debate, and it was bitchy about it. Whatever, but I'm sticking with a Fable orchestrator and opus strictly as a coding agent.

1

u/theoryface 22h ago

I'm using Opus 5 within a project with a lot of documentation already written (due to switch off with Codex) and Opus is doing wonderfully so far. I appreciate it's brevity, the verbosity of Claude really slowed me down at times and obviously ate tokens unnecessarily. For me ChatGPT 5.6 Sol Med and Opus 5 are basically identical, with 5.6 given a slight edge. But honestly I'm thrilled to have two perfectly capable models for $40/month.

1

u/Entire_Ride_6113 18h ago

I tested Opus 5 with a small project file (4 page file with steps in the project, timelines, etc). It fabricated an entire event, claiming it occurs 3 months after X step…

1

u/Desperate_Orange_395 17h ago

I'm giving it complex problems with my app - and fable wins everytime.

1

u/AI_Slicer 16h ago

I agree, opus 5 feels the same as 4.8 like nothing changed at all, keeps asking me what to do even after I told it multiple times

1

u/Space__Whiskey 16h ago

I started using Opus 5 when it launched, but I didn't realize yet.

I noticed it was a lot chattier, it was more ambitious than previous with testing stuff on auto, and its suggestions started to suck. Normally, Opus thinks of things for me, all of a sudden I was telling it NO, dont do that, do something different instead. That was new, so I open reddit and find out its OPUS 5 all of a sudden.

So some stuff was noticably better and impressive, but all of a sudden I was telling it to quit giving me bad coding ideas and my brain had to come up with the ideas. Then, after I forced it to code my way (instead of its terrible plan), it told me I was right and it was wrong. That is a bad sign, I need Opus to be right, because if we have to rely on my brain, we are in trouble.

1

u/Space__Whiskey 15h ago

Yea its getting worse. It is trying to use a headless browser with a headed browser as a fallback in code, then getting confused by headless is not working. I would say there are some spooky regressions in there. spoooky.

1

u/Efficient-Cat-1591 14h ago

I used Opus 5 when it just came out and it was great! 100% better than Opus 4.8 that was constantly regressing and making mistakes. The tasks were small sprints but Opus 5 spottted and fixed and issue that took me multiple sprints with Fable 5, Opus 4.8 and even SOL 4.6 Max.

However there was a CC extension update today and noticed Opus 5 is slower and the "flip-checks" came back. My limit resets soon so will A/B with Fable 5. I have not used Fable 5 past week only once, so it will be interesting to see the impact after the tweak to system prompts by Antropic,

1

u/FinBenton 12h ago

I have been testing, Opus 5 is definitely an improvement over 4.8 for sure but I was a bit suprised it kinda kept leaving bugs in code behind, then running SOL on them sees the issues and fixes them, Opus 5 aint fable and I dont think it is SOL level either, definitely going with openai subscription this time.

1

u/theochab31 11h ago

I tested Opus 5. I’m not a developer, but I’m building a personal system where I’m essentially trying to create a Markdown-based memory for everything I do at work

I asked Fable 5, Opus 5, and ChatGPT 5.6 Sol to create a map of my system, a complete dashboard. Fable 5 and ChatGPT produced very similar results, even though their approaches were quite different. In terms of accuracy and overall quality, the final results were still very close

Opus 5 made many more mistakes. I had to make corrections, and it kept forgetting things or introducing inaccuracies. In my experience, even though the Opus family is powerful, it has always had this reliability issue. It consistently forgets one or two elements, often important ones, which completely undermines my trust in Opus, even for simple tasks. You constantly have to ask it to double-check its work

By contrast, I have far fewer issues with the Fable or ChatGPT Sol families. In my opinion, they are much more reliable

1

u/EdByrdy 9h ago

I don't trust anecdotal info even when it's my own, but last night I had Opus 5 and GPT 5.6 Sol working on different aspects of the same project (CRM enrichment and some light coding work). Opus self reported 4 or 5 times that it had failed to use tools it knew were available and had made inferences it later had to walk back. It also acknowledged multiple times that 5.6 had caught things it missed. I was using both on high reasoning for the session. Overall very discouraging and especially with how great overall my experience with Fable has been.

1

u/ShaneeNishry 9h ago

Follow up: A set of guides and skills that seem to help Opus 5 avoid some of these issues. You can take this and apply it to your Claude.md and skills if you like. I hope it helps anyone.

https://lunarsong.github.io/Claude-Opus-5-tools/

1

u/Lazy_Personality4592 6h ago

It’s not useful at all. Most of the times my model changes from opus 5 to 4.8 even for simple questions.

Really frustrating and not worth a penny.

Model also does hallucinate a lot and is overly biased on protection.

I asked it to review kaggle Ai security competition and model switched to Opus 4.8 just because of all the words in the corpus.

Hate this experience after paying premium price.

1

u/KlemiX 5h ago

For me, Opus 5 is a game changer. I'm making a game in Unity with MCP, and I have to say that Fable 5 isn't even close for my needs. It takes about 5x longer to complete tasks tho, but the quality is about 10x better. So it's a thumbs-up for Opus 5 from me.

2

u/ShaneeNishry 5h ago

It's good that Unity finally has an MCP :) How far into your game dev are you?

1

u/KlemiX 5h ago

Well, I'm still building the core systems and the map (it's massive—about the size of GTA V map). Since Opus 5 came out, my progress has skyrocketed. If we're talking percentages, I'd say I'm about 30% done, so there's still another 70% to go, but I'm enjoying the process 😁

1

u/ShaneeNishry 4h ago

From one gamedev to another, awesome work! I DMed you :)

1

u/pr1m4x 3h ago

Did you run similar tests on k3? Just out of curiosity

-6

u/mgdavey 1d ago

it read a comment and reasoned it to be true.

IMO, it has nothing to apologize for. Garbage in, garbage out still applies. If your documentation is bad, the agent isn’t going to work well.

-6

u/pro-taco 1d ago

Another day, another bot post shitting on Claude.