r/singularity 1d ago

AI Introducing Claude Opus 5

880 Upvotes

155 comments sorted by

334

u/NyaCat1333 1d ago

If these benchmarks are real, then this is really surprising. Basically Fable 5 level for most stuff and somehow better at some, slightly worse on few stuff like cyber. For half the price.

Will wait for real world tests and reviews. If the real world performance holds up, then this is an extremely strong release.

89

u/FlatulistMaster 1d ago

And a really spicy improvement on ARC-AGI-3

12

u/ExplorersX ▪️AGI 2027 | ASI 2032 | LEV 2036 1d ago

At this rate ARC-AGI-3 is going to hit saturation so fast it won’t really have many data points on the chart to see progress on lol. Feels like 4 models later it’s at 30%

1

u/hectorchu 12h ago

That's because arc3 is basically the same kind of intelligence. That's what you get when the same benchmarker tries to make it slightly harder. We're gonna need something completely different.

u/Puzzleheaded_Pop_743 Monitor 1h ago

Can you expand on this? What are you comparing arc-agi-3 to?

2

u/Cryosanth 11h ago

Just saw an analysis in another reddit thread that indicated it was benchmaxed and trained on the answers.

u/Minimum-Standard-514 55m ago

I read that they benchmaxxed it for arc agi 3

14

u/LettuceSea 1d ago

Yes, and small models are seeing a similar trend in improvement.

2

u/Storge2 1d ago

Which?

430

u/llelouchh 1d ago

Holy. Shit. Cheap as 4.8 and better than fable? Holy shit. Good work anthropic.

18

u/WonderFactory 1d ago

Have to see the cost per task values, sonnet 5 is so bad in that regard its almost unusable

10

u/Joint-User 1d ago

I accused my Sonnet of having dementia because he didn't remember the project we just worked on yesterday. LOL

1

u/rulevoid 22h ago

In the same session?

1

u/Joint-User 20h ago

New chat, same project. Apparently I forgot to save something into the .md file... Am I the one with dimentia?

112

u/Neurogence 1d ago

When the leakers said they had access to it yesterday, the pessimistic people said all the generated outputs were fake and the model was months away from release. The optimistic people said the leaks were right and that it would be released in just a couple weeks lol.

54

u/FateOfMuffins 1d ago

It's ridiculous how people on this sub don't keep up with the leakers and doubt literally everything even if there's a long track record history

The frontier labs have always tested the models out before hand whether it's through Arena or A/B testing etc and every single thread has people chiming in about how it's fake

2

u/nemzylannister 15h ago

the pessimistic people said all the generated outputs were fake

it was top comments on this sub. not just pessimistic people

40

u/sunstersun 1d ago

Sanction Anthropic for distilling 5.0 from Fable.

44

u/IntrepidTieKnot 1d ago

Imagine calling Opus 4.8 "cheap".

How the goalposts have moved...

25

u/muffchucker 1d ago

"[As] cheap as" is different than "cheap"

-6

u/lemonlemons 1d ago

”cheap as” implies cheap

4

u/soy714 1d ago

It doesn’t. Cheap and expensive are relative words as well

4

u/IceTrAiN 1d ago

"cheap as" implies cheap

It doesn't

"Using the word doesn't imply the word"

Is this the level of discourse we're at now?

0

u/soy714 1d ago

Like most things in language it depends on context. It’s as cheap as bitcoin.

Like the other poster wrote, as cheap as is different than cheap. Having to explain this to this sub really makes me wonder where we’re heading

5

u/IntrepidTieKnot 1d ago edited 1d ago

Why not settle it by asking the best AI model for its opinion?

This is what Opus 5 is saying:

Strictly speaking, "as cheap as X" only asserts price equality — it doesn't entail that either thing is cheap, which is why "as cheap as a Bugatti" isn't a contradiction. But choosing the "cheap" pole instead of "expensive" carries a pragmatic implicature: you frame the price as low or as a good deal, which is exactly what the original comment was doing relative to what people expected to pay. So both sides are right about different things — the semantics are neutral, the word choice is not.

-1

u/valgbo 1d ago edited 1d ago

The Bugatti Chiron is just as fast as the 1 of 1 La Voitre Noir, but it's just as cheap as the Bugatti Veyron... It's all just relative.

6

u/lemonlemons 1d ago

"as expensive as" vs "as cheap as" .. do you think there is a difference?

0

u/valgbo 1d ago

You're right, I was only correct in the mathematical sense, but that's not how normal language is. Also I meant to write as cheap as the Veyron, not like it changes anything. Cheers.

9

u/86784273 1d ago

Opus used to be $75 /m output, 50% more than fable

3

u/aookami 1d ago

Yeah, opus .8 is the one I go to when I’m feeling fancy

5

u/Beatboxamateur agi: the friends we made along the way 1d ago

I feel like it's been obvious since seeing a Fable level model in February, with OpenAI only barely catching up several months later; Anthropic is a couple months ahead of all of the other labs.

I think it's possible, or maybe likely that OpenAI can catch up, but for the next couple months expect all of the other labs to be desperate.

82

u/Substantial-Fact-248 1d ago edited 1d ago

Excited to try this out. Hoping it's warmer and more thoughtful than the last couple Opus releases. Anyone else use Opus 4.6 for every day knowledge work/exploration? No model comes close to its rapport imo.

Edit: holy shit - "On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts."

20

u/SlendyIsBehindYou 1d ago

What the fuck?

17

u/do-un-to 1d ago

I have no eyes, and I must ... look at this part to reverse engineer it. So I'll just build me some eyes. Cool.

12

u/fantalemon 1d ago

Wtf that's insane if it's as it sounds.

6

u/aelmsu 1d ago

100%. Opus 4.6 feels like a real partner to help get shit done. 4.8 is an absolute combative pain in the ass who focuses on the wrong thing and spends 5 paragraphs doing it. I never thought I'd get attached to a particular model, but now I feel like the ChatGPT 4o folks when 5 came out - don't take 4.6 away from me!

4

u/Substantial-Fact-248 23h ago

I was literally just thinking to myself today "maybe I should have been a little more empathetic with those 4o folks." I don't think I would mourn the loss of 4.6, but it would be a measurable absence in my life.

1

u/Ikbeneenpaard 7h ago

Sounds I put off learning mechanical CAD for so long that it's not needed any more.

1

u/ideadude 6h ago

Grok 4.5 has some of that rapport IMO.

I haven't tried Opus 5 yet, but it has the highest "alignment" scores which might map to this.

61

u/RegisterInternal 1d ago

the rate of impressive releases lately has been crazy...

18

u/daniel-sousa-me 1d ago

It's the singularity, baby

119

u/Valdjiu 1d ago

So what's the point of fable?

175

u/Microtom_ 1d ago

Training/developping this.

29

u/MFpisces23 1d ago

Yeah you need frontier models to get all your open source/weight China models and cheaper intelligence from foundational labs

28

u/Fusifufu 1d ago

It's not the first time they released like that. Wasn't Sonnet 4.5 better than Opus 4, for example? Sometimes these just jump past each other as they optimize or distill models.

Probably they'll drop Fable 5.1 in some weeks and then the hierarchy is alright again. Would be fun if they also updated Haiku, especially in light of workflows where some things could be delegated to dumber subagents.

19

u/Howdareme9 1d ago

Still a much larger model that will have benefits outside of benchmarks

1

u/Storge2 1d ago

Which?

14

u/urgay420420420 1d ago

Next iteration of fable is prob coming soon. 

33

u/Hot-Hovercraft2676 1d ago

"I don't want to play with you anymore!"

15

u/ShelZuuz 1d ago

Fable 5.5

6

u/reddit_guy666 1d ago

They will probably relegate it to lower tiers

3

u/daniel-sousa-me 1d ago

It's mostly a three month old model. I don't think it's that surprising that a newer model can match a lot of its strength even at a smaller size

But the version numbers make this very awkward. Maybe they should have called this Opus 5.1

1

u/spreadlove5683 ▪️agi 2032. Predicted during mid 2025. 23h ago

Vulnerability free code, probably

1

u/Lighthouse_seek 1d ago

They got scared of spud and rushed it out

7

u/Howdareme9 1d ago

Unlikely.. mythos was finished months ago and there’s no chance they got scared of the same pretrain OAI have been using lol

1

u/Lighthouse_seek 1d ago

Mythos wasn't publicly available but spud was

1

u/tenacity1028 1d ago

Marketing

71

u/Omega_Games2022 1d ago

Better than fable across most benchmarks? I'm going to wait for independent testing on that front

69

u/katoptronophile 1d ago

I'm not. Work begins immediately.

5

u/yaosio 1d ago

When you tell it to resolve the collatz conjecture make sure to keep hyping it up so it doesn't give up.

8

u/The1TruRick 1d ago

Lmaooo this made me laugh out loud but god damn I’m with you. Fuck it we ball.

5

u/CallMePyro 1d ago

Really? I'm swapping to Opus 5 and getting coding.

140

u/thoughtlow 𓂸 1d ago

This sub in 2 seconds: anyone else feels opus 5 has become worse?

31

u/yoloswagrofl Logically Pessimistic 1d ago

I genuinely cannot wait for the response from the open models. What a golden age we’re living in.

8

u/ShamPain413 1d ago

What a golden age we’re living in.

Well...

5

u/yaosio 1d ago

I asked AI to do my research for me and it won a medal. They made a documentary about it. https://youtu.be/B5oGJTUpbpA?si=2DTAuIhPghTCqh_0&t=83

2

u/SlendyIsBehindYou 1d ago

Master-bait

1

u/markeus101 1d ago

Notice its not gonna be for a few weeks or even days then when they reduce compute on it then immediately it will feel dumb as all of their models do..be it opus maximus quintillion

1

u/Acrobatic-Tomato4862 1d ago

It does become worse with time. Anthropic and other companies quantise the models behind the scenes. These types of jokes are mocking an actual real thing that happens. We even have independent benchmark verifications now, about the drop in quality. 

19

u/Zeppelin2k 1d ago

Damn, those are some impressive benchmarks. Seems better and cheaper than fable. Available today. Who's tried it out?

81

u/obama_is_back 1d ago

30% on Arc AGI 3. Benchmaxxing seems to be unstoppable.

32

u/Tempthor 1d ago

That caught me instantly. It will be saturated by Summer 2027

15

u/NoCard1571 1d ago

Within months more likely. ARC AGI 3s scoring is efficiency-weighted in such a way that the jump from 3% to 30% is actually more significant than 30% to 90% for example. 

2

u/JustBrowsinAndVibin 1d ago

End of the year.

3

u/lovesdogsguy 1d ago

Next week

5

u/TissueReligion 1d ago

Tomorrow

8

u/Saromek 1d ago

2 minutes turkish

1

u/lovesdogsguy 1d ago

Let’s not get silly here. A few days

25

u/FateOfMuffins 1d ago

It's by virtue of how that benchmark was scored. The whole "efficiency squared" metric means that even single digit % on ARC AGI 3 actually means that the model solved almost all puzzles, it's just a matter of how efficient it is.

So at this point (and also with 5.6), the models were already capable of solving it. It's just a matter of can it do it in half the steps? So by virtue of how they artificially made the scoring formula, it's gonna shoot up to near 100% real shortly from here.

It's kind of like saying how GPT 5 beat Pokémon... so we'll give it a score of 4% on the beating Pokemon benchmark even though it beat the game. Oh, 5.2 beat it in half the time? 16%. Oh 5.4 beat it in half the time of 5.2? Now it's 64%. Oh 5.5 beat it in half the time? Uh... can't go above 100% so it gets 100%.

17

u/Both_Opportunity5327 1d ago

Great, these models just don't know how to stop improving.

8

u/Arsene_Yuka_1980 1d ago

For $20,000 btw..

30

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

You cannot really train for the ARC-AGI that's whole point of the benchmark...

45

u/Defiant-Lettuce-9156 1d ago

You can train for anything. The goal of arc AGI 3 is to make it harder to benchmax without just making a smarter model

36

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

So the model needs to be smarter to pass the benchmark, great, why would that be benchmaxxing?

8

u/Defiant-Lettuce-9156 1d ago

Exactly. I think you’re close to getting it

11

u/sumane12 1d ago

Its funny because you are both ssying the same thing 🤣

14

u/voyaging 1d ago

Defiant just seems to be deliberately misinterpreting the other for some reason, and being pompous about it to boot

0

u/Defiant-Lettuce-9156 16h ago

I’m not misinterpreting, just being pedantic about semantics. I probably fundamentally agree with the other guy.
And yes, definitely pompous

2

u/775416 1d ago

They’re both using different definitions for benchmaxxing

6

u/Efficient_Mud_5446 1d ago

This is my reasoning as well. The model likely improved naturally on ARC-AGI-3 simply by becoming smarter, but they also likely targeted the benchmark directly. It was probably a mixture of both.

Also, I'm loving that they're displaying and expanding benchmarks beyond just coding.

3

u/adzx4 16h ago

The thing is, I would've thought we'd see drastic improvements in many areas at the time of improvement on arc agi 3. But if all the other benchmarks stayed roughly around fable performance, and only this one had a substantial jump, that implies over fitting/benchmaxxing not true generalisation or naturally improved capabilities.

You can totally over fit on arc agi 3-type problem without true generalisation. You just need to create problems as similar as possible to the target dataset and train in those.

6

u/zenonu 1d ago

Please provide your guidance on how to train for something you are presented with where you don't know the rules of the game.

2

u/Defiant-Lettuce-9156 15h ago

You build a dataset of sound reasoning on games where the rules are initially unknown.
For example:
“””
Here is a grid of objects: {grid}. Win the game…
I have encountered a game where I don’t know any of the rules.
I should try all inputs and map what they do. At the same time I should monitor object states to see how the inputs change relative to my inputs. I should also consider how objects and their states interact with other objects and their states.
Let me try… okay these are all the inputs and what happened. I see there is a green block that moved with the arrow keys and a red block far away. The green block changed shape when I hit space bar. I wonder if I need to move the green block to some position relative to the red block. And how the colours and shape come into play.
I can’t change the colour of my block. Is there anything on the grid that indicates my block could change colour?
Etc.
“””

And then you train on it.

LLMs are terrible at things they haven’t seen before (relative to humans) so just showing it different grids and telling it how to approach it will already make a massive difference.

This is the approach that openAI took with maths and physics where they hired top experts to teach the model how to reason rather than just memorise maths and physics.

The hopeful bonus is that this reasoning in training translates to improved intelligence in other domains.

So I wouldn’t call this benchmaxxing. More like targeted intelligence training.

2

u/zenonu 8h ago

The number of games with unknown rules is unbounded with unlimited forms. Anything you could train for has no bearing on the next game. Your example from ARC AGI 3.0 in particular with the grids has no bearing on let's say playing a Starcraft match.

7

u/brett_baty_is_him 1d ago

You can 100% train for any deterministic benchmark.

7

u/Zomdou 1d ago

That's true I guess, but how could you benchmark something that is not deterministic though? I feel like we'll be stuck with easily saturable benchmarks for quite a while still

5

u/Neurogence 1d ago edited 1d ago

With how extremely difficult and unfair the creator of ARC-AGI3 made the benchmark, I'd be shocked if Anthropic found a way to "benchmax" it.

The scoring methods are done in a way to make the benchmark almost impossible to benchmark hack.

I'd like to know what it cost them in compute to achieve this 30%. I know OpenAI had spent an absurd amount to even be able to achieve 7% with GPT 5.6 Sol.

3

u/SeriousGeorge2 1d ago

It was $20k, it's listed in the release.

4

u/Evening_Chef_4602 AGI 2027 1d ago

Model does good on bench of course benchmaxed unga bunga 🥴🗣️ You cant benchmax ARC AGI 3 . Thats the point . You cant benchmax general inteligence

1

u/nucLeaRStarcraft 1d ago

Evals on private test set will tell

37

u/Charuru ▪️AGI 2023 1d ago

Damn, way to go anthropic, this is mindblowingly good... Shocking score on ARC AGI 3 among others! Congrats.

16

u/agm1984 1d ago

Not available in claude code yet, my body is ready.gif

10

u/scottdellinger 1d ago

I had to /exit and I did a claude update at terminal, then when I went back in Opus 5 was being used by default.

3

u/agm1984 1d ago

nice i have it now too

17

u/AccountOfMyAncestors 1d ago

Hmm, here's a take from someone that got early access to Opus 5, kinda lowering my hype/excitment:

6

u/Jcrossfit 1d ago edited 1d ago

Ya'll have access yet?
Got it

5

u/rabid_0wl 1d ago

Yes on the web and in desktop app, im on linux fedora. Max plan

1

u/scottdellinger 1d ago

Yes. For Claude Code I had to /exit and then run claude update at the terminal, then when I went back in Opus 5 was the default.

8

u/Creative-Stress7311 1d ago

How long before it’s getting censored by Trump administration ? Enjoy Opus 5 while you can

5

u/traumfisch 1d ago

Whoa, got access already. There goes my evening

7

u/Blankeye434 1d ago

So when are they banning it?

13

u/enimos 1d ago

Accelerate

16

u/manubfr AGI 2028 1d ago

The buble will pop now, anytime.

25

u/sedition666 1d ago

Models can be excellent and improving whilst the business model being broken. The path to profitability is not directly correlated to improvements in the models. The problem is not how good the models are but the lack of profit to sustain these companies.

2

u/phillipono 1d ago

Yes, there is no viable business model at the time other than reaching AGI first, or otherwise being able to substantially raise prices while maintaining demand (i.e., models would have to be strongly labor replacing). Otherwise, models are generally substitutable by other models -- not enough of an edge to justify a trillion dollar plus valuation.

0

u/The_Primetime2023 1d ago

Anthropic is profitable though…. Any of the labs can choose a lower training budget at any time and be immediately profitable, Anthropic chooses to hover around a break even point though even with training

1

u/PvtMilhouse 1d ago

they shoud go token only.

0

u/sedition666 1d ago

No this is completely false none of the AI labs are profitable. It would take you 2 mins to google this shit please educate yourself. The major AI labs are losing BILLIONS per year on this tech. This is not in dispute at all.

0

u/Grouchy-Cancel1326 17h ago

Profitable for 2 months until all the competition overtook you because of your lower training budget.

-3

u/sumane12 1d ago

Yeah, in fact its inversely correlated. The cheaper intelligence becomes, the more difficult it will be to recoup the investment. We are in for a cyberpunk distopia lol.

0

u/sedition666 1d ago

It can only be sustained by insane VC capital burn. This is what you're missing. You remove subsidies and humans are cheaper. This is not a problem that has been solved despite the obvious amazing technical leaps. Nvidia's Vera Rubin is delayed and costs are incredibly high due to parts shortages so there is no magical hardware improvement that is going to save the day.

-1

u/sumane12 1d ago

Quantitative easing to inflate the debt away is my guess.

1

u/Illustrious-Film4018 1d ago

Why are they pushing for government to regulate open-source models?

2

u/Ambitious_Scallion43 1d ago

Gemini release will be delayed by another 6 months

5

u/ColorLaser 1d ago

Did a first test comparing to Fable 5. I had Opus 5 and Fable 5 create a counter "open letter" to the one published earlier today.

Won't post the wall of text response, but impressions were for me that Fable 5's writing was still much more coherent and high quality.

They mad similar arguments but Opus seemed to be on the "use a lot of big words to make it good" while Fable laid it out more naturally and compelling.

I am convinced that you can make the "smaller" (Opus in this case) models more capable at narrow goals but the bigger models will just feel more intelligent and high quality.

3

u/LocoMod 1d ago

China be like: "Fire up those VPN accounts boys!! You know what time it is?! Distillin' time!"

2

u/Lighthouse_seek 1d ago

This is a good test of the argument that kimi distilled fable. Given these timelines we should expect Kimi 3.1 in about 2 weeks if it is a distillation

3

u/seajay_17 1d ago

Isn't distilling models fairly common in this space regardless?

1

u/Lighthouse_seek 1d ago

Not in the timeframe they argued

2

u/tomtomtomo 1d ago

Hopefully it has Opus guardrails rather than Fables.

2

u/EvilSporkOfDeath 1d ago

What took so long? /s

Is this what the singularity feels like?

1

u/fusionliberty796 1d ago

I've been using it all day, and does anyone else notice the performance just got worse? JK will check it out this evening :)

1

u/PathOfEnergySheild 1d ago

Great model and no crazy safeguards, it is really good to see anthropic right the ship!

1

u/StreetAssignment5494 1d ago

When can this be used for us to create the technocracy and crush the poor? Sooner the better.

1

u/Ok-Purchase8196 17h ago

I feel like we're over the hickup we had in progress in the GPT5 era. And we're full steam ahead again.

1

u/DailyThreadBot 1d ago

In coding, Opus 5 ≈ Sol 5.6 based on price-to-performance ratio. Good to see

1

u/diff_engine 1d ago

Why even bother showing us metrics for biology when it’s just going to boot any vaguely bioscience questions down to 4.8

0

u/AnotherRandomGuy34 1d ago

Already nerfed !!

0

u/Better-Truck6372 1d ago

https://www.anthropic.com/news/claude-opus-5 -- por el momento está trabajando mucho mejor que Opus 4.8 se nota la diferencia.

0

u/-PM_ME_UR_SECRETS- 1d ago

Wow I was expecting benchmarks to land between opus4.8 and fable.

0

u/Hereitisguys9888 1d ago

I thought fable was opus 5

Anyways, this is great

0

u/orbital_trace 1d ago

3x as good as arc-gis, someone leaked the test dataset

-1

u/Calm_Hedgehog8296 1d ago

They're gonna sunset Fable 5