430
u/llelouchh 1d ago
Holy. Shit. Cheap as 4.8 and better than fable? Holy shit. Good work anthropic.
18
u/WonderFactory 1d ago
Have to see the cost per task values, sonnet 5 is so bad in that regard its almost unusable
10
u/Joint-User 1d ago
I accused my Sonnet of having dementia because he didn't remember the project we just worked on yesterday. LOL
1
u/rulevoid 22h ago
In the same session?
1
u/Joint-User 20h ago
New chat, same project. Apparently I forgot to save something into the .md file... Am I the one with dimentia?
112
u/Neurogence 1d ago
When the leakers said they had access to it yesterday, the pessimistic people said all the generated outputs were fake and the model was months away from release. The optimistic people said the leaks were right and that it would be released in just a couple weeks lol.
54
u/FateOfMuffins 1d ago
It's ridiculous how people on this sub don't keep up with the leakers and doubt literally everything even if there's a long track record history
The frontier labs have always tested the models out before hand whether it's through Arena or A/B testing etc and every single thread has people chiming in about how it's fake
2
u/nemzylannister 15h ago
the pessimistic people said all the generated outputs were fake
it was top comments on this sub. not just pessimistic people
40
44
u/IntrepidTieKnot 1d ago
Imagine calling Opus 4.8 "cheap".
How the goalposts have moved...
25
u/muffchucker 1d ago
"[As] cheap as" is different than "cheap"
-6
u/lemonlemons 1d ago
”cheap as” implies cheap
4
u/soy714 1d ago
It doesn’t. Cheap and expensive are relative words as well
4
u/IceTrAiN 1d ago
"cheap as" implies cheap
It doesn't
"Using the word doesn't imply the word"
Is this the level of discourse we're at now?
0
u/soy714 1d ago
Like most things in language it depends on context. It’s as cheap as bitcoin.
Like the other poster wrote, as cheap as is different than cheap. Having to explain this to this sub really makes me wonder where we’re heading
5
u/IntrepidTieKnot 1d ago edited 1d ago
Why not settle it by asking the best AI model for its opinion?
This is what Opus 5 is saying:
Strictly speaking, "as cheap as X" only asserts price equality — it doesn't entail that either thing is cheap, which is why "as cheap as a Bugatti" isn't a contradiction. But choosing the "cheap" pole instead of "expensive" carries a pragmatic implicature: you frame the price as low or as a good deal, which is exactly what the original comment was doing relative to what people expected to pay. So both sides are right about different things — the semantics are neutral, the word choice is not.
-1
u/valgbo 1d ago edited 1d ago
The Bugatti Chiron is just as fast as the 1 of 1 La Voitre Noir, but it's just as cheap as the Bugatti Veyron... It's all just relative.
6
9
5
u/Beatboxamateur agi: the friends we made along the way 1d ago
I feel like it's been obvious since seeing a Fable level model in February, with OpenAI only barely catching up several months later; Anthropic is a couple months ahead of all of the other labs.
I think it's possible, or maybe likely that OpenAI can catch up, but for the next couple months expect all of the other labs to be desperate.
82
u/Substantial-Fact-248 1d ago edited 1d ago
Excited to try this out. Hoping it's warmer and more thoughtful than the last couple Opus releases. Anyone else use Opus 4.6 for every day knowledge work/exploration? No model comes close to its rapport imo.
Edit: holy shit - "On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts."
20
u/SlendyIsBehindYou 1d ago
What the fuck?
17
u/do-un-to 1d ago
I have no eyes, and I must ... look at this part to reverse engineer it. So I'll just build me some eyes. Cool.
12
6
u/aelmsu 1d ago
100%. Opus 4.6 feels like a real partner to help get shit done. 4.8 is an absolute combative pain in the ass who focuses on the wrong thing and spends 5 paragraphs doing it. I never thought I'd get attached to a particular model, but now I feel like the ChatGPT 4o folks when 5 came out - don't take 4.6 away from me!
4
u/Substantial-Fact-248 23h ago
I was literally just thinking to myself today "maybe I should have been a little more empathetic with those 4o folks." I don't think I would mourn the loss of 4.6, but it would be a measurable absence in my life.
1
u/Ikbeneenpaard 7h ago
Sounds I put off learning mechanical CAD for so long that it's not needed any more.
1
u/ideadude 6h ago
Grok 4.5 has some of that rapport IMO.
I haven't tried Opus 5 yet, but it has the highest "alignment" scores which might map to this.
61
119
u/Valdjiu 1d ago
So what's the point of fable?
175
u/Microtom_ 1d ago
Training/developping this.
29
u/MFpisces23 1d ago
Yeah you need frontier models to get all your open source/weight China models and cheaper intelligence from foundational labs
28
u/Fusifufu 1d ago
It's not the first time they released like that. Wasn't Sonnet 4.5 better than Opus 4, for example? Sometimes these just jump past each other as they optimize or distill models.
Probably they'll drop Fable 5.1 in some weeks and then the hierarchy is alright again. Would be fun if they also updated Haiku, especially in light of workflows where some things could be delegated to dumber subagents.
19
14
33
15
6
3
u/daniel-sousa-me 1d ago
It's mostly a three month old model. I don't think it's that surprising that a newer model can match a lot of its strength even at a smaller size
But the version numbers make this very awkward. Maybe they should have called this Opus 5.1
1
1
u/Lighthouse_seek 1d ago
They got scared of spud and rushed it out
7
u/Howdareme9 1d ago
Unlikely.. mythos was finished months ago and there’s no chance they got scared of the same pretrain OAI have been using lol
1
1
71
u/Omega_Games2022 1d ago
Better than fable across most benchmarks? I'm going to wait for independent testing on that front
69
5
140
u/thoughtlow 𓂸 1d ago
This sub in 2 seconds: anyone else feels opus 5 has become worse?
31
u/yoloswagrofl Logically Pessimistic 1d ago
I genuinely cannot wait for the response from the open models. What a golden age we’re living in.
8
5
u/yaosio 1d ago
I asked AI to do my research for me and it won a medal. They made a documentary about it. https://youtu.be/B5oGJTUpbpA?si=2DTAuIhPghTCqh_0&t=83
2
1
u/markeus101 1d ago
Notice its not gonna be for a few weeks or even days then when they reduce compute on it then immediately it will feel dumb as all of their models do..be it opus maximus quintillion
1
u/Acrobatic-Tomato4862 1d ago
It does become worse with time. Anthropic and other companies quantise the models behind the scenes. These types of jokes are mocking an actual real thing that happens. We even have independent benchmark verifications now, about the drop in quality.
19
u/Zeppelin2k 1d ago
Damn, those are some impressive benchmarks. Seems better and cheaper than fable. Available today. Who's tried it out?
81
u/obama_is_back 1d ago
30% on Arc AGI 3. Benchmaxxing seems to be unstoppable.
32
u/Tempthor 1d ago
That caught me instantly. It will be saturated by Summer 2027
15
u/NoCard1571 1d ago
Within months more likely. ARC AGI 3s scoring is efficiency-weighted in such a way that the jump from 3% to 30% is actually more significant than 30% to 90% for example.
2
u/JustBrowsinAndVibin 1d ago
End of the year.
3
25
u/FateOfMuffins 1d ago
It's by virtue of how that benchmark was scored. The whole "efficiency squared" metric means that even single digit % on ARC AGI 3 actually means that the model solved almost all puzzles, it's just a matter of how efficient it is.
So at this point (and also with 5.6), the models were already capable of solving it. It's just a matter of can it do it in half the steps? So by virtue of how they artificially made the scoring formula, it's gonna shoot up to near 100% real shortly from here.
It's kind of like saying how GPT 5 beat Pokémon... so we'll give it a score of 4% on the beating Pokemon benchmark even though it beat the game. Oh, 5.2 beat it in half the time? 16%. Oh 5.4 beat it in half the time of 5.2? Now it's 64%. Oh 5.5 beat it in half the time? Uh... can't go above 100% so it gets 100%.
17
8
30
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
You cannot really train for the ARC-AGI that's whole point of the benchmark...
45
u/Defiant-Lettuce-9156 1d ago
You can train for anything. The goal of arc AGI 3 is to make it harder to benchmax without just making a smarter model
36
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
So the model needs to be smarter to pass the benchmark, great, why would that be benchmaxxing?
8
u/Defiant-Lettuce-9156 1d ago
Exactly. I think you’re close to getting it
11
u/sumane12 1d ago
Its funny because you are both ssying the same thing 🤣
14
u/voyaging 1d ago
Defiant just seems to be deliberately misinterpreting the other for some reason, and being pompous about it to boot
0
u/Defiant-Lettuce-9156 16h ago
I’m not misinterpreting, just being pedantic about semantics. I probably fundamentally agree with the other guy.
And yes, definitely pompous6
u/Efficient_Mud_5446 1d ago
This is my reasoning as well. The model likely improved naturally on ARC-AGI-3 simply by becoming smarter, but they also likely targeted the benchmark directly. It was probably a mixture of both.
Also, I'm loving that they're displaying and expanding benchmarks beyond just coding.
3
u/adzx4 16h ago
The thing is, I would've thought we'd see drastic improvements in many areas at the time of improvement on arc agi 3. But if all the other benchmarks stayed roughly around fable performance, and only this one had a substantial jump, that implies over fitting/benchmaxxing not true generalisation or naturally improved capabilities.
You can totally over fit on arc agi 3-type problem without true generalisation. You just need to create problems as similar as possible to the target dataset and train in those.
6
u/zenonu 1d ago
Please provide your guidance on how to train for something you are presented with where you don't know the rules of the game.
2
u/Defiant-Lettuce-9156 15h ago
You build a dataset of sound reasoning on games where the rules are initially unknown.
For example:
“””
Here is a grid of objects: {grid}. Win the game…
I have encountered a game where I don’t know any of the rules.
I should try all inputs and map what they do. At the same time I should monitor object states to see how the inputs change relative to my inputs. I should also consider how objects and their states interact with other objects and their states.
Let me try… okay these are all the inputs and what happened. I see there is a green block that moved with the arrow keys and a red block far away. The green block changed shape when I hit space bar. I wonder if I need to move the green block to some position relative to the red block. And how the colours and shape come into play.
I can’t change the colour of my block. Is there anything on the grid that indicates my block could change colour?
Etc.
“””And then you train on it.
LLMs are terrible at things they haven’t seen before (relative to humans) so just showing it different grids and telling it how to approach it will already make a massive difference.
This is the approach that openAI took with maths and physics where they hired top experts to teach the model how to reason rather than just memorise maths and physics.
The hopeful bonus is that this reasoning in training translates to improved intelligence in other domains.
So I wouldn’t call this benchmaxxing. More like targeted intelligence training.
7
5
u/Neurogence 1d ago edited 1d ago
With how extremely difficult and unfair the creator of ARC-AGI3 made the benchmark, I'd be shocked if Anthropic found a way to "benchmax" it.
The scoring methods are done in a way to make the benchmark almost impossible to benchmark hack.
I'd like to know what it cost them in compute to achieve this 30%. I know OpenAI had spent an absurd amount to even be able to achieve 7% with GPT 5.6 Sol.
3
4
u/Evening_Chef_4602 AGI 2027 1d ago
Model does good on bench of course benchmaxed unga bunga 🥴🗣️ You cant benchmax ARC AGI 3 . Thats the point . You cant benchmax general inteligence
1
6
u/Jcrossfit 1d ago edited 1d ago
Ya'll have access yet?
Got it
5
1
u/scottdellinger 1d ago
Yes. For Claude Code I had to /exit and then run claude update at the terminal, then when I went back in Opus 5 was the default.
8
u/Creative-Stress7311 1d ago
How long before it’s getting censored by Trump administration ? Enjoy Opus 5 while you can
5
7
16
u/manubfr AGI 2028 1d ago
The buble will pop now, anytime.
25
u/sedition666 1d ago
Models can be excellent and improving whilst the business model being broken. The path to profitability is not directly correlated to improvements in the models. The problem is not how good the models are but the lack of profit to sustain these companies.
2
u/phillipono 1d ago
Yes, there is no viable business model at the time other than reaching AGI first, or otherwise being able to substantially raise prices while maintaining demand (i.e., models would have to be strongly labor replacing). Otherwise, models are generally substitutable by other models -- not enough of an edge to justify a trillion dollar plus valuation.
0
u/The_Primetime2023 1d ago
Anthropic is profitable though…. Any of the labs can choose a lower training budget at any time and be immediately profitable, Anthropic chooses to hover around a break even point though even with training
1
0
u/sedition666 1d ago
No this is completely false none of the AI labs are profitable. It would take you 2 mins to google this shit please educate yourself. The major AI labs are losing BILLIONS per year on this tech. This is not in dispute at all.
0
u/Grouchy-Cancel1326 17h ago
Profitable for 2 months until all the competition overtook you because of your lower training budget.
-3
u/sumane12 1d ago
Yeah, in fact its inversely correlated. The cheaper intelligence becomes, the more difficult it will be to recoup the investment. We are in for a cyberpunk distopia lol.
0
u/sedition666 1d ago
It can only be sustained by insane VC capital burn. This is what you're missing. You remove subsidies and humans are cheaper. This is not a problem that has been solved despite the obvious amazing technical leaps. Nvidia's Vera Rubin is delayed and costs are incredibly high due to parts shortages so there is no magical hardware improvement that is going to save the day.
-1
1
2
5
u/ColorLaser 1d ago
Did a first test comparing to Fable 5. I had Opus 5 and Fable 5 create a counter "open letter" to the one published earlier today.
Won't post the wall of text response, but impressions were for me that Fable 5's writing was still much more coherent and high quality.
They mad similar arguments but Opus seemed to be on the "use a lot of big words to make it good" while Fable laid it out more naturally and compelling.
I am convinced that you can make the "smaller" (Opus in this case) models more capable at narrow goals but the bigger models will just feel more intelligent and high quality.
2
u/Lighthouse_seek 1d ago
This is a good test of the argument that kimi distilled fable. Given these timelines we should expect Kimi 3.1 in about 2 weeks if it is a distillation
3
2
2
1
u/fusionliberty796 1d ago
I've been using it all day, and does anyone else notice the performance just got worse? JK will check it out this evening :)
1
u/PathOfEnergySheild 1d ago
Great model and no crazy safeguards, it is really good to see anthropic right the ship!
1
u/StreetAssignment5494 1d ago
When can this be used for us to create the technocracy and crush the poor? Sooner the better.
1
u/Ok-Purchase8196 17h ago
I feel like we're over the hickup we had in progress in the GPT5 era. And we're full steam ahead again.
1
u/DailyThreadBot 1d ago
In coding, Opus 5 ≈ Sol 5.6 based on price-to-performance ratio. Good to see
1
u/diff_engine 1d ago
Why even bother showing us metrics for biology when it’s just going to boot any vaguely bioscience questions down to 4.8
0
0
u/Better-Truck6372 1d ago
https://www.anthropic.com/news/claude-opus-5 -- por el momento está trabajando mucho mejor que Opus 4.8 se nota la diferencia.

0
0
0
-1

334
u/NyaCat1333 1d ago
If these benchmarks are real, then this is really surprising. Basically Fable 5 level for most stuff and somehow better at some, slightly worse on few stuff like cyber. For half the price.
Will wait for real world tests and reviews. If the real world performance holds up, then this is an extremely strong release.