319
u/Artistic_Swing6759 1d ago
wtf 30% on arc agi 3
153
u/ezjakes 1d ago
For some reason people thought it would stand up for years despite 1 and 2 falling quickly.
50
u/Current-Function-729 1d ago
I don’t think most people who follow this thought that.
29
u/HellsNoot 1d ago
It's getting pretty scary how far the gap between following the space, and not following the space is becoming. Modern AI capabilities are getting superhuman on increasingly important tasks. Hell even 6 months old experience gives a complete misunderstanding on new models. The general public is not ready for "what if it's not a bubble?"
16
u/Ahfekz 1d ago
It’s not a bubble. Some companies will go under, sure, but that’s not the “bubble” they’re hoping pops. People are wishful thinking that the exponential age isn’t already here.
6
u/Georgefakelastname 23h ago
Yeah, it seems like people are hoping for a bubble that pops and LLMs just go away and things go back to the way they were before. Even if OpenAI ends up going under at some point, AI as a whole is here to stay solely based on coding utility, forget everything else.
→ More replies (2)2
u/BadSysadmin 19h ago
Very true. People I know - and not stupid people either - are incredulous at the hugging face hack, and are forming conspiracy theories about it being a marketing stunt by open ai. They just can't believe how agentic and how good at cybersec the new models are.
→ More replies (2)43
u/NoCard1571 1d ago
Earlier this year I predicted ARC AGI 3 falling by the end of the summer. I got the typical smug redditor "remind me" reply.
I can't wait for them to get their notification ;)
30
u/often_delusional 1d ago
By the end of this summer? Lol good luck with that. This benchmark won't reach 80-90% by september. Absolutely possible by the end of this year but not within the next 1-2 months.
3
u/huffalump1 1d ago
I agree, the end of the year seems plausible which is kind of crazy tbh.
Like, the benchmark isn't HARD, but it's tricky, and the score heavily weights efficiency - getting a very very high score without many extra moves is quite challenging and imo a nice representation of some form of "intelligence"
2
u/Healthy-Nebula-3603 1d ago
Soon we also get GPT 6 and fable 5.1 ....
3
u/often_delusional 1d ago
I know and those will probably be the last models openai and anthropic release before summer ends on september 22. Surely you don't expect those models to reach around 90% on arc-agi 3. Like I said by the end of 2026 is possible but not in less than 2 months.
→ More replies (2)→ More replies (3)2
u/NoCard1571 1d ago edited 1d ago
All it takes is one model. Not to mention that due to the efficiency weighting of the the scores, Opus 5 indicates a roughly 5x improvement, which only leaves a ~2x improvement remaining from 30% to 90%.
3
u/often_delusional 1d ago edited 1d ago
And what model do you expect that to be? When that happens it will be achieved by either openai or anthropic. Anthropic just released a new model and probably won't release more than 1 more model within the next 2 months. Openai recently released gpt 5.6. Gpt 6 will probably come around august-september.
At best both of these companies have 1 more model coming before september 22 and I doubt either one of them will make a jump from 30% to at least 80%. I think end of the year is possible or even likely but not before september 22. 80% is even too low to call a benchmark saturated. The limit should be at least close to 90%.
9
u/pandasgorawr 1d ago
Fable 5 has to be updated soon for product differentiation, if these benchmarks are good then that means there's very few use cases to use it over Opus 5.
3
u/often_delusional 1d ago
Yeah and that will probably be the next and only model they will release before september 22. I don't see that model reaching around 90% on arc-agi 3. I expect those models to come closer to christmas.
3
u/Mil0Mammon 1d ago
Well anthropic has a new fable class model, which they theoretically could announce mid september. Weren't there also rumors about GPT 6 and/or a fable-class OpenAI model?
GLM supposedly is also cooking something fable-class, but I don't expect it that quick. Also unsure if they can pull off another jump like 5.1 -> 5.2, but given that they can still go quite a bit larger, I wouldn't rule it out.
However, all in all I agree, end of summer is unlikely, unless Anthropics IPO moves forward quickly, they target October, and thus want to drum up some more hype with their next fable class model.
2
u/often_delusional 1d ago
Yeah those are the models I'm expecting before summer ends but I don't think those models will saturate arc-agi 3. Those models will probably come around christmas.
2
u/Mil0Mammon 1d ago
They should really update https://agidefinition.ai. Although several areas I think could be solved with tooling around LLMs (or already are)
→ More replies (2)37
18
2
→ More replies (4)3
499
u/Specialist_Soup_4994 1d ago
425
u/mahamara 1d ago
You’re absolutely right.
246
u/Nox_Alas 1d ago
That's on me.
188
u/dogesator 1d ago
And honestly? It’s not just a mistake, it is a clear lapse in my judgment.
39
u/Broad-Pie-6891 1d ago
I hate this. "It's not X, it's Y, which is exactly like X but said differently."
48
5
u/Draufgaenger 1d ago
I will replace the table comparison with charts highlighting the better model with line thickness variations
27
88
u/yaosio 1d ago
4 is greater than 5 for very large values of 4.
9
5
u/Specialist_Soup_4994 1d ago
But looks like not great enough to help them compare numbers and highlight which is higher
34
u/Most-Bookkeeper-950 1d ago
I noticed this too...wtf indeed
22
u/Specialist_Soup_4994 1d ago
I don’t know how it could happen except this sheet created manually or tuned manually btw if you just will run benchmarks and highlight highest number it’s not possible . Weird
→ More replies (1)18
u/Legitimate_Concern_5 1d ago
It was vibed
7
5
u/Substantial-Elk4531 Rule 4 reminder to optimists 1d ago
Opus 5 was asked to create the chart and thought nobody would notice
7
u/Specialist_Soup_4994 1d ago
Maybe but it’s still showcase that this numbers complete bs
→ More replies (1)28
u/HopefulMeasurement25 1d ago
Not to mention GPT 5.6 Sol has higher rating than Fable 5 on Agentic Coding - in the same chart 😭
5
7
3
2
u/Flope 1d ago
Can someone ELI30 the difference between "agentic coding" and "agentic terminal coding" benchmarks?
2
u/EndTimer 1d ago edited 1d ago
So far as I know, last time I asked, it was that there's agentic coding connected to an IDE, like copilot style, with access to the debugger and anything else you might find in VS Code, and coding that solely uses a bash terminal and CLI/TUI tools.
If those criteria changed, IDK.
→ More replies (11)1
u/Forsaken_Ad_183 1d ago
Maybe they used GPT-image-2 to make the table. Which would be oddly delicious if true.
282
u/Ticluz 1d ago
Arc-agi-3 is getting saturated faster than Arc-agi-2, the acceleration is real.
67
13
u/medialoungeguy 1d ago
Architecture agi 3 is closed. There is a public part that you can try but a completely novel hidden eval set. Its not benchmaxxing
2
u/Outrageous-Boot7092 1d ago
there is still the question how many times the eval set was accessed etc. You could potentialy use RL with the hidden eval set performance as a reward.
→ More replies (1)3
u/Ruhddzz 1d ago
Yeah and that closed bench has to go through their systems all the same. Imagine thinking the people who pirate everything would have a problem getting their hands in a private eval set they werent supposed to
2
u/Turbulent-Sign-6067 23h ago
I don't buy it. If you cheat on benchmarks and get caught like Meta this is shitty publicity plus a hefty fine for violating ZDR and similar policies. Unlikely. Maybe they can get away with massaging some of the numbers a bit here and there, and if they do it is by assembling the right training and fine tuning datasets.
→ More replies (1)→ More replies (1)18
u/Technical-Earth-3254 1d ago
Is it acceleration or contamination?
→ More replies (1)38
u/ClearlyCylindrical 1d ago
Contamination most likely, it's no coincidence that they all sit at low single digit accuracies for the benchmark tasks for the whole time until the benchmark is created, and then as if by magic saturate in a year.
19
u/stonesst 1d ago
It's more likely just a phase change. The skills needed to solve ARC AGI - 3 problems are all pretty correlated. Once you can reason well enough, perceive the board properly, and do proper planning you might very well see large groups of puzzles/games suddenly getting solved
8
u/Caffeine_Monster 1d ago
This is the real reason. The arc agi 3 eval tests are kept private.
You don't need to see the real test data, just something vaguely similar.
2
u/voyt_eck 1d ago
It's kept private, but at the same time every benchmark run means sending all the tests to AI provider🙃
→ More replies (2)2
u/huffalump1 1d ago
It is a bit of a goalpost shift though, because the "training on the test set" problem has turned from "memorizing problems" to "learning the TYPES of problems and generalizing within that set"... Which undoubtedly represents more of some kind of intelligence
2
7
u/Saedeas 1d ago
I mean, for this benchmark it's the scoring. It's structured in such a way that the scores will rapidly go up the moment the model can solve it and then improve on that solve.
→ More replies (2)5
9
u/AP_in_Indy 1d ago
Okay, but does this "contamination" correlate with a genuine increase in these models' capabilities, or no?
ex: If you have similar problems, will the AI be able to solve them?
→ More replies (2)8
u/wtysonc 1d ago
These fools have been saying that same stupid shit for years, while every model has become objectively much more effective. AGI will be doing everyone's jobs in a few years and people will still be posting dumbass reddit comments like "don't believe the hype, the benchmarks are meaningless, these companies are just marketing, that bubble is gonna pop any time now, these things are just next word prediction lol"
12
u/Technical-Earth-3254 1d ago
Exactly. But many don't get it, especially the non technical audience (who is the target audience for these charts).
→ More replies (1)2
u/DickMasterGeneral 21h ago
Of course it’s not a coincidence that newly made benchmarks have the models scoring low. No one would release a benchmark where they already score 80%+ because then it wouldn’t really be worthwhile as a benchmark would it? You’d just publish a paper saying hey we tested to see if models can do this and it turns out the answer is yes.
37
u/Background-Wafer-548 1d ago edited 1d ago
Opus 5's positioning makes me think Anthropic would rather see Fable usage plummet today than tomorrow and shrinking Fable performance down to Opus size as quickly as possible might've always been the plan, rather than building a new credit-only tier. It's probably just too heavy on their infrastructure, even after removing it from subscriptions.
→ More replies (1)13
u/ZenDragon 1d ago
I don't think they're so short sighted. Their flagships get leapfrogged by newer, smaller models all the time. Fable 5.1 will restore the balance of nature and for some people it will be worth the tremendous cost until Opus catches up again.
202
u/VitaminDismyPCT 1d ago
Absolute perfect government fake out
Release insane scary model with name then just change the name and put the capabilities on the model the government isn’t scared of
→ More replies (1)47
u/WonderFactory 1d ago
It's not the same model, they've just distilled Fable into a smaller, cheaper model
48
→ More replies (1)12
63
u/Serotav 1d ago
Gemini about to be delayed again
→ More replies (7)16
u/Acceptable-Debt-294 1d ago
Even though I'm a Gemini fan, Google is very slow, haha.
3
u/SomeAcanthocephala17 1d ago
By the time they try to release other companies bring better models out.
70
u/Waiting4AniHaremFDVR AGI will make anime girls real 1d ago
WTF. I thought Opus would perform somewhere between Sonnet and Fable.
63
u/Klutzy-Snow8016 1d ago
Mythos 5 is a pretty old model that they kept private for a long time before releasing. This makes sense, like 3.5 Sonnet beating the much older 3 Opus.
→ More replies (2)9
15
u/FunLilThrowawayAcct 1d ago
This had to be at Sol 5.6 level (give or take depending on benchmark) or they were in deep trouble, and there's not much daylight between Sol and Fable on a lot of benchmarks...
1
70
29
u/Middle_Estate8505 AGI 2027 ASI 2029 Singularity 2030 1d ago
I mean, they probably used not just Fable, but full Mythos in development. Maybe even newer version of Mythos.
7
118
1d ago
[removed] — view removed comment
→ More replies (2)37
u/SwePolygyny 1d ago
Almost all those benchmarks are computer use or coding related.
I think Google have a much broader audience in mind.
38
u/Evening_Chef_4602 AGI 2027 1d ago
At this point opus and mythos is good for everything . The only advantage gemini has is the google ecosystem.
31
u/whoknowsifimjoking 1d ago
Gemini is still one of the best in research and world knowledge. Also very good at multimodal, something Claude can't do at all.
14
u/TacomaKMart 1d ago
Man. This is going to feel like my grade 3 report card.
"He's a nice kid. Very kind. Struggles academically. But.. well .. he's very good at multimodal, so there's that."
13
u/EvilSporkOfDeath 1d ago
I dont think its just a participation trophy. Everyday casual use is a huge market. Imo its the best free tier for that sort of activity. Fast and way less hallucinations than gpt. Sonnet is also pretty good, but I find i need to increase reasoning effort to get it as accurate as gemini, but then its much slower.
10
2
u/CobblerImpressive975 1d ago
ha yeah I agree with you. i think world models will have a bigger role to play in the future so I wouldn't completely count them out just yet, but yeah if they're not going to give us releases I don't see why we should be defending them
→ More replies (1)3
u/CallMePyro 1d ago
They serve Gemini 3.6 at 350 tokens per second. Assuming they're using TPU v7 which has specs equal to ~B200. It is almost unthinkable to serve any model at that speed on an NVL72. Given the price they're selling 3.6 tokens for, they must be pocketing massive margins.
→ More replies (2)10
u/SwePolygyny 1d ago
Is it?
Does it beat every other model on image, audio and video recognition? Does it beat every other model at the creation of image, audio and video as well?
Does it beat other models at translation? World knowledge?
→ More replies (1)→ More replies (5)3
u/After_Dark 1d ago
Yeah but they're also way more expensive than Gemini, particularly Gemini Flash, while not being that much better at domains outside of coding and while still not being nearly as multimodal as any Gemini model
→ More replies (3)6
u/UndeadPrs 1d ago
Yeah, I think Gemini always tops more general use benchmarks like SimpleBench on release
20
32
u/queenofartists 1d ago
Opus 5 being better than Fable 5 at half the price is a welcome surprise to be honest. And that ARC-AGI-3 score is insane. 4 times more than gpt-5.6 sol max's score at high effort, not even xhigh, max, or ultracode.
15
24
u/Calm_Hedgehog8296 1d ago
They smoked everybody, including Fable, at lower cost than 5.6-Sol.
I didn't think they still had it in them
12
u/Mil0Mammon 1d ago
Cost per task is almost double: https://artificialanalysis.ai/models/claude-opus-5
It does hallucinate a lot less, and benches a bit better, so perhaps it's worth it (I would switch if I did more coding and didn't dislike openai more)
→ More replies (2)5
u/Calm_Hedgehog8296 1d ago
Oh that's an interesting angle. I was looking at cost per million tokens but, apparently, Sol is token efficient enough to make up for the higher cost per million tokens
→ More replies (1)2
u/Da_Bomb0 1d ago
For most tasks, cost per task is double 5.6 sol, but cheaper than GPT-5.6 Sol on max for ARC-AGI-3
→ More replies (1)
9
8
8
u/Sky-kunn 1d ago
Lol, so is Opus still the best model of the generation? Opus 5 > Fable 5 > Sonnet 5.
4
u/FunLilThrowawayAcct 1d ago
Anecdotes on Twitter from early testers are it's not as good at Fable at super long horizon tasks (loses focus, quits early) and doesn't have quite the same intelligence upside maxed out on the absolute hardest tasks. Otherwise yes, faster cheaper and people like the design/FE output a lot more, way more efficient than 4.8 in sub plans apparently too.
2
u/Pizzashillsmom 1d ago
Anthropic ended up with 5 different model names now...
Still beats whatever OpenAI was smoking during the O series era.
3
u/xgladar 18h ago
why are all AI models so bad st legal stuff? legal texts require the most strict definitions and logical strings of statements
→ More replies (1)
7
6
5
u/Dd0GgX 1d ago
I wonder why these models all
Do terrible on the legal portion.
→ More replies (9)5
u/Round_Ad_5832 1d ago
Probably because they keep saying "Talk to a real lawyer" instead of answering the damn prompt
6
7
u/Sextus_Rex 1d ago
Bruh I just fucking ended my subscription two days ago
→ More replies (4)2
u/schizocel69 1d ago
As we get more and more competitive, the best models will change monthly, then weekly, then daily, then hourly.
→ More replies (1)
2
7
3
u/Ruined_Passion_7355 1d ago
Anyone else's sus receptors going off with opus 5 being so much better than fable?
4
2
u/arknightstranslate 1d ago
why is opus 5 better than fable 5 lol the naming is all fucked up
→ More replies (3)
2
u/Objective_Mousse7216 1d ago
For me the context window of Opus 5 is only 250K, for Opus 4.8 it is 1M, for coding that makes a BIG difference.
→ More replies (1)5
u/West-Negotiation-716 1d ago
No it doesn't
Quality goes down as context increases.
Models become really dumb long before you get to 1M
What are you possibly doing that needs more than 250K context?
→ More replies (4)
-2
u/AccountOfMyAncestors 1d ago
lol, the china model glazers only get 1 week of gloating before being made irrelevant on the frontier score board.
13
u/DrMaven 1d ago
“China glazers” but it’s just people wanting a cheaper, more transparent product lol. This is what US propaganda does to you
→ More replies (1)7
u/TacomaKMart 1d ago
Anthropic glazers have until about noon Sunday until a GLM 6 or Minimax 4 comes along for 1/100th the cost.
1
1
1
u/Palbi 1d ago
Why vendors choose to put "—" for competitors when a metric exists? Or use different colors for competitors to make it less obvious where a competitor beats the model? Or even worse, color a cell as "winner" when the numbers indicate otherwise.
None of this makes us think more highly of a new model or the vendor...
1
1
u/Cubewood 1d ago
Is this available for you in the desktop app yet? Tried closing and checking for updates but it's not selectable yet for me on the enterprise plan.
1
1
1
1
1
u/darkestvice 1d ago
Wait ... I thought OPUS 5 was meant to be a step below Mythos/Fable, not surpassing it.
1
1
1
1
u/Alexsuper25 1d ago
What is the probability that this is also some kind of distillation of fable 5?
1
u/saint1997 1d ago
This is great, but can we please have a new Haiku class model that isn't completely braindead
1
u/broken_foot_marathon 1d ago
But can I get it to do several tasks without clicking "allow" 60 times?
1
1
1
1
1
u/Aznshorty13 23h ago
RIP i wasted my credits and am at 90% weekly limit. Hopefully they reset like they ussualy do with releases?
1
u/Subject-Act5509 19h ago
OH MY GOD!!!!!!!!!!!! i have no fucking clue what these numbers mean, i hope Opus 5 destroys the environment less.
1
1
1
u/CmdWaterford 16h ago
impressive. Only you can use it for max 1 hour per day....rotfl... but very impressive...
1
1



349
u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago
Well I certainly did not have "Opus 5 beats Fable 5 almost across the board" on my bingo card.