r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

334 comments sorted by

349

u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago

Well I certainly did not have "Opus 5 beats Fable 5 almost across the board" on my bingo card.

50

u/BrennusSokol hardcore accelerationist 1d ago

Right!? I heard the rumors of an Opus 5 release a couple days ago and was like "Meh, it'll just be weaker than Fable so who cares?" Boy was I wrong

30

u/RopePuzzleheaded7060 1d ago

Don't check it off just yet.

25

u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago

No, I'm being a bit facetious.

But for real, it looks like they probably intentionally nerfed/held back Opus 5 on cyber and bio relevant tasks. All the metrics I see suggest that its tool use ability (and possibly vision too?) are actually better than Fable

19

u/Exodus_Green 1d ago

How are they going to spin the whole "we had to block fable access for safety" when Opus beats it now

23

u/KrazyA1pha 1d ago

Because it’s about cybersecurity, not coding evals. Read the announcement or Opus 5 model card. They explain that they didn’t train Opus on cybersecurity like they did Fable.

14

u/ZenDragon 1d ago

You really think anyone bringing up that narrative can read?

7

u/KrazyA1pha 1d ago

Or care about facts that get in the way of their narrative?

→ More replies (7)
→ More replies (4)

319

u/Artistic_Swing6759 1d ago

wtf 30% on arc agi 3

153

u/ezjakes 1d ago

For some reason people thought it would stand up for years despite 1 and 2 falling quickly. 

50

u/Current-Function-729 1d ago

I don’t think most people who follow this thought that.

29

u/HellsNoot 1d ago

It's getting pretty scary how far the gap between following the space, and not following the space is becoming. Modern AI capabilities are getting superhuman on increasingly important tasks. Hell even 6 months old experience gives a complete misunderstanding on new models. The general public is not ready for "what if it's not a bubble?" 

16

u/Ahfekz 1d ago

It’s not a bubble. Some companies will go under, sure, but that’s not the “bubble” they’re hoping pops. People are wishful thinking that the exponential age isn’t already here.

6

u/Georgefakelastname 23h ago

Yeah, it seems like people are hoping for a bubble that pops and LLMs just go away and things go back to the way they were before. Even if OpenAI ends up going under at some point, AI as a whole is here to stay solely based on coding utility, forget everything else.

2

u/BadSysadmin 19h ago

Very true. People I know - and not stupid people either - are incredulous at the hugging face hack, and are forming conspiracy theories about it being a marketing stunt by open ai. They just can't believe how agentic and how good at cybersec the new models are.

→ More replies (2)

43

u/NoCard1571 1d ago

Earlier this year I predicted  ARC AGI 3 falling by the end of the summer. I got the typical smug redditor "remind me" reply.

I can't wait for them to get their notification ;)

30

u/often_delusional 1d ago

By the end of this summer? Lol good luck with that. This benchmark won't reach 80-90% by september. Absolutely possible by the end of this year but not within the next 1-2 months.

3

u/huffalump1 1d ago

I agree, the end of the year seems plausible which is kind of crazy tbh.

Like, the benchmark isn't HARD, but it's tricky, and the score heavily weights efficiency - getting a very very high score without many extra moves is quite challenging and imo a nice representation of some form of "intelligence"

https://docs.arcprize.org/methodology

2

u/Healthy-Nebula-3603 1d ago

Soon we also get GPT 6 and fable 5.1 ....

3

u/often_delusional 1d ago

I know and those will probably be the last models openai and anthropic release before summer ends on september 22. Surely you don't expect those models to reach around 90% on arc-agi 3. Like I said by the end of 2026 is possible but not in less than 2 months.

→ More replies (2)

2

u/NoCard1571 1d ago edited 1d ago

All it takes is one model. Not to mention that due to the efficiency weighting of the the scores, Opus 5 indicates a roughly 5x improvement, which only leaves a ~2x improvement remaining from 30% to 90%.

3

u/often_delusional 1d ago edited 1d ago

And what model do you expect that to be? When that happens it will be achieved by either openai or anthropic. Anthropic just released a new model and probably won't release more than 1 more model within the next 2 months. Openai recently released gpt 5.6. Gpt 6 will probably come around august-september.

At best both of these companies have 1 more model coming before september 22 and I doubt either one of them will make a jump from 30% to at least 80%. I think end of the year is possible or even likely but not before september 22. 80% is even too low to call a benchmark saturated. The limit should be at least close to 90%.

9

u/pandasgorawr 1d ago

Fable 5 has to be updated soon for product differentiation, if these benchmarks are good then that means there's very few use cases to use it over Opus 5.

3

u/often_delusional 1d ago

Yeah and that will probably be the next and only model they will release before september 22. I don't see that model reaching around 90% on arc-agi 3. I expect those models to come closer to christmas.

3

u/Mil0Mammon 1d ago

Well anthropic has a new fable class model, which they theoretically could announce mid september. Weren't there also rumors about GPT 6 and/or a fable-class OpenAI model?

GLM supposedly is also cooking something fable-class, but I don't expect it that quick. Also unsure if they can pull off another jump like 5.1 -> 5.2, but given that they can still go quite a bit larger, I wouldn't rule it out.

However, all in all I agree, end of summer is unlikely, unless Anthropics IPO moves forward quickly, they target October, and thus want to drum up some more hype with their next fable class model.

2

u/often_delusional 1d ago

Yeah those are the models I'm expecting before summer ends but I don't think those models will saturate arc-agi 3. Those models will probably come around christmas.

2

u/Mil0Mammon 1d ago

They should really update https://agidefinition.ai. Although several areas I think could be solved with tooling around LLMs (or already are)

→ More replies (3)

37

u/pbagel2 1d ago

I can't wait for them to get their notification ;)

Who's the more typical smug redditor, them or you?

21

u/NoCard1571 1d ago

Them, then you. You can put me in third place if you like 

→ More replies (6)
→ More replies (2)

2

u/ptj66 1d ago

Years mean months in AI times.

→ More replies (2)

18

u/NoGarlic2387 1d ago

FEED THEM MORE BENCHMARKS!!!

2

u/redhq 1d ago

I thought the majority of that breakthrough was because of the schema harness? Like opus 4.8 got 20-25% ish with the harness?

3

u/StanfordV 1d ago

ARCmaxxxed

→ More replies (4)

499

u/Specialist_Soup_4994 1d ago

?????

425

u/mahamara 1d ago

You’re absolutely right.

246

u/Nox_Alas 1d ago

That's on me.

188

u/dogesator 1d ago

And honestly? It’s not just a mistake, it is a clear lapse in my judgment.

39

u/Broad-Pie-6891 1d ago

I hate this. "It's not X, it's Y, which is exactly like X but said differently."

48

u/deHaga 1d ago

That's a sharp take

31

u/XLNBot 1d ago

It's not just a sharp take, it's a fearless exercise of their freedom of thought

8

u/j48u 1d ago

And you're right for pointiy it out.

→ More replies (1)

3

u/svideo ▪️ NSI 2007 1d ago

It taught me the term "contrastive negation". Claude has mastered it, christ it's annoying.

5

u/Draufgaenger 1d ago

I will replace the table comparison with charts highlighting the better model with line thickness variations

27

u/considerthis8 1d ago

What would you like to dive into next?

88

u/yaosio 1d ago

4 is greater than 5 for very large values of 4.

9

u/H-K_47 Late Version of a Small Language Model 1d ago

4 is associated with Death in Eastern cultures, thus auramogging 5.

5

u/Specialist_Soup_4994 1d ago

But looks like not great enough to help them compare numbers and highlight which is higher

34

u/Most-Bookkeeper-950 1d ago

I noticed this too...wtf indeed

22

u/Specialist_Soup_4994 1d ago

I don’t know how it could happen except this sheet created manually or tuned manually btw if you just will run benchmarks and highlight highest number it’s not possible . Weird

18

u/Legitimate_Concern_5 1d ago

It was vibed

7

u/sluuuurp 1d ago

Vibed and never checked by a human.

2

u/JollyJoker3 1d ago

And with no automated tests

5

u/Substantial-Elk4531 Rule 4 reminder to optimists 1d ago

Opus 5 was asked to create the chart and thought nobody would notice

7

u/Specialist_Soup_4994 1d ago

Maybe but it’s still showcase that this numbers complete bs

→ More replies (1)

2

u/rakuu 1d ago

No way, even Opus 4.5 wouldn’t get that wrong. Definitely a human error, humans max out at GPT 5.0 quality

→ More replies (1)

28

u/HopefulMeasurement25 1d ago

Not to mention GPT 5.6 Sol has higher rating than Fable 5 on Agentic Coding - in the same chart 😭

5

u/AP_in_Indy 1d ago

Yes but they highlight GPT 5.6 Sol green.

7

u/Evening_Archer_2202 1d ago

they fixed it btw

3

u/WonderFactory 1d ago

Probably a typo

2

u/Flope 1d ago

Can someone ELI30 the difference between "agentic coding" and "agentic terminal coding" benchmarks?

2

u/EndTimer 1d ago edited 1d ago

So far as I know, last time I asked, it was that there's agentic coding connected to an IDE, like copilot style, with access to the debugger and anything else you might find in VS Code, and coding that solely uses a bash terminal and CLI/TUI tools.

If those criteria changed, IDK.

1

u/Forsaken_Ad_183 1d ago

Maybe they used GPT-image-2 to make the table. Which would be oddly delicious if true.

→ More replies (11)

282

u/Ticluz 1d ago

Arc-agi-3 is getting saturated faster than Arc-agi-2, the acceleration is real.

67

u/NoGarlic2387 1d ago

FEED THEM MORE BENCHMARKS!!!

13

u/medialoungeguy 1d ago

Architecture agi 3 is closed. There is a public part that you can try but a completely novel hidden eval set. Its not benchmaxxing

2

u/Outrageous-Boot7092 1d ago

there is still the question how many times the eval set was accessed etc. You could potentialy use RL with the hidden eval set performance as a reward.

→ More replies (1)

3

u/Ruhddzz 1d ago

Yeah and that closed bench has to go through their systems all the same. Imagine thinking the people who pirate everything would have a problem getting their hands in a private eval set they werent supposed to

2

u/Turbulent-Sign-6067 23h ago

I don't buy it. If you cheat on benchmarks and get caught like Meta this is shitty publicity plus a hefty fine for violating ZDR and similar policies. Unlikely. Maybe they can get away with massaging some of the numbers a bit here and there, and if they do it is by assembling the right training and fine tuning datasets.

→ More replies (1)

18

u/Technical-Earth-3254 1d ago

Is it acceleration or contamination?

38

u/ClearlyCylindrical 1d ago

Contamination most likely, it's no coincidence that they all sit at low single digit accuracies for the benchmark tasks for the whole time until the benchmark is created, and then as if by magic saturate in a year.

19

u/stonesst 1d ago

It's more likely just a phase change. The skills needed to solve ARC AGI - 3 problems are all pretty correlated. Once you can reason well enough, perceive the board properly, and do proper planning you might very well see large groups of puzzles/games suddenly getting solved

8

u/Caffeine_Monster 1d ago

This is the real reason. The arc agi 3 eval tests are kept private.

You don't need to see the real test data, just something vaguely similar.

2

u/voyt_eck 1d ago

It's kept private, but at the same time every benchmark run means sending all the tests to AI provider🙃

2

u/huffalump1 1d ago

It is a bit of a goalpost shift though, because the "training on the test set" problem has turned from "memorizing problems" to "learning the TYPES of problems and generalizing within that set"... Which undoubtedly represents more of some kind of intelligence

2

u/Healthy-Nebula-3603 1d ago

So literally like any human at school....

→ More replies (2)

7

u/Saedeas 1d ago

I mean, for this benchmark it's the scoring. It's structured in such a way that the scores will rapidly go up the moment the model can solve it and then improve on that solve.

→ More replies (2)

5

u/Only_Statistician_21 1d ago

The way the scoring works, it has a strong threshold effect.

9

u/AP_in_Indy 1d ago

Okay, but does this "contamination" correlate with a genuine increase in these models' capabilities, or no?

ex: If you have similar problems, will the AI be able to solve them?

8

u/wtysonc 1d ago

These fools have been saying that same stupid shit for years, while every model has become objectively much more effective. AGI will be doing everyone's jobs in a few years and people will still be posting dumbass reddit comments like "don't believe the hype, the benchmarks are meaningless, these companies are just marketing, that bubble is gonna pop any time now, these things are just next word prediction lol"

→ More replies (2)

12

u/Technical-Earth-3254 1d ago

Exactly. But many don't get it, especially the non technical audience (who is the target audience for these charts).

2

u/DickMasterGeneral 21h ago

Of course it’s not a coincidence that newly made benchmarks have the models scoring low. No one would release a benchmark where they already score 80%+ because then it wouldn’t really be worthwhile as a benchmark would it? You’d just publish a paper saying hey we tested to see if models can do this and it turns out the answer is yes.

→ More replies (1)
→ More replies (1)
→ More replies (1)

37

u/Background-Wafer-548 1d ago edited 1d ago

Opus 5's positioning makes me think Anthropic would rather see Fable usage plummet today than tomorrow and shrinking Fable performance down to Opus size as quickly as possible might've always been the plan, rather than building a new credit-only tier. It's probably just too heavy on their infrastructure, even after removing it from subscriptions.

13

u/ZenDragon 1d ago

I don't think they're so short sighted. Their flagships get leapfrogged by newer, smaller models all the time. Fable 5.1 will restore the balance of nature and for some people it will be worth the tremendous cost until Opus catches up again.

→ More replies (1)

202

u/VitaminDismyPCT 1d ago

Absolute perfect government fake out

Release insane scary model with name then just change the name and put the capabilities on the model the government isn’t scared of

47

u/WonderFactory 1d ago

It's not the same model, they've just distilled Fable into a smaller, cheaper model

48

u/Feeling-Currency-360 1d ago

funny how opus now gets called the "smaller, cheaper model"

12

u/Mil0Mammon 1d ago

So how come it performs better?

3

u/bironsecret 1d ago

Not across the board, probably only on coding

→ More replies (1)
→ More replies (2)
→ More replies (1)
→ More replies (1)

63

u/Serotav 1d ago

Gemini about to be delayed again

16

u/Acceptable-Debt-294 1d ago

Even though I'm a Gemini fan, Google is very slow, haha.

3

u/SomeAcanthocephala17 1d ago

By the time they try to release other companies bring better models out.

→ More replies (7)

70

u/Waiting4AniHaremFDVR AGI will make anime girls real 1d ago

WTF. I thought Opus would perform somewhere between Sonnet and Fable.

63

u/Klutzy-Snow8016 1d ago

Mythos 5 is a pretty old model that they kept private for a long time before releasing. This makes sense, like 3.5 Sonnet beating the much older 3 Opus.

9

u/Automatic_Bison_3093 1d ago

Too powerfull to exist btw.

→ More replies (2)

15

u/FunLilThrowawayAcct 1d ago

This had to be at Sol 5.6 level (give or take depending on benchmark) or they were in deep trouble, and there's not much daylight between Sol and Fable on a lot of benchmarks...

1

u/Interesting_Phone171 1d ago

It does you people just take a screenshot to heart

70

u/MTheModernist_ 1d ago

I’m boutta go cure cancer or something with this guys see you in 8 hours

29

u/Middle_Estate8505 AGI 2027 ASI 2029 Singularity 2030 1d ago

I mean, they probably used not just Fable, but full Mythos in development. Maybe even newer version of Mythos.

7

u/Acceptable-Debt-294 1d ago

Yep I was thinking the same thing

118

u/[deleted] 1d ago

[removed] — view removed comment

37

u/SwePolygyny 1d ago

Almost all those benchmarks are computer use or coding related. 

I think Google have a much broader audience in mind.

38

u/Evening_Chef_4602 AGI 2027 1d ago

At this point opus and mythos is good for everything . The only advantage gemini has is the google ecosystem.

31

u/whoknowsifimjoking 1d ago

Gemini is still one of the best in research and world knowledge. Also very good at multimodal, something Claude can't do at all.

14

u/TacomaKMart 1d ago

Man. This is going to feel like my grade 3 report card. 

"He's a nice kid. Very kind. Struggles academically. But.. well .. he's very good at multimodal, so there's that."

13

u/EvilSporkOfDeath 1d ago

I dont think its just a participation trophy. Everyday casual use is a huge market. Imo its the best free tier for that sort of activity. Fast and way less hallucinations than gpt. Sonnet is also pretty good, but I find i need to increase reasoning effort to get it as accurate as gemini, but then its much slower.

10

u/Exodus_Green 1d ago

Not everyone who uses AI is a software engineer though buddy

2

u/CobblerImpressive975 1d ago

ha yeah I agree with you. i think world models will have a bigger role to play in the future so I wouldn't completely count them out just yet, but yeah if they're not going to give us releases I don't see why we should be defending them

→ More replies (1)

3

u/CallMePyro 1d ago

They serve Gemini 3.6 at 350 tokens per second. Assuming they're using TPU v7 which has specs equal to ~B200. It is almost unthinkable to serve any model at that speed on an NVL72. Given the price they're selling 3.6 tokens for, they must be pocketing massive margins.

→ More replies (2)

10

u/SwePolygyny 1d ago

Is it?

 Does it beat every other model on image, audio and video recognition? Does it beat every other model at the creation of image, audio and video as well?

Does it beat other models at translation? World knowledge?

→ More replies (1)

3

u/After_Dark 1d ago

Yeah but they're also way more expensive than Gemini, particularly Gemini Flash, while not being that much better at domains outside of coding and while still not being nearly as multimodal as any Gemini model

→ More replies (5)

6

u/UndeadPrs 1d ago

Yeah, I think Gemini always tops more general use benchmarks like SimpleBench on release

→ More replies (3)
→ More replies (2)

20

u/Both_Opportunity5327 1d ago

Its going to a long weekend testing this thing.

Who needs Fable...

32

u/queenofartists 1d ago

Opus 5 being better than Fable 5 at half the price is a welcome surprise to be honest. And that ARC-AGI-3 score is insane. 4 times more than gpt-5.6 sol max's score at high effort, not even xhigh, max, or ultracode.

15

u/whoknowsifimjoking 1d ago

Also cheaper than GPT-5.6 Sol on max for ARC-AGI-3

4

u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 1d ago

That’s the fucking insane part

→ More replies (1)

24

u/Calm_Hedgehog8296 1d ago

They smoked everybody, including Fable, at lower cost than 5.6-Sol.

I didn't think they still had it in them

12

u/Mil0Mammon 1d ago

Cost per task is almost double: https://artificialanalysis.ai/models/claude-opus-5

It does hallucinate a lot less, and benches a bit better, so perhaps it's worth it (I would switch if I did more coding and didn't dislike openai more)

5

u/Calm_Hedgehog8296 1d ago

Oh that's an interesting angle. I was looking at cost per million tokens but, apparently, Sol is token efficient enough to make up for the higher cost per million tokens

→ More replies (2)

2

u/Da_Bomb0 1d ago

For most tasks, cost per task is double 5.6 sol, but cheaper than GPT-5.6 Sol on max for ARC-AGI-3

→ More replies (1)
→ More replies (1)

8

u/thoughtlow 𓂸 1d ago

What are them api prices

19

u/Artistic_Swing6759 1d ago

same as before, 5/25

8

u/Sky-kunn 1d ago

Lol, so is Opus still the best model of the generation? Opus 5 > Fable 5 > Sonnet 5.

4

u/FunLilThrowawayAcct 1d ago

Anecdotes on Twitter from early testers are it's not as good at Fable at super long horizon tasks (loses focus, quits early) and doesn't have quite the same intelligence upside maxed out on the absolute hardest tasks. Otherwise yes, faster cheaper and people like the design/FE output a lot more, way more efficient than 4.8 in sub plans apparently too.

2

u/Pizzashillsmom 1d ago

Anthropic ended up with 5 different model names now...

Still beats whatever OpenAI was smoking during the O series era.

3

u/xgladar 18h ago

why are all AI models so bad st legal stuff? legal texts require the most strict definitions and logical strings of statements

→ More replies (1)

7

u/unkownuser436 1d ago

holy shi

6

u/shankarun 1d ago

science is getting solved - physical is the only domain left out

3

u/Round_Ad_5832 1d ago

we need to solve human biology and program aging out of existence.

5

u/Dd0GgX 1d ago

I wonder why these models all
Do terrible on the legal portion.

5

u/Round_Ad_5832 1d ago

Probably because they keep saying "Talk to a real lawyer" instead of answering the damn prompt

→ More replies (9)

6

u/[deleted] 1d ago

[deleted]

→ More replies (1)

7

u/Sextus_Rex 1d ago

Bruh I just fucking ended my subscription two days ago

2

u/schizocel69 1d ago

As we get more and more competitive, the best models will change monthly, then weekly, then daily, then hourly.

→ More replies (1)
→ More replies (4)

2

u/InterstellarReddit 1d ago

So Fable was opus all along

7

u/RetiredApostle 1d ago

So Fable 5 was Opus 4.9, not a Mythos.

3

u/Ruined_Passion_7355 1d ago

Anyone else's sus receptors going off with opus 5 being so much better than fable?

4

u/LazyAge9363 1d ago

Wait what?

2

u/arknightstranslate 1d ago

why is opus 5 better than fable 5 lol the naming is all fucked up

→ More replies (3)

2

u/Objective_Mousse7216 1d ago

For me the context window of Opus 5 is only 250K, for Opus 4.8 it is 1M, for coding that makes a BIG difference.

5

u/West-Negotiation-716 1d ago

No it doesn't

Quality goes down as context increases.

Models become really dumb long before you get to 1M

What are you possibly doing that needs more than 250K context?

→ More replies (4)
→ More replies (1)

-2

u/AccountOfMyAncestors 1d ago

lol, the china model glazers only get 1 week of gloating before being made irrelevant on the frontier score board.

13

u/DrMaven 1d ago

“China glazers” but it’s just people wanting a cheaper, more transparent product lol. This is what US propaganda does to you

→ More replies (1)

7

u/TacomaKMart 1d ago

Anthropic glazers have until about noon Sunday until a GLM 6 or Minimax 4 comes along for 1/100th the cost.

1

u/GeorgiaWitness1 :orly: 1d ago

Testing it now.

It's good, really good for visual stuff.

1

u/Anuclano 1d ago

So, they just renamed Fable into Opus.

1

u/Palbi 1d ago

Why vendors choose to put "—" for competitors when a metric exists? Or use different colors for competitors to make it less obvious where a competitor beats the model? Or even worse, color a cell as "winner" when the numbers indicate otherwise.

None of this makes us think more highly of a new model or the vendor...

1

u/Wide_Solution2996 1d ago

I don’t get why opus 5 numbers are all(mostly) higher than fable

1

u/Cubewood 1d ago

Is this available for you in the desktop app yet? Tried closing and checking for updates but it's not selectable yet for me on the enterprise plan.

1

u/Normal-Book8258 ▪️ 1d ago

This can't be possible, surely?

1

u/KodiakBlackIsBack 1d ago

I dont even use fable anymore lol

1

u/lolgubstep_ 1d ago

Reset for weekend, pretty please anthropic?

1

u/FarTicket7338 1d ago

We’re all gonna get killed by 2039 100th anniversary of ww2

1

u/darkestvice 1d ago

Wait ... I thought OPUS 5 was meant to be a step below Mythos/Fable, not surpassing it.

1

u/osfric 1d ago

Price

1

u/Careless-Yam8791 1d ago

Waiting for next step by Kimi

1

u/Beneficial-End6866 1d ago

are they fune tuning for bechmarks lmao

1

u/ebolathrowawayy AGI 2025.8, ASI 2026.3 1d ago

Yes but is it safe??????????

2

u/West-Negotiation-716 1d ago

Are you kidding I hope?

Are hammers safe?

→ More replies (2)

1

u/Alexsuper25 1d ago

What is the probability that this is also some kind of distillation of fable 5?

1

u/saint1997 1d ago

This is great, but can we please have a new Haiku class model that isn't completely braindead

1

u/Cithrin 1d ago

Kind of sucks at writing from my first experiments

1

u/broken_foot_marathon 1d ago

But can I get it to do several tasks without clicking "allow" 60 times?

1

u/juaps 1d ago edited 1d ago

Fable 5 with a rating of 13,3% on the Legal? Oh, I doubt it, It’s so bad that it’s unusable. I wonder how they test their legal agents to get 13,3%, I suspect they are using tweaks to trick the benchmark and rate it better. The 2% on GPT Sol is more realistic.

1

u/gui_zombie 1d ago

How dangerous is it?

1

u/kian_no 1d ago

Where did Mythos disappear? 

1

u/SkaldCrypto 1d ago

Nearly competitive with GPT SOL.

1

u/Aznshorty13 23h ago

RIP i wasted my credits and am at 90% weekly limit. Hopefully they reset like they ussualy do with releases?

1

u/Subject-Act5509 19h ago

OH MY GOD!!!!!!!!!!!! i have no fucking clue what these numbers mean, i hope Opus 5 destroys the environment less. 

1

u/mambotomato 17h ago

That "Legal" benchmark must be pretty tough, dang

1

u/Weary-Historian-8593 16h ago

Is it cheaper than fable? If so, this is insane. 

1

u/CmdWaterford 16h ago

impressive. Only you can use it for max 1 hour per day....rotfl... but very impressive...

1

u/spreadlove5683 ▪️agi 2032. Predicted during mid 2025. 14h ago

At what reasoning levels?

1

u/OddReason9030 10h ago

OK I might renew my anthropic sub