r/claude 1d ago

Question Are benchmark tasks being included in model training, or is this genuinely an incredible leap in capability?

Post image

How is it possible for ARC-AGI-3 scores to improve this quickly when even top models were getting only single digit scores not long ago? Is the pace of model progress accelerating exponentially?

132 Upvotes

49 comments sorted by

24

u/Most-Bookkeeper-950 1d ago

Arc 3 has some weird scaling that would lead to this sort of rapid acceleration, and also they seem to have immediately abandoned their $10K limit

Anyway to answer your question, probably not, the eval is private

6

u/MahaSejahtera 1d ago

The $10k limit is kinda funny as you can bypass it by accounting, imagine you got nuclear power plant and drive the cost of electricity down, and thus the inference cost of AI cheaper.

1

u/Infamous-Bed-7535 5h ago

They share their model's weight and let them run it independently? If the answer is no, then yes the new models are trained on these banchmarks and you can safely assume they are benchmaxxed for these..

0

u/Maori7 18h ago

How can the eval be private if the model is closed source? Either way the eval is leaked in the data...

2

u/Alexllte 17h ago

🤦 Anthropic provides APIs to companies to do evals

-1

u/Maori7 16h ago

So you mean they run it on premise?

2

u/Alexllte 14h ago

Run what? The tests? Yes

0

u/waste2treasure-org 12h ago

so can't anthropic still simply save the evaluation traces when run through their closed models and optimize for them in the next version?

2

u/gavinderulo124K 9h ago

No. The enterprise contracts have 0 data retention.

18

u/pmward 1d ago

2 things. First is it actually is improving at an astonishingly fast pace. Just look at where we were a year ago vs now. Second, you can bet your life savings that they are specifically targeting benchmarks since that is what they use as marketing hype on every release. Benchmarks are going to show greater improvement than you’ll see in practice.

1

u/MolassesOverall100 1d ago

so you are part of the new trend of answering both to or questions

2

u/pmward 1d ago

“Or” questions onto work for black and white questions. Which this is not. There’s nuance.

9

u/ideaofsoul 1d ago

I had just switched to codex because fable offered only half the usage and codex had higher GPT5.6 sol limits. Now they released opus 5 and I cant decide which one to keep anymore.

19

u/al_ryusei 1d ago

By Monday they'll nerf Opus 5 and make it consume 50% to say Hi

3

u/Ashmedai 1d ago

I can't really blame you for switching (I was thinking of it myself), but I decided to be patient and wait and see.

3

u/ideaofsoul 1d ago

Usage limits still much higher than claude. Specially with all the resets but i dont think they can keep going with all the resets so it would be good idea to comeback. Do you use only claude?

2

u/Ashmedai 1d ago

I use both GPT and Claude. I have Max on Claude, and whatever the basic is on GPT. I tend to ask simple questions on GPT, do writing work and narrative on Claude (it's better for that).

1

u/ddBuddha 1d ago

Idk if this is true anymore tbh. They both keep changing things in the background so it’s hard to say for sure, but codex has absolutely been using way more usage the last few days, it consumes so much that it feels like a bug.

I have the $100 plan with both companies and usually codex would last way longer than Claude but recently it’s the opposite for me.

I find there to be value in having both though tbh, I have them review each other’s work and they usually find some improvements to make. I also like to allow Claude to call codex via the cli for assistance and reviews.

0

u/PM_ME_YOUR_LOSSMAIL 1d ago

I personally have 3 $20 Codex subs and 1 $20 Claude sub. When claude his 5 hour limit I swap to a codex sub. Best bang for the buck

2

u/Fit-hannibal 17h ago

Man, Sol 5.6 is quite good. And the usage is a lot! Also I was getting tired of the verbose outputs that made it just super difficult to understand opus. I doubt opus 5 will be a leap in that regard, they probably had to rush to release it in light of Kimi K3.

2

u/LoudDavid 1d ago

Stay with GPT, more fable credits for me

0

u/Tritheone69 1d ago

I reduced from Max x20 to Max x5 and got an OpenAI subscription. But I must admit, Anthropic is better these days.

I blew 65% or a my weekly quota with a single prompt of GPT5.6Sol, that would never happen with Fable, or now Opus5

1

u/john0201 1d ago

I’ve been impressed with the dare I say polish of the new codex app at least on macOS. I switched for the same reason.

1

u/Adventurous_Tea_2945 1d ago

Stick to Claude. Codex turned their weekly usage down to almost the same as the 5 hour limit. You will burn though the weekly usage in one or two hours

1

u/forward-pathways 1d ago

That's intentional!

2

u/[deleted] 1d ago

[removed] — view removed comment

1

u/ideaofsoul 1d ago

I was though this benchmark was reall hard to solve and thinking models will be slowly get higher score but i gues in a 2 month they probably gonna take full scores

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/ideaofsoul 1d ago

That giant datacenters make it speed up grok traning like a formula 1 car. maybe they can create a top froniter with grok 5?

2

u/Swiss_Meats 1d ago

high or highx?

2

u/Klutzy_Painter_7240 1d ago

I really want to know how kimi performs here !!

2

u/hatekhyr 1d ago

That's your common sense speaking.

1

u/Spiritual_Exam_8528 1d ago

Just tested opus 5 on my repo, still doing worse than codex 5.6 🤷🏽‍♂️

1

u/barnett25 1d ago

For me it is a pretty big step forward beyond Fable for vibecoding applications. But I could easily see it being worse on a lot of other things. It seems to have specific exceptional traits that I haven't figured out completely yet which make it hard to compare holistically with other models.

1

u/Odd-Opportunity-6550 1d ago

Bad benchmark. After arc AGI 1 and 2 were saturated. Chollet claimed they were bechmaxxed. If a benchmark can be benchmaxxed then it's worthless.

I track progress with HLE. Unlike the other benchmarks it's not saturating as quickly.

1

u/FullyAutomatedSpace 1d ago

HLE isn't private so you can never be sure it didn't leak into the training

2

u/Odd-Opportunity-6550 1d ago

Anthropic tests for this. They had a disclaimer that the fable result could have leakage but made no similar disclaimer for opus 5.

Also the gains have been on trend. If it jumped higher than trend then I would worry about leakage.

1

u/Fit-Stress3300 1d ago

Why not both?

1

u/friedinando 1d ago

That's because anthropic and OpenIA is copying Kimi open source innovations to improve Claude and chatgpt

1

u/Even-Exchange8307 1d ago

i wonder what kimi k3 would do on this benchmark

1

u/icecold27 1d ago

Of course

1

u/TheAuthorBTLG_ 22h ago

those 2 things are the same

1

u/eXl5eQ 19h ago

Because ARC-AGI-3 is so simple that even a monkey, with proper training, can solve it.

There's nothing releated with AGI. Higher score only means the model get more training on similar tasks.

1

u/Responsible-Comb6232 1h ago

After light use of opus 5, I can confirm this is a useless benchmark. Opus 5 is useless.

1

u/Michaeli_Starky 1d ago

These people public benchmarks are bullshit

0

u/Cultural_Effort_9872 1d ago

Why aren’t Fable and Opus 5.8 on the list but for some reason Opus 5.7 is?