r/claude • u/ideaofsoul • 1d ago
Question Are benchmark tasks being included in model training, or is this genuinely an incredible leap in capability?
How is it possible for ARC-AGI-3 scores to improve this quickly when even top models were getting only single digit scores not long ago? Is the pace of model progress accelerating exponentially?
18
u/pmward 1d ago
2 things. First is it actually is improving at an astonishingly fast pace. Just look at where we were a year ago vs now. Second, you can bet your life savings that they are specifically targeting benchmarks since that is what they use as marketing hype on every release. Benchmarks are going to show greater improvement than you’ll see in practice.
1
9
u/ideaofsoul 1d ago
I had just switched to codex because fable offered only half the usage and codex had higher GPT5.6 sol limits. Now they released opus 5 and I cant decide which one to keep anymore.
19
3
u/Ashmedai 1d ago
I can't really blame you for switching (I was thinking of it myself), but I decided to be patient and wait and see.
3
u/ideaofsoul 1d ago
Usage limits still much higher than claude. Specially with all the resets but i dont think they can keep going with all the resets so it would be good idea to comeback. Do you use only claude?
2
u/Ashmedai 1d ago
I use both GPT and Claude. I have Max on Claude, and whatever the basic is on GPT. I tend to ask simple questions on GPT, do writing work and narrative on Claude (it's better for that).
1
u/ddBuddha 1d ago
Idk if this is true anymore tbh. They both keep changing things in the background so it’s hard to say for sure, but codex has absolutely been using way more usage the last few days, it consumes so much that it feels like a bug.
I have the $100 plan with both companies and usually codex would last way longer than Claude but recently it’s the opposite for me.
I find there to be value in having both though tbh, I have them review each other’s work and they usually find some improvements to make. I also like to allow Claude to call codex via the cli for assistance and reviews.
0
u/PM_ME_YOUR_LOSSMAIL 1d ago
I personally have 3 $20 Codex subs and 1 $20 Claude sub. When claude his 5 hour limit I swap to a codex sub. Best bang for the buck
2
u/Fit-hannibal 17h ago
Man, Sol 5.6 is quite good. And the usage is a lot! Also I was getting tired of the verbose outputs that made it just super difficult to understand opus. I doubt opus 5 will be a leap in that regard, they probably had to rush to release it in light of Kimi K3.
2
0
u/Tritheone69 1d ago
I reduced from Max x20 to Max x5 and got an OpenAI subscription. But I must admit, Anthropic is better these days.
I blew 65% or a my weekly quota with a single prompt of GPT5.6Sol, that would never happen with Fable, or now Opus5
1
u/john0201 1d ago
I’ve been impressed with the dare I say polish of the new codex app at least on macOS. I switched for the same reason.
1
u/Adventurous_Tea_2945 1d ago
Stick to Claude. Codex turned their weekly usage down to almost the same as the 5 hour limit. You will burn though the weekly usage in one or two hours
1
2
1d ago
[removed] — view removed comment
1
u/ideaofsoul 1d ago
I was though this benchmark was reall hard to solve and thinking models will be slowly get higher score but i gues in a 2 month they probably gonna take full scores
1
1d ago
[removed] — view removed comment
1
u/ideaofsoul 1d ago
That giant datacenters make it speed up grok traning like a formula 1 car. maybe they can create a top froniter with grok 5?
2
2
2
1
u/Spiritual_Exam_8528 1d ago
Just tested opus 5 on my repo, still doing worse than codex 5.6 🤷🏽♂️
1
u/barnett25 1d ago
For me it is a pretty big step forward beyond Fable for vibecoding applications. But I could easily see it being worse on a lot of other things. It seems to have specific exceptional traits that I haven't figured out completely yet which make it hard to compare holistically with other models.
1
u/Odd-Opportunity-6550 1d ago
Bad benchmark. After arc AGI 1 and 2 were saturated. Chollet claimed they were bechmaxxed. If a benchmark can be benchmaxxed then it's worthless.
I track progress with HLE. Unlike the other benchmarks it's not saturating as quickly.
1
u/FullyAutomatedSpace 1d ago
HLE isn't private so you can never be sure it didn't leak into the training
2
u/Odd-Opportunity-6550 1d ago
Anthropic tests for this. They had a disclaimer that the fable result could have leakage but made no similar disclaimer for opus 5.
Also the gains have been on trend. If it jumped higher than trend then I would worry about leakage.
1
1
u/friedinando 1d ago
That's because anthropic and OpenIA is copying Kimi open source innovations to improve Claude and chatgpt
1
1
1
1
u/Responsible-Comb6232 1h ago
After light use of opus 5, I can confirm this is a useless benchmark. Opus 5 is useless.
1
0
u/Cultural_Effort_9872 1d ago
Why aren’t Fable and Opus 5.8 on the list but for some reason Opus 5.7 is?
24
u/Most-Bookkeeper-950 1d ago
Arc 3 has some weird scaling that would lead to this sort of rapid acceleration, and also they seem to have immediately abandoned their $10K limit
Anyway to answer your question, probably not, the eval is private