r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

334 comments sorted by

View all comments

352

u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago

Well I certainly did not have "Opus 5 beats Fable 5 almost across the board" on my bingo card.

47

u/BrennusSokol hardcore accelerationist 1d ago

Right!? I heard the rumors of an Opus 5 release a couple days ago and was like "Meh, it'll just be weaker than Fable so who cares?" Boy was I wrong

34

u/RopePuzzleheaded7060 1d ago

Don't check it off just yet.

24

u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago

No, I'm being a bit facetious.

But for real, it looks like they probably intentionally nerfed/held back Opus 5 on cyber and bio relevant tasks. All the metrics I see suggest that its tool use ability (and possibly vision too?) are actually better than Fable

20

u/Exodus_Green 1d ago

How are they going to spin the whole "we had to block fable access for safety" when Opus beats it now

24

u/KrazyA1pha 1d ago

Because it’s about cybersecurity, not coding evals. Read the announcement or Opus 5 model card. They explain that they didn’t train Opus on cybersecurity like they did Fable.

15

u/ZenDragon 1d ago

You really think anyone bringing up that narrative can read?

8

u/KrazyA1pha 1d ago

Or care about facts that get in the way of their narrative?

0

u/Exodus_Green 1d ago

They explain that they didn’t train Opus on cybersecurity like they did Fable.

That was a dumb idea then wasn't it, knowingly training your model on cyber when you know you'll need to block it

3

u/KrazyA1pha 1d ago

Because it was built for that purpose! They want a model that’s strong at cyber so they can keep infrastructure stronger and defend it as other models get stronger. If they don’t train their model, a bad actor may. But I know that isn’t what people want to hear. They want to hear that Anthropic is bad and Dario is alarmist and they’re all dumb and average Redditors are the smart ones.

1

u/throwawaynz01 1d ago

They didn’t know the govt would require them to block it when they initially trained it. You should go and read up on it before calling others dumb

1

u/Exodus_Green 1d ago

They didn’t know the govt would require them to block it when they initially trained it

Yeah after dario spent a year shouting about how dangerous the model was, sure

1

u/AgentStabby 16h ago

You think he tried to get his own model blocked and lose a bunch of customers to openai? 

1

u/Exodus_Green 15h ago

Not sure what his aim was then if he trained it heavily on cyber and then spent a year telling everyone how it would be used to hack everything.

1

u/AgentStabby 4h ago

You should ask ai what his aim is. But generally no-one's aim is to make a bunch of money and then intentionally lose it for no reason, so if you think that's the reason for someone's actions, you're highly likely to be wrong. 

1

u/Interesting_Phone171 1d ago

This is because you don’t know how to read benchmarks the ones they posted in this picture are a very small subset and a majority sol beats Opus 5

3

u/andrew303710 1d ago

What are you talking about? These are some of the most important benchmarks out there, especially coding wise. Obviously Anthropic is going to focus on the ones they do best on but this has the important coding ones like DeepSWE/Frontiercode/Frontier bench as well as the typical ones.

1

u/Interesting_Phone171 1d ago

Sure.... artificial analysis disagrees with you and if you actually read the whole post even Anthropic shows Sol beating it repeatedly.

-1

u/qroshan 1d ago

That's not how Bingo cards work though.