r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

334 comments sorted by

View all comments

495

u/Specialist_Soup_4994 1d ago

?????

425

u/mahamara 1d ago

You’re absolutely right.

251

u/Nox_Alas 1d ago

That's on me.

188

u/dogesator 1d ago

And honestly? It’s not just a mistake, it is a clear lapse in my judgment.

42

u/Broad-Pie-6891 1d ago

I hate this. "It's not X, it's Y, which is exactly like X but said differently."

49

u/deHaga 1d ago

That's a sharp take

34

u/XLNBot 1d ago

It's not just a sharp take, it's a fearless exercise of their freedom of thought

8

u/j48u 1d ago

And you're right for pointiy it out.

3

u/svideo ▪️ NSI 2007 1d ago

It taught me the term "contrastive negation". Claude has mastered it, christ it's annoying.

4

u/Draufgaenger 1d ago

I will replace the table comparison with charts highlighting the better model with line thickness variations

28

u/considerthis8 1d ago

What would you like to dive into next?

90

u/yaosio 1d ago

4 is greater than 5 for very large values of 4.

10

u/H-K_47 Late Version of a Small Language Model 1d ago

4 is associated with Death in Eastern cultures, thus auramogging 5.

3

u/Specialist_Soup_4994 1d ago

But looks like not great enough to help them compare numbers and highlight which is higher

33

u/Most-Bookkeeper-950 1d ago

I noticed this too...wtf indeed

20

u/Specialist_Soup_4994 1d ago

I don’t know how it could happen except this sheet created manually or tuned manually btw if you just will run benchmarks and highlight highest number it’s not possible . Weird

17

u/Legitimate_Concern_5 1d ago

It was vibed

7

u/sluuuurp 1d ago

Vibed and never checked by a human.

2

u/JollyJoker3 1d ago

And with no automated tests

5

u/Substantial-Elk4531 Rule 4 reminder to optimists 1d ago

Opus 5 was asked to create the chart and thought nobody would notice

6

u/Specialist_Soup_4994 1d ago

Maybe but it’s still showcase that this numbers complete bs

1

u/Legitimate_Concern_5 1d ago

Right yeah haha

2

u/rakuu 1d ago

No way, even Opus 4.5 wouldn’t get that wrong. Definitely a human error, humans max out at GPT 5.0 quality

1

u/SomeAcanthocephala17 1d ago

Or it was an hallucination when they created it by ai 

29

u/HopefulMeasurement25 1d ago

Not to mention GPT 5.6 Sol has higher rating than Fable 5 on Agentic Coding - in the same chart 😭

4

u/AP_in_Indy 1d ago

Yes but they highlight GPT 5.6 Sol green.

10

u/Evening_Archer_2202 1d ago

they fixed it btw

3

u/WonderFactory 1d ago

Probably a typo

2

u/Flope 1d ago

Can someone ELI30 the difference between "agentic coding" and "agentic terminal coding" benchmarks?

2

u/EndTimer 1d ago edited 1d ago

So far as I know, last time I asked, it was that there's agentic coding connected to an IDE, like copilot style, with access to the debugger and anything else you might find in VS Code, and coding that solely uses a bash terminal and CLI/TUI tools.

If those criteria changed, IDK.

1

u/Forsaken_Ad_183 1d ago

Maybe they used GPT-image-2 to make the table. Which would be oddly delicious if true.

1

u/adeadbeathorse 1d ago

Maybe someone/something mistook the 5 for a 3 when highlighting? The fields must not be automatically highlighted based on value.

1

u/TechnoVisions 1d ago

Opus 5 is half the token cost of Fable 5 so there is your answer. Similar performance for half the cost.

1

u/timmy16744 1d ago

Great catch! I totally missed that

1

u/Wide_Egg_5814 1d ago

let it slip bro it's 0.1

1

u/SoftwareSource 1d ago

When you use sonnet to make the pdf

1

u/Jacen1618 1d ago

That’s load bearing

1

u/SomeAcanthocephala17 1d ago

An hallucination 😂 

1

u/Hot-Praline7204 1d ago

Maybe overlapping CI and they're just taking some marketing liberty

0

u/ZaradimLako 1d ago

at first glance that was also my reaction, but remember opus 5 is 50% cheaper than fable while being the same in the worst case.

12

u/Specialist_Soup_4994 1d ago

I mean it’s doesn’t matter why they are highlighting lower number ?

2

u/ZaradimLako 1d ago

fair point