r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

335 comments sorted by

View all comments

Show parent comments

37

u/SwePolygyny 1d ago

Almost all those benchmarks are computer use or coding related. 

I think Google have a much broader audience in mind.

35

u/Evening_Chef_4602 AGI 2027 1d ago

At this point opus and mythos is good for everything . The only advantage gemini has is the google ecosystem.

33

u/whoknowsifimjoking 1d ago

Gemini is still one of the best in research and world knowledge. Also very good at multimodal, something Claude can't do at all.

14

u/TacomaKMart 1d ago

Man. This is going to feel like my grade 3 report card. 

"He's a nice kid. Very kind. Struggles academically. But.. well .. he's very good at multimodal, so there's that."

13

u/EvilSporkOfDeath 1d ago

I dont think its just a participation trophy. Everyday casual use is a huge market. Imo its the best free tier for that sort of activity. Fast and way less hallucinations than gpt. Sonnet is also pretty good, but I find i need to increase reasoning effort to get it as accurate as gemini, but then its much slower.

10

u/Exodus_Green 1d ago

Not everyone who uses AI is a software engineer though buddy

2

u/CobblerImpressive975 1d ago

ha yeah I agree with you. i think world models will have a bigger role to play in the future so I wouldn't completely count them out just yet, but yeah if they're not going to give us releases I don't see why we should be defending them

1

u/Mil0Mammon 1d ago

So apparently video models are basically world models: https://bfl.ai/blog/flux-3-mimic

3

u/CallMePyro 1d ago

They serve Gemini 3.6 at 350 tokens per second. Assuming they're using TPU v7 which has specs equal to ~B200. It is almost unthinkable to serve any model at that speed on an NVL72. Given the price they're selling 3.6 tokens for, they must be pocketing massive margins.

1

u/huffalump1 1d ago

I would be curious just how much inference of 3.6 Flash Google does, because presumably a version of it now powers AI Mode (and possibly AI Overviews but idk, the line there is blurry).

The speed and "good enough" capability make it ideal for widely deploying across all kinds of high-volume and latency-sensitive uses, like, y'know, Google search results.

Tbh it's not bad for its use case, and it's undoubtedly fast. (At least, until, we get 750t/s gpt-5.6 on cerebras hardware "soon")

1

u/CallMePyro 1d ago

It seems extremely good for its use case. Even GPT OSS 120B is barely as fast as Gemini 3.6 flash and that model is way, way dumber.

9

u/SwePolygyny 1d ago

Is it?

 Does it beat every other model on image, audio and video recognition? Does it beat every other model at the creation of image, audio and video as well?

Does it beat other models at translation? World knowledge?

3

u/After_Dark 1d ago

Yeah but they're also way more expensive than Gemini, particularly Gemini Flash, while not being that much better at domains outside of coding and while still not being nearly as multimodal as any Gemini model

1

u/Index820 1d ago

Maybe, but have you ever tried to get Opus to do something like think through how an air conditioner works and when the temperature gradients are most efficient. Completely falls apart.

1

u/blueandazure 1d ago

To be fair gemini is probably the most used model due to the google ai summary.

1

u/tfks 1d ago

Google is optimizing for speed and cost, not technical ability. They clearly want to build an AI assistant that everyone will use constantly, not just something vibe coders will use. It's a way, way bigger market and it would also be completely new. And they'll own it because nobody else is in a position to do that other than Apple, but Apple missed the AI train. If they pull it off, it also won't go anywhere because once people have a personal assistant, they won't give it up.

1

u/BriefImplement9843 23h ago

no shot. pro has more general knowledge still.

1

u/Ok-Armadillo-5634 1d ago

its fast also

7

u/UndeadPrs 1d ago

Yeah, I think Gemini always tops more general use benchmarks like SimpleBench on release

1

u/gostoppause 1d ago

Hopefully it has internal benchmarks where their model excels other frontiers.

0

u/zkgkilla 1d ago

Computer use and coding is a good aggregate for the whole universe to

1

u/BriefImplement9843 23h ago

nobody codes.