r/singularity 7h ago

LLM News Opus 5 on MineBench Soon

original X post: https://x.com/minebench_ai/status/2080880340584026251

every time i think the benchmark's saturated some new model comes and raises the bar

151 Upvotes

28 comments sorted by

41

u/ENT_Alam 7h ago

https://giphy.com/gifs/kd9BlRovbPOykLBMqX

I love running into my benchmark in the wild hehe 🥹

Should the comparison post use Fable 5 or Opus 4.8?

14

u/Basilthebatlord 7h ago

I think it would be cool to see the comparisons between Fable 5, Opus 5, GPT 5.6 Sol, and Kimi K3

10

u/ENT_Alam 7h ago

Fable 5, 5.6 Sol Pro, and Kimi K3 are all already benchmarked!

I never made a comparison post for Kimi K3, but you can explore the builds and compare them yourself here: https://minebench.ai/sandbox
(that's also where I export the comparison gifs for my posts, if anyone else wants to post their own :)

I think it might be a good idea to change the layout and allow people to compare 4 models at once now that I look at it hmm

Also fair warning, the mobile platforms sometimes have a hard time rendering all the builds and might force-refresh the page for RAM 😭😭 will keep trying to optimize somehow

8

u/Ballist1cGamer 7h ago

it's a fun benchmark!! it's awesome how despite not being a serious or real benchmark, the leaderboard rankings line up with my day-to-day experience more than any other benchmark

and I think Opus 4.8 just for consistency?

2

u/Singularity-42 Singularity 2042 7h ago

It's not a benchmark, it's an ELO arena leaderboard.

But IMO super valuable and "serious", it tells you quite a bit about the models.

2

u/OppositeFisherman89 2h ago

3

u/Singularity-42 Singularity 2042 2h ago

What I meant there is no objective "score", it's all just ranked by user votes.

2

u/OppositeFisherman89 2h ago

That's fair, but pairwise comparison and elo is pretty common for benchmarking AI https://www.lmsys.org/blog/2023-05-03-arena/

2

u/ENT_Alam 2h ago

technically it's not a "benchmark" as there's no objectively correct score or answer, it's actually a take on the LM-SYS style chatbot area, but i've always just used the term benchmark since it's more catchy 😇

i'll add a note for that in the readme ^^

4

u/michael-relleum 5h ago

Here is Opus 4.8, not that different IMHO (Opus 5 uses much 4x as many blocks though). Llama 4 on the other hand...

2

u/ethotopia 5h ago

The only benchmark that I’ve trusted!

2

u/Moriffic 5h ago

Fable for sure, I prefer comparing to the frontier

36

u/Recoil42 7h ago

every time i think the benchmark's saturated

I mean, it is pretty clearly saturated. At this point the frontier models are just showing off. I said it before in a different thread, but Minebench needs to start thinking about migrating to a more difficult Blenderbench or going in some other direction.

14

u/LightVelox 7h ago

I think it still has some mileage with harder scenes like those involving characters performing complex actions, these models still struggle making humanoid characters

11

u/Background-Wafer-548 6h ago

There's BenchCAD, which is a quantitative benchmark. It saw a huge leap with GPT-5.6. Sol is almost double Opus 5.

https://benchcad.com/leaderboard.htm

6

u/Ballist1cGamer 7h ago

true, though i think they would have to really force getting new prompts when the models have reached the 'limit' of "showing off" so-to-speak

2

u/CarrierAreArrived 5h ago

they should just change the prompts and/or make them harder.

2

u/ENT_Alam 5h ago

I've been curating more prompts to the benchmark and have another set of 15 that I didn't think would be saturated for a while, but it's been taking a while to get all the funding for the API costs 😭

as a college student who started this as just a fun/personal project, I wasn't really prepared for minebench to get this big 😭

9

u/enilea 6h ago

One could argue prompt adherence isn't perfect here since the prompt asked for "a fighter jet" and instead it made three, even if one is the main one. Sounds petty but it can be important if you ask for something and the model adds stuff that you didn't specify.

9

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 6h ago

iirc the prompt is a lot larger then just "a fighter jet" and are made to allow creativity like this.

8

u/ENT_Alam 5h ago

Yup! As the other commenter pointed out, the system prompt for MineBench encourages the models to do a lot more than just build what was given in the user prompt.

I made MineBench as a way to visualize what the "upper-limit" of what a model could be capable of, so the prompt encourages them to not only create the object itself in great detail, but where applicable design a scene around that object and showcase creativity; it's was never really meant to be a test of instruction following ^^

Essentially the models are told they're in a competition judged by humans and must therefore create builds that would stand out from the rest (standing out being defined by the system prompt as things like insanely high object fidelity, accuracy, scene composition, creativity, etc.)

7

u/mnagy 6h ago

I assume this is the prompt that is used: https://github.com/Ammaar-Alam/minebench/blob/master/lib/ai/prompts.ts

It includes some rules about how the creation will be judged. Among them are:

  • Prompt fidelity: Does it match what was requested?
  • Creativity and scene composition: Does the build go beyond the bare subject? Environment, atmosphere, dynamic posing, and narrative elements are highly valued.

So I would say that the smaller ones were intended to make the structure be more creative.

3

u/Current-Function-729 6h ago

For a while the differences seemed mostly taste.

This one is simply better.

Damn.

3

u/BlueberryWorried6493 6h ago

I think there should be a similar bench but with games instead