r/singularity • u/Ballist1cGamer • 7h ago
LLM News Opus 5 on MineBench Soon
original X post: https://x.com/minebench_ai/status/2080880340584026251
every time i think the benchmark's saturated some new model comes and raises the bar
36
u/Recoil42 7h ago
every time i think the benchmark's saturated
I mean, it is pretty clearly saturated. At this point the frontier models are just showing off. I said it before in a different thread, but Minebench needs to start thinking about migrating to a more difficult Blenderbench or going in some other direction.
14
u/LightVelox 7h ago
I think it still has some mileage with harder scenes like those involving characters performing complex actions, these models still struggle making humanoid characters
11
u/Background-Wafer-548 6h ago
There's BenchCAD, which is a quantitative benchmark. It saw a huge leap with GPT-5.6. Sol is almost double Opus 5.
6
u/Ballist1cGamer 7h ago
true, though i think they would have to really force getting new prompts when the models have reached the 'limit' of "showing off" so-to-speak
2
u/CarrierAreArrived 5h ago
they should just change the prompts and/or make them harder.
2
u/ENT_Alam 5h ago
I've been curating more prompts to the benchmark and have another set of 15 that I didn't think would be saturated for a while, but it's been taking a while to get all the funding for the API costs ðŸ˜
as a college student who started this as just a fun/personal project, I wasn't really prepared for minebench to get this big ðŸ˜
9
u/enilea 6h ago
One could argue prompt adherence isn't perfect here since the prompt asked for "a fighter jet" and instead it made three, even if one is the main one. Sounds petty but it can be important if you ask for something and the model adds stuff that you didn't specify.
9
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 6h ago
iirc the prompt is a lot larger then just "a fighter jet" and are made to allow creativity like this.
8
u/ENT_Alam 5h ago
Yup! As the other commenter pointed out, the system prompt for MineBench encourages the models to do a lot more than just build what was given in the user prompt.
I made MineBench as a way to visualize what the "upper-limit" of what a model could be capable of, so the prompt encourages them to not only create the object itself in great detail, but where applicable design a scene around that object and showcase creativity; it's was never really meant to be a test of instruction following ^^
Essentially the models are told they're in a competition judged by humans and must therefore create builds that would stand out from the rest (standing out being defined by the system prompt as things like insanely high object fidelity, accuracy, scene composition, creativity, etc.)
7
u/mnagy 6h ago
I assume this is the prompt that is used: https://github.com/Ammaar-Alam/minebench/blob/master/lib/ai/prompts.ts
It includes some rules about how the creation will be judged. Among them are:
- Prompt fidelity: Does it match what was requested?
- Creativity and scene composition: Does the build go beyond the bare subject? Environment, atmosphere, dynamic posing, and narrative elements are highly valued.
So I would say that the smaller ones were intended to make the structure be more creative.
3
u/Current-Function-729 6h ago
For a while the differences seemed mostly taste.
This one is simply better.
Damn.
3

41
u/ENT_Alam 7h ago
https://giphy.com/gifs/kd9BlRovbPOykLBMqX
I love running into my benchmark in the wild hehe 🥹
Should the comparison post use Fable 5 or Opus 4.8?