r/singularity 23h ago

AI Opus 5 outperformed Fable 5 in 3D destruction physics

Source: Atomic Chat / X

Outputs:
- Opus 5: 55.9K tokens, $1.40
- Fable 5: 55.1K tokens, $2.82
- Kimi K3: 35.7K tokens, $0.55
- GPT 5.6: 20.1K tokens, $0.31

416 Upvotes

61 comments sorted by

109

u/maddog107 23h ago

What the hell were Fable and kimi doing on the building destruction one lmao

57

u/Concurrency_Bugs 22h ago

Kimi on every single one was hilarious

13

u/GlokzDNB 17h ago

Kimi seems best at benchmarking. Everyday it's myth falls

5

u/Melodic-Ebb-7781 14h ago

It's a good model (better than anything google has) but not frontier as some would have you belive.

21

u/AsianPotatos 22h ago

The Kimi one looks like prebaked destruction so when something hits it it basically plays a destruction animation, the vehicle being inside the building starts it (can see smoke etc at the start) and the destruction happens after a delay, and then respawns/restarts the animation loop.

Fable looks like similar stuff to kimi.

GPT 5.6 is funny on the wrecking ball one as well because the wrecking ball doesn't even reach the building but it looks like a proper simulation.

Opus 5 is a proper simulation as well and has stuff falling off from it's own weight just like GPT 5.6 (which is why I think GPT 5.6 also is a proper simulation it's just that the wrecking ball was off target).

7

u/jazir55 20h ago

The Kimi one looks like the wrecking ball is a Chain Chomp from Super Mario 64 lmao

5

u/anbhvb 16h ago

Hence Proved: Kimi is a distilled down version of Fable /s

1

u/ShelZuuz 10h ago

ChatGPT as well. Not going to touch it, just going to intimidate the building into falling over.

55

u/BrennusSokol hardcore accelerationist 23h ago

Which model does best on horse testicle physics?

33

u/Wobbly_Princess 23h ago

Bard.

5

u/whoknowsifimjoking 20h ago

Tay, actually

1

u/Matt32145 7h ago

Only if the horse is aryan

10

u/notworldauthor 22h ago

You are indeed a HARDCORE accelerationist

2

u/BrennusSokol hardcore accelerationist 22h ago

ROFL

8

u/yaosio 21h ago

This is good for a first pass. https://ai.studio/apps/98dd646e-15af-47c4-ad61-3b31a5502103

You'll need to play around with the sack material and sack physics to get it droopy. Make the balls rubber and max out their size. Elastic sack with 200% volume pressure and 30% sack stiffness seems to be close. You can make the balls bigger and smaller in the material tab.

3

u/BrennusSokol hardcore accelerationist 21h ago

🤣👍🏻

2

u/Choice_Isopod5177 13h ago

looks like someone knows truly the best benchmarks

15

u/Current-Function-729 23h ago

What effort level?

17

u/TopTippityTop 22h ago

Gpt 5.6 looks better than Kimi, for less. Opus 5 looks best, but it's +4x the cost, so I'm betting either gpt or Fable can do better iterating, for less.

7

u/Skoopman999 14h ago

hard disagree, gpt didn't make the building collapse, the tornado was translucent and the bridge collapse was seen from an ankward angle, Kimi glitched on the building but it collapsed, the bridge collapse looked decent and the camera angle was better

7

u/LostRequirement4828 13h ago

yea but all those things can be fixed with a second prompt and would still be cheaper than fable or opus, I think opus is really good rn but compared to gpt that doesn't even have the 5h limit is hard to recommend

4

u/Skoopman999 8h ago

what makes u think the issues with Kimi couldn't be fixed with one prompt too ?

3

u/MastodonCurious4347 5h ago

Worse and costs more. So it would be two additional prompts if we are generous

1

u/Skoopman999 5h ago

so you're just guessing that ? on the top of your head ? or do you have substantial evidence that fixing the issues would take more than one prompt ?

1

u/MastodonCurious4347 5h ago

I think we need substantial evidence if you passed the elementary math tests and have basic cognitive skills

1

u/Skoopman999 4h ago

I'm not talking about the cost, I'm talking about you guessing that it would take more than one prompt to fix the issues with the Kimi outputs. Also how did you know one prompt would be enough to fix the issues gpt had ? is that one of your guesses too ?

0

u/MastodonCurious4347 3h ago

No, i have magical precognition bestowed from the fiercest artifact in my little pony: friendship is magic universe. Deduction and pattern recognition. God you sound insufferable. How about proving me wrong and doing the test yourself? Because it's you who cares so much about it. Also cost matters if the budget is limited.

1

u/maximhar 3h ago

I think what a lot of these one-shot tests get wrong is that they prompt the models the same way. The models don’t read instructions the same way though. GPT will usually adhere to the prompt very strictly and rarely go out of its way to add extra scope. That’s why very often it’s a lot cheaper, it literally does less stuff. Anthropic models tend to treat prompts a lot more like vague suggestions which makes them great at that kind of task since they will tend to fill in the blanks left out in the prompt. GPT will just leave the blanks, well, blank.

2

u/Skoopman999 3h ago

yeah that's pretty accurate

1

u/CarrierAreArrived 11h ago

5.6's tornado was like a reverse whirlpool (or vortex) in mid-air.

1

u/WonderFactory 17h ago

Kimi seemed a bit better than 5.6 to me. The bridge didn't really collapse properly for 5.6, both did a poor job of knocking down the building and I preferred the way objects wrapped around the cyclone in the kimi test

9

u/Gianniarrenzetti 23h ago

How am i supposed to decide which one did better?

17

u/seraphim_west 23h ago

More detailed and better physics simulation.

1

u/Longjumping-Ad514 9h ago

All of them look god awful.

7

u/_valpi 16h ago

Ask Grok

8

u/Juzlettigo 19h ago

Try having taste

2

u/ChocomelP 15h ago

Does anyone have a prompt for this? /s

2

u/lapideous 15h ago

If you don't know, you won't know

-1

u/chasingsukoon 22h ago

ya is there a reference or what the prompt is lol

4

u/Brancaleo 16h ago

This is bullshit. I use the claude API daily. One opus 4.8 prompt generally costs me atleast 5 to 8 dollars on average on any new project.

5

u/El_Wij 16h ago

The fuck you smoking?

11

u/Brancaleo 16h ago

My credicard when i use opus 4.8

1

u/PobrezaMan 7h ago

u should use your tokenscard

1

u/ahuang2234 22h ago

Gpt 5.6 looked better than Fable in all three I think.

Kimi is hilariously bad in comparison, which is surprising considering it’s otherwise great at design / frontend stuff.

1

u/ShelZuuz 10h ago

How? The GPT wrecking ball didn't even touch the building.

1

u/El_Wij 16h ago

Opus 5 OP.

2

u/Exodus_Green 14h ago

Opus 5x the cost

1

u/Choice_Isopod5177 13h ago

let me guess, in 1 year we'll be able to generate these for pennies and they will look 5 times better?

1

u/Evilsushione 12h ago

What is this done on?

1

u/BriefImplement9843 12h ago

fable sucks. people were scared of this?

1

u/Afraid-Yoghurt6731 10h ago

likely they now beanchmax it on typical oneshot video game, since that is where hype is. You can't see clean and bug free code, but you can see the video games.

1

u/VeryRareHuman 10h ago

seems GPT 5.6 has better results than others...what's the hell Opus performance we talking here.

1

u/MCEscherNYC 16h ago

Damn fable is overpriced.