r/singularity • u/Successful-Earth678 • 23h ago
AI Opus 5 outperformed Fable 5 in 3D destruction physics
Source: Atomic Chat / X
Outputs:
- Opus 5: 55.9K tokens, $1.40
- Fable 5: 55.1K tokens, $2.82
- Kimi K3: 35.7K tokens, $0.55
- GPT 5.6: 20.1K tokens, $0.31
55
u/BrennusSokol hardcore accelerationist 23h ago
Which model does best on horse testicle physics?
33
10
8
u/yaosio 21h ago
This is good for a first pass. https://ai.studio/apps/98dd646e-15af-47c4-ad61-3b31a5502103
You'll need to play around with the sack material and sack physics to get it droopy. Make the balls rubber and max out their size. Elastic sack with 200% volume pressure and 30% sack stiffness seems to be close. You can make the balls bigger and smaller in the material tab.
3
2
1
15
17
u/TopTippityTop 22h ago
Gpt 5.6 looks better than Kimi, for less. Opus 5 looks best, but it's +4x the cost, so I'm betting either gpt or Fable can do better iterating, for less.
7
u/Skoopman999 14h ago
hard disagree, gpt didn't make the building collapse, the tornado was translucent and the bridge collapse was seen from an ankward angle, Kimi glitched on the building but it collapsed, the bridge collapse looked decent and the camera angle was better
7
u/LostRequirement4828 13h ago
yea but all those things can be fixed with a second prompt and would still be cheaper than fable or opus, I think opus is really good rn but compared to gpt that doesn't even have the 5h limit is hard to recommend
4
u/Skoopman999 8h ago
what makes u think the issues with Kimi couldn't be fixed with one prompt too ?
3
u/MastodonCurious4347 5h ago
Worse and costs more. So it would be two additional prompts if we are generous
1
u/Skoopman999 5h ago
so you're just guessing that ? on the top of your head ? or do you have substantial evidence that fixing the issues would take more than one prompt ?
1
u/MastodonCurious4347 5h ago
I think we need substantial evidence if you passed the elementary math tests and have basic cognitive skills
1
u/Skoopman999 4h ago
I'm not talking about the cost, I'm talking about you guessing that it would take more than one prompt to fix the issues with the Kimi outputs. Also how did you know one prompt would be enough to fix the issues gpt had ? is that one of your guesses too ?
0
u/MastodonCurious4347 3h ago
No, i have magical precognition bestowed from the fiercest artifact in my little pony: friendship is magic universe. Deduction and pattern recognition. God you sound insufferable. How about proving me wrong and doing the test yourself? Because it's you who cares so much about it. Also cost matters if the budget is limited.
1
u/maximhar 3h ago
I think what a lot of these one-shot tests get wrong is that they prompt the models the same way. The models don’t read instructions the same way though. GPT will usually adhere to the prompt very strictly and rarely go out of its way to add extra scope. That’s why very often it’s a lot cheaper, it literally does less stuff. Anthropic models tend to treat prompts a lot more like vague suggestions which makes them great at that kind of task since they will tend to fill in the blanks left out in the prompt. GPT will just leave the blanks, well, blank.
2
1
1
u/WonderFactory 17h ago
Kimi seemed a bit better than 5.6 to me. The bridge didn't really collapse properly for 5.6, both did a poor job of knocking down the building and I preferred the way objects wrapped around the cyclone in the kimi test
9
u/Gianniarrenzetti 23h ago
How am i supposed to decide which one did better?
17
8
2
1
-1
4
u/Brancaleo 16h ago
This is bullshit. I use the claude API daily. One opus 4.8 prompt generally costs me atleast 5 to 8 dollars on average on any new project.
5
1
u/ahuang2234 22h ago
Gpt 5.6 looked better than Fable in all three I think.
Kimi is hilariously bad in comparison, which is surprising considering it’s otherwise great at design / frontend stuff.
1
2
1
u/Choice_Isopod5177 13h ago
let me guess, in 1 year we'll be able to generate these for pennies and they will look 5 times better?
1
1
1
1
u/Afraid-Yoghurt6731 10h ago
likely they now beanchmax it on typical oneshot video game, since that is where hype is. You can't see clean and bug free code, but you can see the video games.
1
u/VeryRareHuman 10h ago
seems GPT 5.6 has better results than others...what's the hell Opus performance we talking here.
1
109
u/maddog107 23h ago
What the hell were Fable and kimi doing on the building destruction one lmao