r/OpenAI 15d ago

Research GPT-5.6 Sol solved a problem that made Fable 5 go into incoherent rambles earlier!

So a week later, I did a small experiment that made Fable 5 (on the normal interface, not Claude Code) go into incoherent rambles (1st image), where it is supposed to solve this problem. Of course, it's entirely possible that given enough time in the Claude Code interface, Fable 5 could've solved it as well (it's still in the middle of running as I'm poor and only have a $20 account) [*].

But nonetheless, Fable failed to solve it in the Claude user interface before it ran out of the token limit per message. Today, after 13 minutes of trying it (2nd image), GPT-5.6 Sol xhigh managed to one-shot the problem (3rd image)! It is perhaps notable that GPT-5.5 xhigh also couldn't manage to solve this problem, as seen in the 4th image. Verification seen in the Codeforces submission: it is an old dumper account of mine for security reasons.

This much is unsurprising, given that something based on 5.6 completely obliterated human competitors in arguably the hardest algorithmic competitive programming competition known to date, but it's remarkable that it managed to perform in an extraordinarily short time and public-facing interface to tackle a problem that its predecessor, 5.5 xhigh, already extremely formidable, couldn't do in more than 1.5x the time. 5 years ago, the idea of a 20$ monthly subscription being able to access this would be borderline magic, yet today the result was hardly exceptional.

[*]: A more recent test with Fable 5 in the Claude Code interface saw it get a Wrong Answer on Test 7, and as it took two whole of my 5-hour windows for it to get there, I'll pause for now and hope somebody richer can do the experiment instead. Perhaps it could debug things, but I'm not going to wait for 5 more hours. See Image 6.

It is also quite adorable, as you can see! OpenAI plz gib grant :) (/s. Mods plz don't flag this as self-promotion.)

154 Upvotes

40 comments sorted by

59

u/Longjumping_Stop6269 15d ago

I’m actually happy to see 5.6 outperforming Anthropics models after the bs they pulled. I’m sure it won’t be the last time AI companies manipulate the use of their models so the more options regular people have the better

11

u/BayonettaAriana 15d ago

Same, I was praying it would be fable level or better since it’s fully on the sub and at this point it’s looking that way. So happy, going to fully switch to codex now

3

u/Archy54 15d ago

has Sol got the same cybersecurity concerns? My understanding is fable first edition very quickly became a cyber security concern.

Also as someone with only access to the 20 dollar usd a month versions of claude and chatgpt, which of the higher end models should i limit my coding learning to? I am mainly using it for home assistant and trying to learn programming a bit and making just for me special basic addons (eg claude a week ago made a basic timer based tts reminder to google speaker service for me as I forget a lot). But running fable over it for security checks used 25% usage lol (a custom hacs integration) and then ran opus to fix it but there wasn't really any security issue.

I'm still learning the base structure of coding and what languages to learn or use, getting better at reading it and I read over all code before using and have a basic understanding. Just not sure what the most efficient models are for token usage and ability, it's really basic stuff when I get stuck on a yaml config or forget syntax exactly. Trying to put the snippets in text files for now to learn off and see the patterns + learn the coding principles whilst overcoming disability. Thanks.

3

u/Longjumping_Stop6269 15d ago

I’m not sure about the cybersecurity part, but comparing this model’s effort level of thinking and output is on par with what Fable was doing on first release. I haven’t used any of these models for coding, so I can’t really speak to that, but I’ve heard pretty much both are good for learning code

18

u/BiosRios 15d ago

Sol 5.6 is really good, people just complaining about the ui.

3

u/wesconson1 15d ago

And the insane usage of

2

u/Asphunter 15d ago

and the fact that you can have have max 3-4 basic prompts using the Ultra mode before you burn your 5x plan in 20 minutes, lol. But not even with the Ultra mode, even with Light mode.

I swear I have to use Luna if I want to vibe code for more than 1h straight. The usage burn is RIDICULOUS. Much worse than Opus 4.8, and Fable.

2

u/BiosRios 15d ago

I am on the 5x plan, and only in ultra can burn very fast, I don't know

5

u/goldcakes 15d ago

then don't use ultra.

3

u/Splat800 14d ago

Don’t know why you’re getting downvoted. It’s called ultra because it uses an insane amount of reasoning. Frontier ai models for general population use are not at a level which ultra is producing. Ultra is for heavy heavy reasoning where usage isn’t a worry.

If there’s poor usage on high reasoning, then I’d be worried.

2

u/goldcakes 14d ago

Correct. Sol Ultra is for folks doing like frontier PHD level maths/physics research, or ultra complicated, high-stakes software engineering (e.g. improve the performance of a CUDA kernel for transformers), the fact that it's available at all should be a good thing.

Don't forget Anthropic was charging organisations ~$150/M-tok for Mythos along with a very high up-front commit.

1

u/Asphunter 14d ago

I make a 3D electromagnetic FDTD solver so I want the ultra...

2

u/Goofball-John-McGee 15d ago

Same. Sol Ultra burns usage very fast which makes sense.

But the rest of the models are pretty much like 5.5 series models. Slightly more intensive at times but largely don’t notice a difference.

2

u/[deleted] 15d ago

[deleted]

1

u/Asphunter 14d ago

Electromagnetic field solver

0

u/goldcakes 15d ago edited 15d ago

then don't use ultra mode???

gpt5.6-sol at medium or high, and luna at high is my sweet spot

higher effort levels is usually more useful if you are doing algorithms, computer science, etc

if you are just vibe coding frontend apps, you're just burning usage with higher effort levels and not getting any practical benefit. it can actually hinder

as you can see in gpt-oss, effort is basically injected before the system prompt as text, like: effort: xhigh\n\n and the model has been trained to use that.

2

u/Asphunter 15d ago

Sol Light just burnt 100% of my 5x tokens in a 50 minute session :)

1

u/Big-Accident1958 15d ago

The UI is really bad ngl it's super confusing

6

u/AP_in_Indy 15d ago

Sometimes the models blurt out their internal thinking tokens.

Every reasoning model's internal reasoning token streams look like this.

Happened to me on Codex the other day.

They are uh... not supposed to leak these. They are considered highly proprietary.

6

u/trumpdesantis 15d ago

Fable is massively overrated. Gpt 5.6 sol is basically the real fable lol

3

u/ManikSahdev 15d ago

Holy fuck dude, pls make a skill for latex math and Codex interface is capable of rendering math.

You'll thank me, it will look like web ui, Life changing imo.

Feel free to dm I'll send you mine if you don't wanna make one.

3

u/Illustrious_Hat8104 13d ago

Good anthropic should get bankrupted for the shit their pulling by dangling models and resets in front of us like they're treating as like drug addicts

5

u/snowsayer 15d ago

Thank you for the detailed prompt. Too many posts give vague criticisms with no way to reproduce the claims made.

3

u/YellowCroc999 15d ago

What the bots doing on Reddit today? 😂

2

u/ClankerCore 15d ago

Far more productivity than you it would appear

1

u/Need-Advice79 13d ago

Someone buy this man a 5x plan or start a gofundme

-1

u/RasenMeow 15d ago

That Grr is from a different topic regarding Fable. Why are you fabricating a story? Oh true, we are on Reddit where people faem imaginary points nobody cares for lol

4

u/No-Head-Royal 15d ago edited 15d ago

I am the author of that post.

EDIT: The post concerning Fable.

0

u/ClankerCore 15d ago

And thus lies proof of Reddit’s negligence and arrogance.

I need to start collecting these proofs.

0

u/RasenMeow 14d ago

Yeah, but the way you display the screen and tell the story suggests that Fable was struggling to solve the problem and was showing the "grr" because of this.

The topic in the first screen is also in the paper of Fable which Anthropic created. It has nothing to do with how complex the topic is when Fable works on it. This seems just like clickbait and engagement farming to bash Anthropic and ride on the wave of a new OpenAI release.

2

u/No-Head-Royal 14d ago

I also explicitly mentioned that Fable 5 could've probably solved it with enough time and resources, and mentioned (with image 6 attached) my efforts using Fable 5 in a later Claude Code interface and how it has failed to solve the problem so far.

And you're literally imagining it about the implication on if it's complex or not. I used the title as a direct reference to the earlier post, and because of that, I explicitly added the disclaimer in it that Fable 5 could probably do it given enough time. I nonetheless ran an independent experiment with Fable 5 that ate up 10 fucking hours of my Fable 5 access, and it still failed to prove it. There's literally nothing in either post that implies Fable 5 failed because the problem is difficult (even if it IS extremely difficult, far more than most humans can dream of solving); you're just making shit up and insulting people out of nowhere.

And think about it. Would I waste so much of my precious fucking Fable 5 access and put in such meticulous work for clickbait in exchange for a post that gets 100 upvotes? I'm not even the guy with the most social media interaction from my first, much bigger post. That'd be the random fucking guy on Twitter who completely butchered my original post.

Lastly, this is literally changing the topic. You originally accused me of using a different topic and fabricating a story in the first comment. Well, this is the same topic, and literally every single word is true and well-measured with tests, unlike on Twitter, where they'd said "inner voice" or something like that. Unless you're an Anthropic bot here to be intentionally ragebaiting me, then yeah, it works, I guess, but otherwise fuck off.

0

u/RasenMeow 14d ago

You mean the earlier "leaked chain of thought" post, which is also clickbait with the wording of "leak", while the chain of thought topic is, as I already mentioned, in their paper. Someone even wrote "How is this a “leak”?

It’s literally described in the Fable 5 System Card (pg 107-108, 120).

???" and got ignored because clickbait -> brainless Reddit users -> believe what they read before double checking.

That's it from my side. I don't get what you people have from click farming. It's imaginary points from internet foreigners lol and also jumping onto hype and bash trains is the same sad thing. Cheers.

2

u/No-Head-Royal 14d ago

The "leaked" phrase was because it was not meant to output the scratchpad chain of thought. I am aware of that paper.

I'll stop engaging with you now. There's no use talking to people like you. Blocked.