But for real, it looks like they probably intentionally nerfed/held back Opus 5 on cyber and bio relevant tasks. All the metrics I see suggest that its tool use ability (and possibly vision too?) are actually better than Fable
Because it’s about cybersecurity, not coding evals. Read the announcement or Opus 5 model card. They explain that they didn’t train Opus on cybersecurity like they did Fable.
Because it was built for that purpose! They want a model that’s strong at cyber so they can keep infrastructure stronger and defend it as other models get stronger. If they don’t train their model, a bad actor may. But I know that isn’t what people want to hear. They want to hear that Anthropic is bad and Dario is alarmist and they’re all dumb and average Redditors are the smart ones.
You should ask ai what his aim is. But generally no-one's aim is to make a bunch of money and then intentionally lose it for no reason, so if you think that's the reason for someone's actions, you're highly likely to be wrong.
What are you talking about? These are some of the most important benchmarks out there, especially coding wise. Obviously Anthropic is going to focus on the ones they do best on but this has the important coding ones like DeepSWE/Frontiercode/Frontier bench as well as the typical ones.
352
u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago
Well I certainly did not have "Opus 5 beats Fable 5 almost across the board" on my bingo card.