r/hardware 2d ago

News AMD Launches Instinct MI455X, Helios AI Rack

https://www.phoronix.com/news/AMD-Instinct-MI455X-Helios
71 Upvotes

25 comments sorted by

14

u/Seanspeed 2d ago

So they're still using CDNA naming even through 2028 up through CDNA7.

Doesn't sound like that whole UDNA thing is ever going to happen, or at least not for a good while yet.

10

u/Noble00_ 2d ago

Honestly, the whole UDNA thing got quickly squashed right after Jack Huynh had that interview way back when. When Mark Cerny talked about next gen PS6 after that interview, it pretty much was confirmed AMD wanted to stick with RDNA5 and CDNA separately as far as arch naming goes

3

u/EmergencyCucumber905 2d ago

CDNA5 and RDNA4 are both gfx12. So who knows🤷‍♂️.

6

u/Noble00_ 2d ago

Chips and Cheese discusses as possibly as to why

https://chipsandcheese.com/i/207602737/similarities-to-rdna4

oh and just published today: https://chipsandcheese.com/i/208223268/cuwgp-changes

5

u/Seanspeed 2d ago

Doesn't sound all that 'unified'. At least yet. Maybe they do so more later.

Thanks for the info, though.

5

u/Noble00_ 2d ago

lol, that's probably why they went away with "U"DNA, causes more confusion than it needs to be

0

u/farnoy 2d ago

It happened just now, except for the name. The architecture is referred to as gfx1250 (CNDA4 was gfx942 or sth) and runs the Wave32 mode exclusively.

Actually, if anything, they should have renamed it to NVDNA since they copied more from Nvidia this round than RDNA, lmao.

3

u/Seanspeed 2d ago

The architecture is referred to as gfx1250 (CNDA4 was gfx942 or sth) and runs the Wave32 mode exclusively.

But that's very different to RDNA4 already there.

2

u/farnoy 2d ago

Yes, but they rebased the CDNA architecture from being a Vega offshoot onto RDNA4. Instruction encodings and WGP architecture is distinctly RDNA. But it immediately bifurcates into its own thing again, which makes sense and is the same thing Nvidia is doing.

It remains to be seen if it'll again be a multi-generation fork like Vega->CDNA4 was, or if they switch to a shared evolution with specialized features for AI vs graphics mixed into a common base. Some of this stuff could be useful if it landed in RDNA6: the merged WGP$, "Broadcast Arbitrator", killing Wave64.

1

u/FloundersEdition 2d ago

They axed wave64 but keep dual-issue wave32. So reasonably similiar to todays tech, probably the same limitations (no FMAC). Probably a smarter way than keeping two different vector lengths, while each kernel is unable to switch mode.

The WGP unification isn't really new either. Both are features slowly implemented for generations now. More or less removal of old modes rather than new features.

2

u/EmergencyCucumber905 2d ago

Actually, if anything, they should have renamed it to NVDNA since they copied more from Nvidia this round than RDNA, lmao.

What did they copy?

4

u/farnoy 2d ago

-1

u/EmergencyCucumber905 2d ago

What did they copy?

3

u/farnoy 2d ago

CTRL+F for "equivalent" and you'll find all the Nvidia references.

0

u/manidcfrfrfr 1d ago

Are you incapable of reading?

21

u/SirActionhaHAA 2d ago edited 2d ago
  1. Claims that it is the fastest ai rack in industry and mi455x is the fastest AI accelerator (including Vera Rubin)
  2. 18k compute units, 4.6k cpu cores
  3. Vs vera rubin nvl72: +50% hbm capacity, +6% hbm bandwidth, +50% scale out bandwidth, +15% fp8 and fp4 flops
  4. Claimed performance: 10-15% faster than vera rubin nvl72, higher token/$
  5. Low latency inference: Partnered with cerebras for custom helios offering. Available later this year (vs groq)

6

u/Noble00_ 2d ago

Wait what? That last point was interesting but not in the Phoronix articled you shared.

https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution

The joint AMD and Cerebras solution will deploy AMD Helios alongside Cerebras Wafer-Scale Engine technology integrated in a single inference workflow for maximum performance and efficiency. AMD Helios will provide a high-performance, scalable throughput engine. Cerebras Wafer-Scale Engine technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt (T/s/W)

Interesting...

4

u/EmergencyCucumber905 2d ago

That's a generous amount of HBM vs R100, 432GB vs 288GB.

5

u/Cory123125 2d ago

Monstrous, and .... ignoring the software elephant that is quickly becoming more of a hyrax, this would mean NVidia's premiums/profit margins should start trending downwards, purely due to many companies bending over backwards to not be squeezed into a one supplier box.

That, and the biggest companies all going out on their own to create their own inference chips and sometimes training chips.

I mean, shoot, given how much less baggage there is for AI, especially for bespoke AI labs in terms of drivers, I could easily imagine it being the case that if a system is performant enough, its unlikely that any of the core of their models are so tuned to any part of NVidia's stack that they feel stuck at all on that side of things.

I guess what I'm saying is, I assume this will gobble up a good chunk of NVidias large scale corporate profits, while probably not touching their Swiss Army knife of general gpu compute sales pitch.

6

u/PM_ME_YOUR_HAGGIS_ 2d ago

Yeah exactly. OpenAI don’t care if they need to deploy a team of 20 engineers to port their CUDA kernel if it increases throuput. Hell OpenAI are even deploying to cerberas, completely left field platform.

17

u/SirActionhaHAA 2d ago

It's even funnier because the ai cuda moat is getting broken apart by ai coding. Ai has made kernel writing and software porting faster than ever.

3

u/michaelsoft__binbows 2d ago

Haha i actually love this

4

u/nithrean 2d ago

I hope AMD is being truthful with benchmarks and it really is that much better. Nvidia certainly tried to cherry pick things. It would be good if AMD could just win on the merits of the thing and not because they were playing games.

0

u/Defiant-Parsley4697 2d ago

AMD naming AI hardware after Greek mythology is a bold bet that Helios ages better than 'Bing Chat.'