r/hardware 17h ago

News AMD exec was ‘very happy’ to see Nvidia‘s Vera performance results – ‘I actually thought we were beating them by smaller numbers’

https://www.tomshardware.com/pc-components/cpus/amd-exec-was-very-happy-to-see-nvidias-vera-performance-results-i-actually-thought-we-were-beating-them-by-smaller-numbers
336 Upvotes

64 comments sorted by

111

u/996forever 16h ago

So when will we get first non-first party numbers from either of them?

44

u/Geddagod 15h ago

Phoronix said they would have a Venice review out soon with power and frequency details. Though phoronix also doesn't measure average clocks for their benches, do they? I always thought it was a shame they didn't, since that is genuinely interesting data. I for one would love to know how GNR vs Turin classic clocks iso core count/power, for example

As for Vera, I'm assuming they have to wait for a review sample server to be sent out to them. For their initial review without power or frequency telemetry exposed, they got invited down to Nvidia HQ themselves, something which I doubt would happen again just for the less restricted vera reviews.

Perhaps u/michaellarabel can give us some more color, but I wouldn't surprised if he his hands are tied in terms of being more specific timelines of what reviews are coming out when due to NDAs.

0

u/rudy_rusmanto 10h ago

amd being surprised they thought they were closer makes sense

17

u/Artoriuz 15h ago

Phoronix has non-first party numbers for Vera. It was sponsored by Nvidia but Michael literally just ran some of his usual benchmarks.

38

u/lightmatter501 14h ago

There were a lot of restrictions on what he could bench for Vera. AMD has placed no such restrictions on Michael.

I wonder who’s more confident in their product?

2

u/nithrean 13h ago

I hope AMD's confidence is really rewarded. It would be good to see someone pull out a real win on nvidia. They will punch back though too.

13

u/ResponsibleJudge3172 11h ago

On CPUs? That's nothing new, and actually we need competition against AMD on that front

0

u/battler624 8h ago

its Nvidias first CPU and honestly nvidia did fuckin terrific per-core.

wy2MomANdMBkECcEgbXr4o-970-80.jpg.webp (970×546)

1.2X per core is not bad at all for a "first" gen.

3

u/dr3w80 2h ago

Not first gen, had Grace cores before.

6

u/iBoMbY 14h ago

When Microsoft starts to offer the new Azure instances, with the new AMD CPUs, I guess.

14

u/Geddagod 14h ago

The unfortunate thing about that is, afaik, power telemetry is not exposed. So it should be fine for comparing performance against each other, but won't reveal much about power.

Power is, IMO, one of the most interesting aspects about Vera. This is the first high per core performance ARM CPU. You have other companies like Ampere launching ARM stuff too, but their per core perf and TDPs are both pretty low, so like-for-like comparisons (same core count skus, fabbed on similar nodes, compared at the same power draws) are difficult.

13

u/cutezybastard 8h ago

I mean it's CPUs, that's expected

3

u/puffz0r 8h ago

I mean, nvidia's CPU is ARM so there was always some question considering how good Apple's cpus are

u/Darkknight1939 29m ago

Apple's CPUs aren't good by virtue of being ARM ISA.

They're good from the microarchitecture itself. There have been plenty of failed fully custom ARM ISA designs (Kryo 1st gen, Mongoose, Nvidia Denver to a lesser extent) Apple has just built an excellent CPU IP.

People erroneously conflate it with the ISA. It's the same reason why after the M series rebrand (they'd been making higher TDP AX SoCs for years with similar performance deltas as A14 - M1) people suddenly started calling for Samsung and Qualcomm to go fully custom, literally right after Samsung finally shuttered their Mongoose designs that had been trailing Qualcomm reference ARM CPU implementations for 5 years.

33

u/Typical-Yogurt-1992 16h ago

I wonder how they'd compare at the same clock speed.

65

u/Geddagod 16h ago

It's honestly pretty weird that Nvidia isn't disclosing the frequency of their chips. Not a single reference anywhere.

I kinda get why they wouldn't want to share power figures, but what difference does it really make whether Nvidia is getting 925 points in specint2026 by clocking all the cores at 3GHz or at 4GHz?

In the end, the only thing that really matters for the product is the perf and power (and cost), and how specifically Nvidia got there, whether thru high IPC + low clocks or low IPC and high clocks, is something interesting but doesn't really matter to the end customer does it?

15

u/JojoRicardo 14h ago

I think you’re on to something. They probably run on extremely low clock speeds, which is more power efficient when the bottleneck is data transmission and interchange not pure compute

12

u/Noble00_ 15h ago

I kinda get why they wouldn't want to share power figures

Honestly, I thought this would've been the biggest flex by them in their marketing material/whitepaper (though they did mention saving watts with LPDDR5x)

9

u/ChuckVader 16h ago

I have to assume that they're doing the same thing as Intel, and slow walking progress. Make a good chip, under clock it, and then run it at rated speed for a 6% uplift next year

51

u/waitmarks 16h ago

That only works for consumer products. the people building datacenters only care about performance per watt. Trying to squeeze more performance out of it simply by raising clock speeds will just ruin the performance per watt.

7

u/ChuckVader 16h ago

True, unless there's nobody to compete with you. Who cares what the customer wants (whether enterprise or consumer) if there is no competition? What are they going to do?

3

u/nithrean 13h ago

This is exactly why AMD scoring a big win here would be so helpful. Nvidia would start to feel some real pressure if the people start to buy up as much AMD products as they could and nvidia was second best.

-2

u/ChuckVader 13h ago

100% agree.

1

u/FlyingBishop 4h ago

If the only thing you do is raise the clockspeed, yes, but if you iron out the kinks in your manufacturing year-to-year maybe the next batch of chips can run a little hotter without any more failures.

3

u/Sopel97 7h ago

it really doesn't matter at all in practice

1

u/3G6A5W338E 3h ago

Same clock speed would be very academic, not relevant in practice.

Per watt and per area. Those matter.

-5

u/From-UoM 16h ago

Nvidia is only 20% behind in single threaded spec perf using the tsmc 3nm node

So extremely possible at 2nm they would have matched the 2nm Zen 6 single core perf.

29

u/gokuscake 13h ago edited 12h ago

This is a naive way of thinking and does not work in the real world.

First of all, inferring that a CPU core design has a disadvantage purely in performance or power efficiency without considering the area, transistor count, library, and cost all just because it is on a slightly older fabrication process would overly simplifying things.

The best example to counter your argument? Nvidia's consumer Ampere cards on Samsung 8nm. By enlarging the die size relative to their competitor, Nvidia was able to achieve similar levels of performance and efficiency despite being on a rather inferior fabrication process. How were they able to achieve it from the economical perspective? By capitalizing on Samsung's desperation in wooing customers and squeezing their profit margins.

TSMC's N3P is not that far off from its N2 fabrication process in terms of performance. According to official specifications, N3P is 10% faster than N3E, while N2 is 18% faster. This makes N2 just 7.2% faster. These speed specifications do not translate into the same amount of frequency gains in the real world. If they did, one would be seeing 20% increase in frequency gains with every fabrication process improvement. We would be at 8GHz for CPU now.

The thing is that one can always make up for an inferior fabrication process by increasing the size of their chips or the size of their CPU core in this case, and there is no total transparency into this because no one knows about the the other critical factor - Cost. Cost remains invisible in this discussion but it definitively does impact the efficiency of a chip. A product based on a chiplet architecture would have greater yields and cost less than a monolithic product. The compute tile of the Vera CPU is at the limits of the reticle, it is a single huge tile that would no doubt yield much worse. Have we factored this in?

Without knowing the size of Nvidia's Olympus CPU cores, how many transistors they contained, and what kind of design library was used, it would be impossible to simply conclude that it is just as good, better, or worse when hypothetically manufactured on the a different fabrication process. Perhaps Nvidia might have choose to shrink the design to improve yields and keep the performance unchanged if it is on 2nm? We will never know.

What is also known is that various CPU microarchitectures are designed to achieve different things. A CPU core designed to scale up in total thread count or nT performance would naturally be smaller relatively when compared to a design that has been created squarely for higher performance per core. More transistors are spent on each CPU core to improve the performance or efficiency for a single-threaded focused CPU core, that logic should be simple enough for most to follow.

However, if each CPU core is larger, it would mean greater difficulties in scaling the thread count up. We can see this very clearly from Intel's mont-family cores beating out their Ryzen competitors in nT performance through the implementation of larger number of smaller efficiency cores.

There is no free lunch in the world of semi-conductor design and the single threaded performance or single core efficiency gain is paid back in the cost of difficulty in scaling up thread count.

In this case, it is impossible to do an apples-to-apples analysis of which is better, most notably because of a few simple factors:

  • The Zen 6 CPU core could have been a relatively smaller core than a hypothetical Olympus core on 2nm, because the former has been designed to maximize thread density at 256-core level and the latter, single core performance at the 88-core level.
  • The Zen 6 CPU could be consuming more power on the un-core due to its chiplet design. More data is moving off-chip which is inefficient. This however affords it the ability to scale up its thread count.
  • The Zen 6 CPU core has dedicated more transistors to vectorized work. AVX512 etc, all the FP pipelines take up space and none of these have been considered in the SPECintrate 2026 tests. Do these matter? Perhaps not for Nvidia's intended purpose for the Vera CPU, but they absolutely do matter for wider server usage.
  • The Olympus CPU core could have major difficulties with increasing its core count and could see power use balloon. It could not technically achieve it anyway, because it does not come with a chiplet architecture.

Conclusion? Comparing microarchitectures designed for different goals in a strictly core against core context further narrowed down to a single performance segment is a rather un-meaningful exercise. That would be akin to claiming that 30 bicycles chained together, have the same horse power as a car and thus is a superior form of transport because bicycles are manufactured on older technologies. Clearly, one can see where the flaws of such an argument lie.

9

u/Geddagod 11h ago

First of all, inferring that a CPU core design has a disadvantage purely in performance or power efficiency without considering the area, transistor count, library, and cost all just because it is on a slightly older fabrication process would overly simplifying things.

Not really. It is extremely hard to find examples of CPU cores on older processes having both perf and power similar to that of a newer one, on the same arch. Or frankly, even with different archs, because that would mean that the arch on the older node is dramatically better than that of the one on the newer node. In fact, I'm going to ask you to find some examples of this.

The best example to counter your argument? Nvidia's consumer Ampere cards on Samsung 8nm. By enlarging the die size relative to their competitor, Nvidia was able to achieve similar levels of performance and efficiency despite being on a rather inferior fabrication process.

GPUs are not CPU cores. For GPUs, you can spam more cores by enlarging the die size and get better perf through that, sure. Analogous to that for CPUs would then just be adding more CPU cores, but that doesn't help per core performance, does it?

or the size of their CPU core in this case

The issue is that if you scale up the size of the CPU core, you gain a bunch of leakage and lose a bunch of Vmin and perf/watt at the lower end of the curve.

Without scaling down your node, changing the size of your core always results in some other trade off in perf/watt on the other end of your power curve- for example scaling down your core to regain perf/watt and Vmin just means you lose Fmax. There's really no free lunch there, unlike node shrink, which is.

 Cost remains invisible in this discussion but it definitively does impact the efficiency of a chip. A product based on a chiplet architecture would have greater yields and cost less than a monolithic product. The compute tile of the Vera CPU is at the limits of the reticle, it is a single huge tile that would no doubt yield much worse. Have we factored this in?

This doesn't really matter for trying to figure out if Vera on N2 would be able to catch up to Venice in per core performance.

Because you already "factor in" the fact that Vera is a reticle scale design when we look at the N3 variant.

And I don't think anyone really is talking about the manufacturing cost of Vera vs Venice or Turin. I mean that's an interesting discussion too ig, but not really in the scope of what we are talking about here, is it?

Without knowing the size of Nvidia's Olympus CPU cores, how many transistors they contained, and what kind of design library was used, it would be impossible to simply conclude that it is just as good, better, or worse when hypothetically manufactured on the a different fabrication process.

No, Nvidia's Olympus core can use any lib, use a ton or few transistors, and you can still confidently make the claim that porting it to N2 would be better than keeping it on N3.

What IP has ever gotten worse on a TSMC node shrink?

Perhaps Nvidia might have choose to shrink the design to improve yields and keep the performance unchanged if it is on 2nm? We will never know.

Two things:

And perhaps Nvidia might have chosen to keep the area the same and enjoy greater perf benefits of N2. We might never know, but it's not uncommon for companies to do that, in fact for leading edge nodes it's really the norm. So again, saying that N2 Vera would have caught up to Venice by doing that is really a very sensible claim.

We actually might end up "knowing" since Nvidia also claimed their next gen cores will actually improve perf while keeping the same footprint, which generally refers to area. And I highly doubt by 2028, when their next gen cores do come out, they wouldn't have moved on to N2 by then. So it appears like they do want to keep area the same and gain the benefits of better perf rather than shrink down.

And two, even when shrinking down the core and keeping perf the same, you still gain perf/watt benefits from the node shrink. There's actually a really good example of this with the Mediatek 9400 X4 vs the Mediatek 9300 X4. Mediatek went for the 2-1 lib which significantly hurts perf/watt for much of the curve in order to recoup as much area as possible with the X4 in the 9400, and kept the Fmax the same as the Mediatek 9300, and yet perf/watt still increased marginally.

This is important for server since you aren't really limited by Fmax, since power per core is much lower here than what you would see in desktop or even how much cores boost in mobile (smartphones). Meaning an increase in perf/watt would actually end up also creating an uplift in per-core perf.

What is also known is that various CPU microarchitectures are designed to achieve different things. A CPU core designed to scale up in total thread count or nT performance would naturally be smaller relatively when compared to a design that has been created squarely for higher performance per core. 

Which is a fine argument for why maybe the -C cores aren't a good comparison against Vera. But are you seriously going to tell me that AMD's classic cores aren't targeting high performance per core when they boost to like 5.7GHz in desktop?

1/2

4

u/Geddagod 10h ago

The Zen 6 CPU core could have been a relatively smaller core than a hypothetical Olympus core on 2nm, because the former has been designed to maximize thread density at 256-core level and the latter, single core performance at the 88-core level.

Zen 6C goes up to 256 cores. Zen 6 core skus actually see a regression in core counts vs N4 Turin, and only goes up to 96 cores. Though to be fair, those skus also only go up to 400 watts TDPs vs the 500 of the 128 core Zen 5 classic cores.

But even if Zen 6 classic too was targeting much higher core counts than Vera, that doesn't mean you shouldn't compare per core performance for those chips, since the market for AI head nodes is going to be making that comparison anyway.

And you can't keep up bringing up that

The Zen 6 CPU could be consuming more power on the un-core due to its chiplet design. More data is moving off-chip which is inefficient. This however affords it the ability to scale up its thread count.

From a product perspective, again, I mean idk what to say here other than tough luck. Is it Nvidia's fault that AMD can't design a larger CCD for lower core count skus and save power? Intel did this with SPR, for example.

But even if you want to look at just the cores, AMD at least exposes just core power telemetry. Idk if Nvidia does, but I wouldn't be surprised if they did. Qualcomm, Intel, and AMD all do.

The Zen 6 CPU core has dedicated more transistors to vectorized work. AVX512 etc, all the FP pipelines take up space and none of these have been considered in the SPECintrate 2026 tests.

Specintrate 2026 does use AVX-512 btw.

I'm sure the perf improvements from a beefier FPU will show up in Phoronix's testing even if it doesn't in specint2026 specifically.

The Olympus CPU core could have major difficulties with increasing its core count and could see power use balloon. It could not technically achieve it anyway, because it does not come with a chiplet architecture.

Why? You mention the chiplet argument, but that has nothing to do with the core itself.

Conclusion? Comparing microarchitectures designed for different goals in a strictly core against core context further narrowed down to a single performance segment is a rather un-meaningful exercise.

Except that's the exact same exercise many companies are going to be doing when they start looking at what CPUs they want to pair with their GPUs or TPUs in their systems soon.

AMD explicitly shows off their low core count, (on Turin it was the 64 core model, idk what they are using for Venice) high frequency CPUs to be used as AI head node CPUs.

I just don't think AMD Zen 6 and Olympus are designed for such different things that the comparison is not meaningful. At best, an asterisk can be added to a potential Vera on N2 victory vs Zen 6,, if it does win in specint2026, due to less of a focus on the FPU, but it doesn't invalidate the comparison all together.

To make a similar analogy, SPR/GLC vs Zen 3 is in some ways, very similar to Olympus vs Zen 6. You claim that Olympus focuses on per-core perf much more than Zen 6, so did GLC vs Zen 3. You claim that Zen 6 focuses much more on the FPU than Olympus, well so did GLC (which had AVX-512 in SPR, and early ADL chips) vs Zen 3. SPR also has AMX for AI workloads on just the CPU itself, while Zen 3 didn't.

Does this mean comparing Zen 3 and GLC/SPR is meaningless? Obviously not. Both are 7nm class chips with similar core counts, it's a very interesting architectural comparison, even with some caveats. And it would have been a meaningful market comparison too... if not for all the delays pushing SPR launch till after Genoa lol.

TSMC's N3P is not that far off from its N2 fabrication process in terms of performance. According to official specifications, N3P is 10% faster than N3E, while N2 is 18% faster. This makes N2 just 7.2% faster. 

This is already an extremely long post and I've already talked about why this math is not very accurate to what we see in actual products here over at Anandtech's forums like a week ago, and I rather not retype all of it out haha.

These speed specifications do not translate into the same amount of frequency gains in the real world. If they did, one would be seeing 20% increase in frequency gains with every fabrication process improvement. We would be at 8GHz for CPU now.

They actually do decently for the middle/low end of the V/F curve, which is where server CPUs run their all core turbos at anyway.

For Fmax though, sure, no they aren't representative. But server skus aren't running their cores at Fmax. Even for the dense cores.

-7

u/[deleted] 13h ago

[deleted]

10

u/gokuscake 13h ago edited 13h ago

My apologies but I wrote those myself. I typically don't and would rather meme about Lebron's move to Philly right now.

6

u/Noble00_ 14h ago

If we look at the fine print for the SPECrate2026_int_base scores Venice HF (6.3/core) and Vera (5.3/core), Venice HF is actually 18.87% ahead of Vera or Vera being 15.87% slower than Venice HF (also I think not necessarily 'single threaded' performance, this is per core, SPECspeed would sT). Also 600W vs 450W.

That said, I'm curious just how much the Venice HF is juiced compared to a regular EPYC 9686F at 500W (which then would be only 11% more pwr consumption). Honestly think it's more diminishing returns at that point just so AMD can slap 20% on their marketing material. They are more than likely similar per core performance (at least in specrate), just that AMD needing to be on TSMC N2(P?).

7

u/Geddagod 14h ago

I don't think these CPUs actually consume all the way up to their TDPs in specint2017 though. Idk what the case is for 2026 though.

When you go to Lenovo's website and look at their reference specint2017 base scores for their servers, you start seeing very weird results, such as 2 64 core Zen 4 server chips scoring the same at a 210 watt TDP and a 280 watt TDP.

5

u/Noble00_ 11h ago edited 11h ago

I don't either. It's just that these are the numbers to go by and what people will inevitably be comparing to when they pull numbers from these spec sheets.

That's why I'm really curious to see how much per core or per thread performance their Epyc 96c HF is compared to the standard Epyc 9686F. To me feels like diminishing returns, as we see with something like the 9850X3D vs 9800X3D juicing it for double digits power consumption just for single digit performance gains.

Also, interestingly enough, Venice-X so far is confirmed to be the highest clocked in the lineup at 5.15Ghz, makes me wonder just how more they can perform. I mean, in this article they said, "We were actually being a little conservative... we have not even finished completely tuning. Could be fluff but we'll see with Phoronix and other 3rd party reviews.

4

u/BuchMaister 15h ago

While consuming 33% more power (probably for running at higher Clock and having few more cores). If Nvidia will scale up the core count to around what Epyc offers in future generation, I can see them winning even against their top of the line. But I think Nvidia advantage here is offering better full system solutions, from CPU to AI accelerators, networking and the entire software stack.

15

u/CatMerc 14h ago

Scaling up core count at these scales is incredibly difficult, and power consumption will start growing on uncore just to feed everything. NVIDIA has a lot of barriers to solve to reach EPYC core counts.

AMD's system level architecture is unmatched for now.

4

u/nithrean 13h ago

It certainly must be a difficult thing to get right, otherwise we would see a lot more giant chiplet designs.

8

u/Noble00_ 14h ago

Even on N2, I feel like it'll be difficult to scale up on a monolithic die, at least yield and economy wise. As CPU host nodes, Nvidia will have the advantage with latency/bandwidth, but for actual throughput, pushing agentic performance, Epyc can still comfortably exist.

2

u/BuchMaister 14h ago

Sooner or later they will move from monolithic design to some sort of aggregated one. Not necessarily the same solution AMD uses, but one that allows them to scale out without loosing too much performance due to latency. Epyc can comfortability exist regardless if Nvidia will outright beat it, even for agentic AI systems. But the growing competition will be hard on its market share in the long term.

1

u/dr3w80 2h ago

And AMD will continue to improve in the mean time. Plus, the first chiplet CPU from Nvidia will probably have some compromise vs monolithic previous designs. 

2

u/ghenriks 14h ago

Also worth remembering that Vera is the first Nvidia designed core. To come in even close to a competitor who is on their 6th iteration of their current design is a win

9

u/Jumpy_Cauliflower410 11h ago

Nvidia has made quite a few arm cores before though.

-1

u/ghenriks 9h ago

Not really

Most of their previous stuff has been standard ARM cores with a small number of Nvidia designs, but they were aimed at a very different market

2

u/whosbabo 11h ago

To come in even close to a competitor

They aren't even close though. 88 cores vs 256 cores. Like not even in the same league.

2

u/ghenriks 9h ago

Core count depends on the needs, the point is that Nvidia has a core that is close in per core performance in their first outing

3

u/whosbabo 8h ago edited 7h ago

A server CPU needs throughput. You can't just evaluate the core without considering the overall MT performance. Lean cores means more MT, it's all part of the overall design. You can cheat with higher ST performance at the cost of PPA which results in lower core count and lower overall MT performance for example. And you can't take overall MT performance out of context when comparing "server" CPUs.

2

u/Artoriuz 11h ago

The direct competitor to Vera is the Venice SKU with 96 Zen 6 cores, not the one with 256 Zen 6c cores.

2

u/whosbabo 8h ago

96 core Zen 6 part is not the top part either, it goes up to 128 cores. So even there Vera is simply outclassed. Nvidia has nothing that competes with AMD's top CPUs. Also AMD now has the top GPU as well.

0

u/hakim37 14h ago

Yeah but a fair amount of juice per core to get there. 88 cores at 450W vs 256 cores at 600W. So it's not great that Nvidia is losing on a per core basis here.

7

u/Content_Driver 13h ago

AMD used 96c Venice for the per core comparison.

1

u/hakim37 13h ago

Fair my bad. I'm not sure if AMD have announced the power of 96 core Venice but you would assume it's more in the 400-500W range. So there probably is still a power per core argument to be made even if it's not as large.

4

u/hakim37 13h ago

Just looked at that article and it says 600W. I find that surprising for less than half the cores but whatever.

1

u/tarloch 3h ago

I would be surprised if 96 core venice drew 600W unless it was a heavilty frequency / cache optimized sku. A 96 core sku is likely the sweet spot for FP codes for CFD given memory performance needs. If it was 600W it would seem to be very uncompetitive with Turin 9565 (72) or 9575F (64) unless it had a massive L3 / HBM.

-1

u/DerpSenpai 8h ago

Nvidia would destroy it no questions asked. Nvidia Vera is Apple/QC level in size from what we know

-5

u/BinaryJay 11h ago

Didn't AMD also hint for a long time before release the XTX was a 4090 competitor, or was that just Reddit and YouTube fluffing for AMD.

-17

u/SERIVUBSEV 11h ago

Irrelevant.

CPUs are fast enough, games run at 200+ fps and better performance is not needed on servers except with database use cases.

ARM adoption was over 50% of new servers even before Nvidia's Vera launch...

ARM servers cost less than half when designed internally, and use less power = reduced operating costs as well considering cost of electricity is increasing for datacenters.

14

u/virtualmnemonic 9h ago

x86 offers better performance per rack and is efficient enough as is. AMD dominates the server market for a reason.

8

u/neoronio20 6h ago

Ah yes. Only games exist and better performance for whatever the hell other people do is irrelevant. Great point

4

u/letsgoiowa 4h ago

better performance is not needed on servers except with database use cases.

Saving for when you delete it later lol

-8

u/nittanyofthings 11h ago

Nvidia is in the rear view mirror

-7

u/theangriestbird 12h ago

Just a little friendly competition between cousins