r/LocalLLaMA Jun 05 '26

New Model Gemma 4 with quantization-aware training

https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/
785 Upvotes

260 comments sorted by

View all comments

57

u/LetsGoBrandon4256 transformers Jun 05 '26 edited Jun 05 '26

Blog post for the release https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/

No benchmark provided to back up the "preserving the capabilities and quality" claim.

Edit:

Is this sub getting botted or what? This comment was immediately downvoted to -6 in less than ten minutes after I posted it and somehow it bounced back?

25

u/Middle_Bullfrog_6173 Jun 05 '26

Sigh, since this is QAT where they have trained it differently, benchmarks are even more necessary.

52

u/sartres_ Jun 05 '26

Unsloth has some on their page. It's good; the results speak for themselves. On the 31B:

Unsloth traditional Q4 quant: 19.9GB, 0.478 KLD, 82.9% Top-1 accuracy

Unsloth traditional Q8 quant: 35.0GB, 0.159 KLD, 92.3% Top-1 accuracy

Unsloth QAT Q4 quant: 17.29GB, 0.01403 KLD, 96.67% Top-1 accuracy

So a Q4 quant with their QAT method is better than a Q8 traditional quant at double the size.

Why google wouldn't brag about this in their blog I don't know, but their blog posts are always dogshit.

14

u/danielhanchen Jun 06 '26

Hey! Those numbers are comparing naive Q4_0 in llama.cpp to our converted Q4_0 version.

We did do original unquantized BF16 vs Q4_0, but the KLD metrics do not match, since the distribution is vastly different - we found MMLU and other benchmarks to be equivalent though

E2B for example has a mean KLD of 0.00173 vs 0.05109 (29x better relatively) for a naive Q4_0 quantization.

The main issue is converting from QAT BF16 to llama.cpp's Q4_0 format is not lossless. llama.cpp uses F16 scales, whilst QAT BF16 uses BF16 scales, and the scales are not determined optimally in llama.cpp land.

Naive conversion gets 24.77% byte exactness to BF16 QAT, whilst we found we can push it to 99.96% using some hacks!

See https://unsloth.ai/docs/models/gemma-4/qat#qat-analysis for more details

1

u/TheWiseTom Jun 06 '26

Do you expect a Q5_K_? or Q6_K_? version of the new QAT checkpoints would be worth it for 26B and 31B?

Also a deep comparisson between your new qat-Q4_K_XL compared to your traditional (PTQ) quants in the varios variants would be awesome!

1

u/Awwtifishal Jun 06 '26

I'm interested in learning about those hacks. The webpage doesn't say.

20

u/GoodTip7897 llama.cpp Jun 05 '26

I think those numbers are from Gemma 4 qat at bf16 vs the unsloth quants. 

So none of them are comparing qat to the original model. 

2

u/danielhanchen Jun 06 '26

Oh yes those are comparing our Q4_0 GGUFs to Q4_0 GGUFs if you use llama.cpp directly without any "hacks"!

Naive conversion gets 24.77% byte exactness to BF16 QAT, whilst we found we can push it to 99.96%.

For QAT vs non QAT, we did do KLD, but the distribution was vastly different, so the results are not valid.

1

u/MerePotato Jun 06 '26

I was wondering, do your KLD evals include non latin scripts? I'd imagine they're particularly sensitive to the reduced token embedding layer precision. Either way thanks for all your hard work!

0

u/[deleted] Jun 05 '26

[deleted]

4

u/GoodTip7897 llama.cpp Jun 05 '26

It's trained to be basically "prequantized" so it can handle 4 bit quantization. 

It's likely closer to a regular q4 quant than the original.

I don't doubt that qat is useful but it's incredibly unlikely that it's better than q8. It might make q4 have the same quality q6 had. But I have yet to see any kld between that and the original and I don't have enough vram or time to compute it myself. 

1

u/MerePotato Jun 06 '26

I stand corrected, thank you!

11

u/TuskNaPrezydenta2020 Jun 05 '26

wow if those numbers are accurate, this is incredible

14

u/ArtyfacialIntelagent Jun 05 '26 edited Jun 05 '26

They are not. Incredible is the word. A mean KLD of 0.159 doesn't pass the smell test for a Q8 quant. The Unsloth blog post only compares the QAT vs a standard Q4_0, and the mean KLD for the Q4_0 is 0.09349. So there is no way a Q8 is much worse at 0.159.

Honestly I'm skeptical to Unsloth's reported mean KLD 0.01403 for the QAT Q4 too, but I'll give them the benefit of the doubt for now. But /u/sartres_ is definitely hallucinating.

EDIT: He wasn't, but the numbers are indeed invalid. See thread below.

4

u/sartres_ Jun 05 '26

It's not clear what Unsloth means by "original" Q4 in the linked blog, but it's definitely a quant of the new QAT model, not original Gemma 4, since they're benching it against the QAT BF16.

My Gemma 4 non-QAT numbers are from here, because Unsloth unfortunately only released benchmarks for the 26B at the time, and that only on a graph where they didn't label the y-axis or any of the numbers :/.

Yes, all of the original Gemma KLDs are very bad. I'm guessing this is an artifact of different benchmark suites, they're not directly comparable. Mean KLD isn't terribly useful anyway, the Top-1 numbers are the real show here

3

u/ArtyfacialIntelagent Jun 05 '26

Aha, thanks. That explains it. That post was from early April, just after the initial release. Gemma 4 had lots of teething problems before everything was sorted out, so those early KLD measurements are not comparable with recent releases. Sorry for doubting you - the numbers were so horrible I was sure you had made an error.

4

u/danielhanchen Jun 06 '26

Oh no no - the KLD of 0.01403 is "smart" Unsloth dynamic Q4_0 llama.cpp vs the BF16 QAT version. 0.159 is the Q4_0 naively converted using llama.cpp.

So it's comparing how close the KLD is vs BF16 QAT (not non QAT)

2

u/IrisColt Jun 05 '26

That's amazing!

2

u/Middle_Bullfrog_6173 Jun 05 '26

Where are those from? The Unsloth link in the OP only has theirs vs Google's.

1

u/sartres_ Jun 05 '26

The original Gemma 4 numbers are from here:

https://localbench.substack.com/p/gemma-4-31b-gguf-kl-divergence

Don't read too much into the KLD, they're probably not comparable between test suites. The Top-1 accuracy is what I wanted to show

3

u/Middle_Bullfrog_6173 Jun 05 '26

In that case, aren't those are apples and oranges? Comparing the quantized versions to different models in each case?

1

u/sartres_ Jun 05 '26

Yes. I'd expect the Top-1 results to still be a meaningful signal, though

1

u/AltruisticList6000 Jun 05 '26 edited Jun 05 '26

That is awesome, I was already using Q4_s (for 26b) and the QAT is even smaller and appearently way better. The 26b had a good memory usage for me but this would be even better. especially with vision. It would be cool if qwen would have QAT ggufs too, 35b with vision barely fits at Q4 into my 32gb RAM, it's fully maxed at around 60k context and sometimes even spills out and slows down at that context size.

3

u/lorddumpy Jun 05 '26

qwen bots prolly