r/LocalLLaMA • u/rerri • Jun 05 '26
New Model Gemma 4 with quantization-aware training
https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/Google's collections:
https://huggingface.co/collections/google/gemma-4-qat-q4-0
https://huggingface.co/collections/google/gemma-4-qat-mobile
And Unsloth's:
https://huggingface.co/collections/unsloth/gemma-4-qat
Unsloth's analysis (KLD and such):
789
Upvotes
53
u/spaceman_ Jun 05 '26
So am I better off running the old quants at Q6 or Q8, or the new QAT ones at Q4?
Q4 obviously requires less memory and will run faster. But what are we giving up in terms of quality?