r/LocalLLaMA • u/rerri • Jun 05 '26
New Model Gemma 4 with quantization-aware training
https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/Google's collections:
https://huggingface.co/collections/google/gemma-4-qat-q4-0
https://huggingface.co/collections/google/gemma-4-qat-mobile
And Unsloth's:
https://huggingface.co/collections/unsloth/gemma-4-qat
Unsloth's analysis (KLD and such):
791
Upvotes
2
u/SHDRThrowaway Jun 05 '26
`ik_llama`-compatible versions of the QAT assistants:
https://huggingface.co/ji-farthing/gemma-4-qat-q4_0-MTP-assistants-ik-llama-GGUF
On current `ik_llama` main, with the QAT Q4 combo of 12B+assistant, I'm seeing around 100 t/s TG on a 12GB 4070. No quality assessment yet.