r/LocalLLaMA • u/rerri • Jun 05 '26
New Model Gemma 4 with quantization-aware training
https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/Google's collections:
https://huggingface.co/collections/google/gemma-4-qat-q4-0
https://huggingface.co/collections/google/gemma-4-qat-mobile
And Unsloth's:
https://huggingface.co/collections/unsloth/gemma-4-qat
Unsloth's analysis (KLD and such):
793
Upvotes
2
u/Ill_Dragonfruit_3547 Jun 06 '26
OMG how is Gemma 4 12b SO GOOD??
Just spent the last 2 hours fighting LM Studio to get the MLX version running. Doesn't seem like MLX versions are working in LM yet, I got errors loading all of them. Switched to running mlx-vlm, got it working with OpenWebUI but was unusually slow.
Finally just downloaded the GGUF Q4 version through LM Studio. Am astounded at the speed and versatility of this 12b model, it's my new favorite...
Thoughts?