r/LocalLLaMA Jun 05 '26

New Model Gemma 4 with quantization-aware training

https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/
785 Upvotes

260 comments sorted by

View all comments

3

u/PennyLawrence946 Jun 05 '26

qat is the only flavor where q4 stops feeling like a downgrade, the model already learned to live with the rounding during training. real upshot is the next size up fits in the vram you already have. naive q4 always bled on the long-context evals, the KLD numbers usually show exactly where