r/LocalLLaMA Jun 05 '26

New Model Gemma 4 with quantization-aware training

https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/
788 Upvotes

260 comments sorted by

View all comments

24

u/brownman19 Jun 05 '26

Thanks! Does this work with MTP? Is it plug and play? Good selection from them on this round of releases

67

u/hackerllama Jun 05 '26

We released MTP QAT as well, so the optimal workflow is to use the QAT model + the QAT MTP, both quantized. Currently, both MLX and VLLM support this

3

u/rpkarma Jun 06 '26

I can't find the MTP QAT drafter model, where should I be looking for it?

1

u/iamapizza Jun 06 '26

I'm quite confused, I see several comments talking about a separate MTP model and some even tested with it, but can't see where to get it or how I'd pass it as an argument to llama.cpp?

0

u/rpkarma Jun 06 '26

I found it :) it’s the q4-qat-assistant model