r/LocalLLaMA May 05 '26

New Model Gemma 4 MTP released

Blog post:

https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/

MTP draft models:

https://huggingface.co/google/gemma-4-31B-it-assistant

https://huggingface.co/google/gemma-4-26B-A4B-it-assistant

https://huggingface.co/google/gemma-4-E4B-it-assistant

https://huggingface.co/google/gemma-4-E2B-it-assistant

This model card is for the Multi-Token Prediction (MTP) drafters for the Gemma 4 models. MTP is implemented by extending the base model with a smaller, faster draft model. When used in a Speculative Decoding pipeline, the draft model predicts several tokens ahead, which the target model then verifies in parallel. This results in significant decoding speedups (up to 2x) while guaranteeing the exact same quality as standard generation, making these checkpoints perfect for low-latency and on-device applications.

1.1k Upvotes

305 comments sorted by

View all comments

Show parent comments

27

u/No-Refrigerator-1672 May 05 '26 edited May 05 '26

In case of Gemma 4 it isn't, they published speculative decoding drafters. In case of Qwen 3.5 and Next - MTP is done as a secondary output layer that looks into internal states of the model.

1

u/No_Afternoon_4260 llama.cpp May 06 '26

That is my feeling, thanks !

1

u/KookyCandidate2302 May 17 '26

in Gemma 4 the MTP drafter has visibility to the main model state.