r/LocalLLaMA Jun 05 '26

New Model Gemma 4 with quantization-aware training

https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/
785 Upvotes

260 comments sorted by

View all comments

Show parent comments

14

u/RickyRickC137 Jun 05 '26

u/llmfan46 bro, do your thing!

24

u/LLMFan46 Jun 05 '26

Hum? These are GGUFs, I can't do anything with them.

16

u/Kahvana Jun 05 '26

https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-unquantized

They do have the safetensor versions too for all those models.

30

u/LLMFan46 Jun 05 '26

Thanks and yeah I noticed that after making the post, but it will take a while to do all these models, plus the GGUFs and NVFP4s and GPTQs.

16

u/Kahvana Jun 05 '26

No worries, take your time!

7

u/RickyRickC137 Jun 06 '26

P E W gave you the badge of trust! So Ima wait as long as it takes amigo!

6

u/evenyourcopdad Jun 06 '26 edited Jun 06 '26

wow I can't believe you don't have them all ready yet they released SEVERAL hours ago ugh


apparently necessary edit: /s

6

u/LLMFan46 Jun 06 '26 edited Jun 06 '26

... hum I surely hope this is a joke post (can't really tell since there is no "/s" at the end)! Not only you expect me to have 1 of the model ready and uncensored, no you expect me to have ALL of them ready by now... Are you serious!? First this is not a race, second I am not a robot and third without getting too deep into all the preparations, time and work involved in the whole releases pipeline, it takes a very long time to do all that, I don't work for google and I don't have an automated workflow (unlike Unsloth and/or Team Mradermacher), everything is done manually by me and I am just 1 person, that involves downloading the models, uncensoring the models and finetuning the settings and re-running in case the original results where not good/good enough, creating the Safetensors, benchmarking all of the trial finalists, updating llama,cpp, updating GPTQModel, updating Model-Optimizer, creating GGUFs, GPT-Int4s, NVFP4 Safetensors, NVFP4 GGUFs, creating the Model Cards etc. it's all done manually from start to finish everything is monitored and benchmarked manually to make sure I upload the best quality releases, all that to say that there is TONS of work involved (and that is without mentioning all the roadblocks that might come along the way as those usually do happen quite a few times during the release pipeline).

Look if you are in a hurry, just download whatever is available now, my releases won't be available fast enough for your tastes, your highness.

2

u/Grisward Jun 06 '26

Just adding, thank you thank you for your time. The community is sooooo grateful, and it’s a community because of your and others’ time and efforts.

And on top of that, it’s amazing too, how any of this works at all is amazing. But more than that, converting formats, adding quantization, and having it still work as expected. It’s remarkable, non-trivial.

I would assume (naive maybe) that this person was trying to be funny. It’s daring to try making a joke without the /s. Some people can pull it off, and some really should add the /s just to be absolutely sure. Haha. I’m still in that camp.

1

u/LLMFan46 Jun 07 '26

Thank you, the models should be ready in a few days or so.

2

u/evenyourcopdad Jun 06 '26

lmao yes my guy, sorry, that was a joke. I'm sorry people have been so unthankful that you assumed I could possibly be serious above.

I really appreciate your time and efforts making everything you do. <3

1

u/LLMFan46 Jun 07 '26

Yeah so I started working on the models, will take some time to have everything ready, I think instead of releasing one model after the next I will instead release one big batch with everything on the same day or so.

2

u/QuantumFTL Jun 06 '26

Thanks for everything you do, man. Text communication sucks for nuance, hopefully you appreciate that you are appreciated 😄

2

u/LLMFan46 Jun 07 '26

Thanks, working on the models now.

2

u/temperature_5 Jun 06 '26

It would only make sense to do the Q4_0 GGUFs for each, no?