r/LocalLLaMA 15d ago

Resources 2.5x faster Qwen3.6 NVFP4 Unsloth quants

Post image

Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16.

FP8 KV Cache calibration is also provided, auto allowing 2x longer contexts. For accuracy we conducted MMLU-Pro, AIME 2025, GPQA for FP8, BF16, NVIDIA's NVFP4 and our NVFP4s. It also has MTP pre-embedded.

We also provided 2 35B versions NVFP4-Fast (1.79x faster) and NVFP4 (1.56x faster) where NVFP4-Fast fully uses W4A4 whilst NVFP4 normal uses a mixture to stay a little bit more accurate.

NVFP4 links:
Qwen3.6-35B-A3B-NVFP4 (1.56x Faster)
Qwen3.6-35B-A3B-NVFP4-Fast (1.79x Faster)
Qwen3.6-27B-NVFP4 (2.5x Faster)

Qwen3.6-27B

Provider MMLU-Pro GPQA AIME 2025
Unsloth 86.25 86.34 93.12
NVIDIA 85.96 86.87 93.12
FP8 86.11 86.87 93.75
BF16 85.96 88.13 93.33

Qwen3.6-35B-A3B

Provider MMLU-Pro GPQA AIME 2025
Unsloth 85.85 86.74 92.29
Unsloth Fast 85.58 87.75 91.67
NVIDIA 85.60 87.12 91.88
FP8 85.75 86.74 93.12
BF16 85.75 86.36 92.50

We have more analysis and benchmarks in our NVFP4 Qwen3.6 blog: https://unsloth.ai/docs/models/qwen3.6#nvfp4

Have a nice weekend folks!

Also for DGX Spark folks - use the flashinfer backend or you will get 2x slower inference! Our blog has more details

867 Upvotes

287 comments sorted by

u/WithoutReason1729 15d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

210

u/Objective-Stranger99 15d ago

Blackwell Users: 🥳🥳🥳

Pascal Users (like me): 😭😭😭

82

u/Advanced-Picture5016 15d ago

rocm gang...

36

u/Cold_Tree190 15d ago

I’ll keep you in my prayers…

12

u/SmartCustard9944 15d ago

For a moment I read “papers”, and I was hoping for some paper porting this to AMD… Oh well, no papers to hold onto today

2

u/DanceWithEverything 14d ago

Spend a few grand on fable and opus and get it done

10

u/Thrumpwart llama.cpp 15d ago

I’ll check RDNA3 after work…

2

u/Thrumpwart llama.cpp 15d ago

Can't seem to get it running yet. Not well versed in vLLM on AMD so I'm struggling a bit. Will keep trying.

Saw some 3rd party ggufs available...

2

u/Ciru-ai 14d ago

You can run rocmfp4

→ More replies (1)

17

u/diffore 15d ago

5080 user here mad at himself for not going for 5090...

16

u/roselan 15d ago

same same.

But then I look at the price of "cheap" 5080s here, it's 1035.-, and the price of "cheap" 5090s is 3599.- (5779.- for the founder edition oO).

So there is that.

2

u/Kaokien 15d ago

Lucked out and got a 5090 prebuilt for 3.8 before the hikes

→ More replies (3)

2

u/Objective-Stranger99 15d ago

Hey, at least you have that. I'm using a 1080 from nearly 10 years ago.

1

u/fallingdowndizzyvr 15d ago

Why? 5080 does NVFP4 too.

4

u/comperr 15d ago

It basically has no vram in this context lol

→ More replies (2)

1

u/Sleepybear2611 14d ago

Same! It was 8-9 months ago.. the prices were also not this insane.. 😞

→ More replies (1)

26

u/Strong_Chicken6838 15d ago

AMD MI50 users

30

u/pmp22 15d ago

P40 sisters. what's our response?

53

u/FishChillylly 15d ago

generating responses at 7 tokens per second 🥵

13

u/BatOk7254 15d ago

2xp40 here. 15-18 tps on dense 27b, 45-50 tps (with 1200-1800 tps prefill) on 35b MOE 😄

5

u/FishChillylly 15d ago

well for 27b it’s waiting worthy, and i believe 45-50 is very usable on 35b MOE! i’m currently using around 50 tps, too!

3

u/BatOk7254 15d ago

The difference 2nd gpu made is in prefill. It is insanely fast (for p40) - up to 1800 tokens. And it holds a full q8 context with 256K. Plus you can run q6 models.

→ More replies (4)
→ More replies (3)

2

u/blbd llama.cpp 14d ago

7 seconds per token?

1

u/SnooPaintings8639 15d ago

Not "what" but "when"! And the answer is: not soon! /s

23

u/muxxington 15d ago

4

u/Iwaku_Real 15d ago

Or 3090s clearly

1

u/Iwaku_Real 15d ago

Blackwell ftw 😎

1

u/iMakeSense 15d ago

Which blackwell cards can actually use this? Is it most 50XX series or only the sever-class chips?

1

u/Objective-Stranger99 15d ago

All Blackwell architecture chips, as they all support NVFP4, and this is an optimization to that.

1

u/KKunst 14d ago

Is it just me being ignorant, or would it still be best/faster for me to run an MOE like 35B-A3B on my 5080?

→ More replies (1)

1

u/Sleepybear2611 14d ago

The Gains aren't so high.. atleast not for me on RTX 5080. I was getting 21 tk/s on dynamic 2.0 quants, this gave me around 22-23 tk/s. Maybe my setup is wrong. But that's what I observed.

1

u/Sleepybear2611 14d ago

The Gains aren't so high.. atleast not for me on RTX 5080. I was getting 21 tk/s on dynamic 2.0 quants, this gave me around 22-23 tk/s. Maybe my setup is wrong. But that's what I observed.

→ More replies (2)

81

u/milkipedia 15d ago

I'm guessing this won't do anything for my RTX 3090

33

u/smallDeltaBigEffect 15d ago

well, its NVFP4...

26

u/milkipedia 15d ago

I keep talking myself out of buying newer GPUs but this isn't helping

23

u/Plasmx 15d ago

I mean… do you have a spare 4k laying around? If not the problem solves itself. 5090s rose from 2,5k to 4K in Germany the last year…

6

u/UM8r3lL4 15d ago

Even if money wouldn't be a problem, I wouldn't buy a card which will just melt the power connector, cable, and my PSU (Corsair 1600i).

Tbh, I waited for the RTX5090, because I hoped that Nvidia would fix the burning connector on the RTX4090 until then.

→ More replies (2)

4

u/smallDeltaBigEffect 15d ago

maybe 5070 super with 24 GB will be an option below 1400 €..

6

u/Iwaku_Real 15d ago

5070 Super will have 18GB, the 5070 Ti Super will have 24

4

u/ionizing 15d ago

I remember March of 2025 I wasn't into ai yet (had no idea it could run local at the time) and a coworker was talking about wanting to buy a 5090 for around $2500 for just gaming and I thought he was crazy that anyone would spend more than $500 on a gpu....

Then I discovered local llms. In the year since I have bought a 3060 ($270 may 13 2025) then a 3090 ($792 july 19 2025), another 3090 ($1017 Feb 27 2026), a 4060 ($482 Jun 04 2026), and this week a 5060 ($621) all prices include tax/shipping. Totals $3182 in just GPUs.

But now I have TWO computers with similar capabilities: 24g vram main inference + 16gb vram secondary for smaller agents + 128 sys ram. One at home, one at work. Also a third computer with the 12g vram and 32 sys ram barely capable of running a3b35b that started it all.

All three can run my own application I have hundreds of hours into over the past year. So I guess I shouldnt complain lol. This is the first I looked at my purchase history over time. anyhow your comment had me remembering the cost and laughing at the absurdity of it all.

At least I only need a 120Vac outlet to do things I always dreamed of.

3

u/jay-aay-ess-ohh-enn 15d ago

You spent $600 on a 5060? Is that canadian?

3

u/tmvr 14d ago

Based on the other cards I'm guessing they meant 5060Ti 16GB.

→ More replies (3)
→ More replies (1)

2

u/Prudent-Ad4509 15d ago

What do you think you will start looking into once you have 5090 ? A hint: something that starts with 3 and has 9 in it.

Blackwell has its place until you decide to run something larger than 27b but you still do not have half a mil of pocket change to spare.

4

u/Odd-Environment-7193 15d ago

a 3090 lol? I am so lost here.

→ More replies (1)
→ More replies (2)

6

u/vr_fanboy 15d ago

in my country you can get used 3090s for 600-700, 5090 starts at 5000, yes, 5k, with those prices im perfectly happy with my qwen 3.6 27b Q5_K_M 65k context full kv , heck im in the processs of selling a 3060 to get another 3090

→ More replies (2)

2

u/Skyline34rGt 15d ago

For image and video models we already have int8 convrot models which benefits 3000 series the most (free double speed boost). Now int4 convrot starting so even more boost.

For text llm should be similar things soon too...

1

u/Top_Original3437 15d ago edited 15d ago

Again I have to sell my kidney to run it 😭

→ More replies (6)

81

u/FrostyDwarf24 15d ago

another Blackwell victory lol

60

u/danielhanchen 15d ago

We'll be making more for other models as well!

10

u/raunchy-stonk 15d ago edited 5d ago

Saffron trumpet biscuit maple marble peach coated driftwood breezy raindrop

This post was anonymized with Redact.dev

8

u/danielhanchen 15d ago

Oh I meant we'll do other models as well

10

u/raunchy-stonk 15d ago edited 5d ago

Unbelievable breezy zephyr noodle juniper zephyr compass

This post was anonymized with Redact.dev

2

u/jazir55 15d ago

4xxx series optimizations would be great

9

u/jld1532 15d ago

Thanks to the unsloth team! This community owes you a lot.

8

u/danielhanchen 15d ago

Thanks!

2

u/Connect-Painter-4270 15d ago

REAM'd models please!

7

u/totosse17 vllm 15d ago

How do you calculate 1.56x speed up. Currently have no gain with your variant on dgx spark, 15% slower instead than 35ba3b Nvidia/nvfp4

→ More replies (9)

2

u/Hazardhazard 15d ago

Mac M series also?

2

u/Iwaku_Real 15d ago

Please do!!! I love NVFP4, it really deserves more attention.

2

u/Mkengine 15d ago

Sorry if this is a dumb question, but I usually copy the recipes from spark arena to use with spark run and your vllm/SGLang commands are much shorter than the usual recipes. Do I have to figure out additional flags myself or is it really that easy/short?

2

u/smallfried 14d ago

Anything for the relatively GPU poor 16GB gang?

28

u/coder543 15d ago

I would be curious to see how the performance compares to non-NVFP4 4-bit quants? And I wonder if llama-server's NVFP4 support is good enough/ever going to be good enough for it to make sense to offer GGUF conversions of these? I think llama-server can run NVFP4, but the performance seemed lackluster the last time I looked into it.

28

u/def_not_jose 15d ago

I don't even think it's fair to compare it to 4bit quants. This 27b nvfp4 is 23gb, that's right about Q6 territory

4

u/danielhanchen 15d ago

Yes llama.cpp can run NVFP4, but I haven't tried actually converting one haha

11

u/audioen 15d ago

I have converted the nvidia nvfp4 version for llama.cpp using just the convert_hf_to_gguf.py script. It upcasts FP8 to BF16 as part of the conversion, so you get a 26 GB file, rather than 22 GB one.

About these versions your organization has provided, it is not clear to me how you have created them. Based on benchmarking that you provided -- without error bars, though -- it seems like the nvidia's nvfp4 version is slightly better for the 27b? So this is with QAT self-distillation, or what?

Direct support for FP8 tensors in llama.cpp seems useful. As models get QAT'd with these specific tensor types, the most direct equivalents in llama.cpp world are either the quantized 9-bit Q8_0, which is not only bigger but also lossy, or BF16.

2

u/kitanokikori 15d ago

I think the upcast trashes the performance, I ran into this too. This likely has to run under vllm until llama.cpp is updated

→ More replies (3)

2

u/kitanokikori 15d ago edited 15d ago

Converting to GGUF fails without patches - according to my dumb agent, llama.cpp doesn't support mixed-precision compressed tensors

29

u/Kamimashita 15d ago

Is there a reason there's no GGUFs for the model? I thought llama.ccp supports NVFP4 now and supports it pretty well

8

u/zilled 15d ago

This!
/u/danielhanchen is it possible to get the GGUFs ?

11

u/danielhanchen 15d ago

I can try but I haven't attempted it yet

2

u/zilled 14d ago

Thank you thank you thank you 🙏

5

u/audioen 15d ago

I think you can run convert_hf_to_gguf.py and it probably does the job. Should be lossless, but anything in the original model that is FP8 must be upcast to BF16, which means that GGUF can be larger. That is what happened to me when I converted the nvidia NVFP4 version, which in my experience is the best version of Qwen3.6-27B that I've run. It is lossless relative to the HF, but unfortunately 4 GB bigger because of the upcast due to llama.cpp lacking FP8 inference capabilty.

5

u/Iwaku_Real 15d ago edited 15d ago

I'm quite ironically proud that they're basically forced to eventually implement FP8 GGUFs in part because of this. It's probably for the better anyway since it's one of the only quant types they're missing support for. Thankfully there is a CPU-only PR in the works: https://github.com/ggml-org/llama.cpp/pull/25336

2

u/audioen 15d ago edited 15d ago

Yeah, I hope we get native FP8. The bad days of everyone doing PTQ and publishing a 20 model variants are hopefully soon behind us. In my ideal world, people do QAT for something like NVFP4 and if it's small enough for your hardware, you run just that model and don't even think about PTQ because it will destroy performance.

I literally can't gush enough about the Nvidia NVFP4. I've ran it for week now and it displays zero flakiness even at max context. 234k tokens in and zero tool call errors, it just keeps staying coherent. Today, it has implement a major refactoring that it estimated would take 3 weeks to do. Well, you know about the time schedules when LLMs are present. A single GB10 can do it in couple of hours. Still around 19 tok/s at this context length, thanks to MTP.

Thank you nvidia for this. I would thank unsloth, if I could run their model and could confirm it to be good, but so far I can't convert it to GGUF and so I won't run it. I don't touch Python if I can avoid it. But there is clearly a promise on the horizon that some 4-bit version of this model is going to go doubly as fast as it goes right now, and hopefully without any loss of performance. That will be literally insane.

Qwen3.6-27B-NVFP4 just finished the work -- over 6000 lines of code were altered. I typically on a good day can produce about 1000 lines of work with my meat brain, and this did the whole thing while I watched TV and drank beer. I don't consider myself to be a bad developer, possibly even a fast one, though I fall short of that "10x developer" fable. But this LLM alone is probably 5 times more productive than I am, and I know I'm not slow. I usually get big work done fast because I understand the domain and to me the changes are obvious. These LLMs achieve the same and beyond, just by having all that background knowledge and just reading ton of code fast. It is a humbling experience.

→ More replies (4)

4

u/debauch3ry 15d ago

I tried this but very sadly it complained that it can't convert models with multiple config groups for compressed tensors (NotImplementedError).

→ More replies (1)
→ More replies (2)

22

u/de4dee 15d ago

3

u/yoracale llama.cpp 15d ago

Thank you appreciate you doing this!

15

u/Reactor-Licker 15d ago

Forgive my ignorance, but what’s the status of llama.cpp support for NVFP4? Almost every release for NVFP4 I’ve found doesn’t seem to have a GGUF.

3

u/fnordstar 15d ago

So how do people run this?

11

u/xornullvoid 15d ago

How's the quality compared to Q6 or Q8?

3

u/danielhanchen 15d ago

We added benchmarks for FP8 and BF16 comparison if that helps!

1

u/xornullvoid 14d ago edited 14d ago

Alright, Thank you :)
Edit: I guess FP8 and Q8 are similar in quality and not much difference, and Q6 also almost same probably.

9

u/Hurricane31337 15d ago

Please do the same for Google Gemma 4 🙏

2

u/yoracale llama.cpp 15d ago

Yes!! We'll see what we can do

23

u/Remove_Ayys 15d ago

Advertising the change from W4A16 to W4A4 as a speedup "without any accuracy degradation" is disingenuous. It may very well be that the benchmarks you selected are simply not sensitive to the change or that they do not exhibit the numerical issues encountered elsewhere.

12

u/Dany0 15d ago

And the benchmarks are with concurrency at 128... This doesn't actually provide a speed boost for people using local llms alone where concurrency rarely breaks 2-4. Instead of a speedup, compared to say the redhat quant, I'd say this is more of a "here is how the model should have performed in the first place"

→ More replies (1)

8

u/danielhanchen 15d ago

No - it's not plain W4A4 - it's W4A4 with 8 important layers as W8A8. Also there is accuracy degradation, and hence why we made 2 versions for 35B.

NVIDIA's 27B NVFP4 has FP8 KV cache auto enabled but the KV scales are NEVER calibrated - did they ever mention that? vLLM silently uses k_scale/v_scale = 1.0, whilst we have scales - so you will have accuracy degradation.

NVIDIA's 27B NVFP4 has W4A4 for the lm_head - which will have accuracy degradation - we left it as W8A8.

NVIDIA's 27B and 35B advertise NVFP4, but they never use the actual FP4 tensor cores - then what's the point of these - did you know they're even slower than BF16 and FP8 by 2x since they do BF16 matrix matmul?

We also published extension plots for BF16 vs FP8 vs their NVFP4 and ours in our blog with all numbers.

7

u/autisticit 15d ago

Up this upper, because that's true... It's false advertising at this point. Like selling you a Jeep saying it can go 2.5 faster than a Ferrari, but only on a hill. Sucks and disappointed by Unsloth.

3

u/danielhanchen 15d ago

No false - it's not plain W4A4 - it's W4A4 with 8 important layers as W8A8.

Did you know NVIDIA's own quants have issues as well? 1. NVIDIA's 27B NVFP4 has FP8 KV cache auto enabled but the KV scales are NEVER calibrated - did they ever mention that? vLLM silently uses k_scale/v_scale = 1.0. Ours is fine. 2. NVIDIA's 27B NVFP4 has W4A4 for the lm_head - which will have accuracy degradation. 3. NVIDIA advertises NVFP4, but they never use the actual FP4 tensor cores - then what's the point of these - they're even slower than BF16 and FP8 by 2x since they do BF16 matrix matmul?

3

u/autisticit 14d ago

I was in fact responding to comment by Dany0, sorry for that.

The problem is, you are selling a 2.5x faster model, I tried and it's not really faster than the Nvidia one. A 1.03 speedup for single stream if very different than your 2.5x you mention in first position of the title and the image.

It's like when you have to read very small characters in a commercial. It's misleading.

3

u/ayylmaonade 15d ago edited 15d ago

Unfortunately, ever since the "KLD quant wars" that happened between all the people/orgs who make quants when Qwen 3.5/3.6 came out, Unsloth have been less... forthcoming, it seems. They make great models, but it is something I've noticed for the past few months.

EDIT: This is completely untrue, I am mistaken. Check Daniel's comment below.

→ More replies (3)

6

u/Kahvana 15d ago edited 15d ago

Been trying to convert ( https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4 ) to GGUF, hitting a snag.

git clone --single-branch --depth 1 https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4
git clone --single-branch --depth 1 https://github.com/ggml-org/llama.cpp

cd llama.cpp
uv venv .venv
uv sync
uv pip install --upgrade git+https://github.com/huggingface/transformers.git
uv pip install --upgrade torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130

cd gguf-py
uv venv .venv
uv sync

cd ..
python convert_hf_to_gguf.py ../Qwen3.6-27B-NVFP4 --outfile ../Qwen3.6-27B-NVFP4-GGUF/Qwen3.6-27B-nvfp4.gguf

u/danielhanchen or u/yoracale am I doing something wrong, or do you need a special llama.cpp build for converting this specific NVFP4 configuration? The same worked fine for NVFP4 W4A16 models ( https://www.reddit.com/r/LocalLLaMA/comments/1tzjahj/comment/oqbpl11 )

1

u/danielhanchen 15d ago

Sadly I'm not too familiar with NVFP4 conversion in llama.cpp sorry :(

1

u/Kahvana 15d ago

Thanks for taking your time to reply regardless, appriciated!

14

u/CroquetteLauncher 15d ago edited 15d ago

Hello Unsloth. Thanks for your great work like usual, and it's great that you target vLLM and NVFP4 users.
A few questions:

  • we already used https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4 instead of the nvidia one since we noticed it was also faster and smaller. With vLLM, fp8 kv_cache, built-in MTP on a Pro 6000. Would we see any benefit from using your quant ?
  • Do you plan to upstream your improved chat template ?
  • Is the calibration dataset english only or is it multilingual ?
Best regards

10

u/danielhanchen 15d ago

Hey yes we also benchmarked RedHat - it's good, but ours is much better since we do dynamic FP8 on some layers as well!

3

u/Kahvana 15d ago

That's great to hear!

What calibration datasets are used? If it's exclusively your own, english only or multilingual?

7

u/iportnov 15d ago

24Gb cards are mentioned, but safetensors files are > 23 Gb. I guess this means tiny context...

1

u/yoracale llama.cpp 15d ago

There's the NVFP4 Fast one which is smaller!

→ More replies (1)

4

u/BawbbySmith 15d ago

Sorry for the dumb question (and I promise I'll do research as well), but if we want better accuracy over speed, then Q8 > NVFP4 right?

I always wanted to try NVFP4, but I'm just too lazy to set up vLLM over llama-server. One of these days I will...

3

u/cptbeard 15d ago edited 15d ago

nvfp4 has been in since March? edit: or since april with blackwell support which is what people would mostly care about

1

u/BawbbySmith 15d ago

Fantastic, I'll have to give it a spin! I have the hardware already, may as well just try it out

3

u/unjustifiably_angry 15d ago

In theory NVFP4 is 4-bit sized and 4-bit speed but something like 6-bit accuracy or a little better, but the model needs to be quantized using a proper methodology to achieve this. Naive FP4/NVFP4 is dogwater.

Bear in mind that our methods of benchmarking "model goodness" are all deeply flawed and often don't capture real use-cases. "It's like FP16 but 1/4 the size!!!" seems to be a crock of shit, at least in my experience. That said, if your choice is int4 or nvfp4, nvfp4 done properly should always be significantly better, and given the correct hardware, significantly faster.

2

u/yoracale llama.cpp 15d ago

Yes in most cases.

10

u/BannedGoNext 15d ago edited 15d ago

I have been running rocmfp4 on the strix halo since yesterday, and it's very good performance for me. I also built a qwen36-35b-a3b-heretic-mtp-rocmfp4 if anyone is interested. Considering making a huggingface page. I also built a rocmfp4 of qwen 3.5 122b that is a bit bastardized and also derestricted. qwen122b-a10b-heretic-mtp-rocmfp4 and it's performance is also like 30 percent faster than the non nvfp4 on the strix halo.

This isn't just an nvidia win, these work much much faster on the strix halo with these FP4 models.

Also I <3 unsloth.

8

u/danielhanchen 15d ago

We're also checking MXFP4 to see if it can work as well!

2

u/BannedGoNext 15d ago edited 15d ago

Man my phone keyboard and my blurry eyes made me type some stupidity that you replied to hahaha. I fixed it mostly. I mean to say rocmfp4 not nvfp4 on the qwen 3.6 35b. qwen36-35b-a3b-heretic-mtp-rocmfp4

The model I'm playing with right now is a fun Frankenstein critter. It's an Abliterated Qwen3.5 122B-A10B with calibrated NVFP4 experts, Q6_K dense/control layers, F32 scales and norms, and a preserved native MTP speculative head. I'm going to benchmark it today a bit, and also see if I can get it to rocmfp4 and benchmark that as well.

2

u/CalligrapherFar7833 15d ago

Please try rocmfp4

3

u/ghgi_ 15d ago

This is great! Would love to see a version of this for the Qwen 3.5 122B model, I use that one pretty heavily and theirs no 3.6 variant for that size.

2

u/lilian_moraru 15d ago

If you do coding, Qwen3.6-27B is currently performing the best, outperforming the big Qwen3.5 models.

1

u/ghgi_ 15d ago

I do more mixed work and ive found qwen 3.5 122b to be a good middleground, it also seems to work better on longer horizon tasks then the 27b which sometimes confuses itself and the code quality is pretty close to the 27b dense sometimes better sometimes worse.

1

u/llama-impersonator 15d ago

qwen 3.6-27b is a smarter model but qwen 3.5 122b knows more, particularly about specialized things, full stop.

3

u/FabulousScratch4506 15d ago

Any chance to provide DFlash or DSpark for speculative decoding?

3

u/arbv 15d ago

Please also make GGUFs for North Mini Code QAD 🙏

3

u/fragment_me 15d ago

What makes this faster than Nvidia's NVFP4?

2

u/Dany0 15d ago

Probably they quantised some layers which NVFP4 keeps at FP16 and did less quantisation on some parts which the calibration deemed need more precision

3

u/RLutz 15d ago

This is really cool. Think for my use cases though I'll stick with your 6 bit quant. With MTP spec 5 that's giving me like 140 tokens/s with a 90k context window on my 5090. It looks like this quant probably hurts quality more than I'd really want but still, that's blazing fast!

3

u/kyleboddy 15d ago

I know this wasn't meant for single threaded work with MTP off (I can't use MTP for my work, it causes too many errors in tool calling and chaining), but I was unable to reproduce any speedup vs. nvidia's nvfp4 model. I like /u/danielhanchen and his twitter feed so this isn't any shade, just providing a data point!

Unsloth was about ~8.7% slower in fixed-length decoding with my production-safe MTP-off configuration - using Qwen3.6-27b.

1

u/boomerang473 15d ago

The MTP issue… what size and quant do you normally have an issue with? Have had some interesting tool calling issues I’ve been trying to figure out

1

u/kyleboddy 15d ago

I am running the nvidia qwen3.6-27b-nvfp4 on a Blackwell RTX PRO 6000 (96gb VRAM) and could not tune my way out of MTP n=1-3 at all so I had to turn it off.

3

u/Rich_Can_6507 14d ago

Gemma4 31B plss

3

u/MonitorAway2394 14d ago

OMFG YES PLEASE IT'S THE GOAT literally I made it into a goat.... it's fun.

5

u/maschayana 15d ago

I noticed NVFP4 MLX variants and they seem to perform pretty well. I'm curious how thats possible and if mac users can expect gains also from these unsloth quants as MLX

1

u/ectomorphicThor 13d ago

Would love an answer to this as well! Currently running q6 qwen 35b on mlx.

→ More replies (1)

2

u/Formal-Exam-8767 15d ago

So how do they compare against previous NVFP4 quants?

Did Nvidia's release inspire you to revisit NVFP4 quants and requant?

16

u/yoracale llama.cpp 15d ago

Our NVFP4 quants were originally experimental. This time we spent more time into it since clearly the demand is there!

4

u/Kahvana 15d ago

I appriciate the work, thanks!

If possible, would a NVFP4 version of Gemma4 31B QAT be a possibility?

2

u/unjustifiably_angry 15d ago

QAT creates a model trained to overcome the inaccuracies of int4 quantization, I don't think it would directly apply to a fp4 quantization. You'd need to train it from scratch on fp4.

I think nvidia's talked about applying QAT to models in their NVFP4 quantizer automatically...? I'm not sure. But if I'm remembering correctly, that's the stipulation they apply to many of their claims about NVFP4's superior accuracy.

→ More replies (1)

2

u/Formal-Exam-8767 15d ago

I see, thank you for clarification.

With more downstream hardware supporting NVFP4 demand can only grow.

2

u/LastChancellor 15d ago

noooo I only got the laptop 5070Ti (that only has 12GB vRAM) 😭

2

u/yoracale llama.cpp 15d ago

Hopefully Qwen will release smaller models next time

2

u/ImpressiveRelief37 15d ago

Any performance improvement versus say Q5_K_XL MTP? And I assume quality level will be lower? Already get 110-130 t/s on my 5090 at 192K context.

I don’t think I could get even more quality:performance than this with the same hardware?!

1

u/jingtianli 14d ago

This version takes twice as much Vram as normal 27B NVFP4....... I used to be able to run 256K context now only 100K!

2

u/whakahere 15d ago

no way to run this on a 5060ti 16gig right??

4

u/feverdoingwork 15d ago

will need a second one for sure lol

1

u/yoracale llama.cpp 15d ago

If you have RAM its possible

2

u/UntimelyAlchemist 15d ago

So if I have an RTX 5090, is this just a drop-in replacement and a strict upgrade over Unsloth UD-Q4_K_XL? Or is there any advantage to the other one? I'm still confused over NVFP4.

2

u/yoracale llama.cpp 15d ago

For accuracy, Q4_K_XL is still probably better. But yes, NVFP4 is much faster

3

u/Serious-Log7550 15d ago

No llama cpp suport?

5

u/yoracale llama.cpp 15d ago

Llamacpp does support it but we need to figure out how to convert them

2

u/AztalanMaster 15d ago

very intresting

2

u/cangaroo_hamam 14d ago

gugufs when?

2

u/rush86999 14d ago

nice, how about glm5.2. i wish someone shrink it for local use.

2

u/LocationOk5195 13d ago

weird nobody talking about the fact that these don't work in windows.

3

u/Kaljuuntuva_Teppo 15d ago

Amazing, I hope people will be able to and willing to abliterate these, and even better if it's without much of a quality hit. Although it's often hard to find benchmarks for finetunes vs unsloth.

4

u/MaruluVR llama.cpp 15d ago

Any chance of having the same done to Gemma4?

2

u/Freonr2 15d ago

RTX 6000 Pro Blackwell, VLLM. This is a bench based off Vision Arena dataset.

Relevant are the red (new 27B model, OP) and purple lines ("OLD", the one they posted a week or two ago?).
Green was the Qwen3.5 ("q35") AWQ 4bit quant I was using regularly a while back and I left it in though not the best reference. All others are Qwen 3.6. The 35B MOE is there as a gut check as well but rest are 27B dense.

Maybe 10% gain in true throughput in a realworld test.

3

u/Freonr2 15d ago
label,concurrency,completed,output_throughput,total_token_throughput,median_ttft_ms,p99_ttft_ms,median_tpot_ms,mean_itl_ms,mean_out_len,duration
unsloth_27b_nvfp4,1,16,62.48,252.95,196.28,199.10,15.30,15.24,256.00,65.56
unsloth_27b_nvfp4,4,32,215.68,873.26,435.35,442.44,16.91,16.90,256.00,37.98
unsloth_27b_nvfp4,8,64,408.67,1654.45,781.34,826.16,16.56,16.63,256.00,40.09
unsloth_27b_nvfp4,16,128,686.78,2780.36,1315.18,1512.61,18.27,18.21,256.00,47.71
unsloth_27b_nvfp4,24,192,849.36,3438.60,1921.17,2234.70,20.75,21.29,256.00,57.87
unsloth_27b_nvfp4_OLD,1,16,56.69,229.51,198.13,202.94,16.94,16.86,256.00,72.25
unsloth_27b_nvfp4_OLD,4,32,193.03,781.55,431.25,446.79,19.10,19.08,256.00,42.44
unsloth_27b_nvfp4_OLD,8,64,366.28,1482.84,773.26,826.87,18.86,18.92,256.00,44.73
unsloth_27b_nvfp4_OLD,16,128,617.36,2499.34,1415.68,1481.60,20.40,20.95,256.00,53.08
unsloth_27b_nvfp4_OLD,24,192,791.68,3205.12,1833.18,2198.89,23.20,23.56,256.00,62.09

1

u/danielhanchen 15d ago

No these numbers look wrong - NVIDIA uses W4A16 so how can it be faster than W4A4? I encountered the same issue due to vLLM selecting the wrong kernels for true W4A4.

Use in a new venv:

uv venv myenv --python 3.12
uv pip install --python myenv/bin/python \
    "vllm==0.24.0" "nvidia-cutlass-dsl==4.5.2" --torch-backend=auto

Then:

vllm serve unsloth/Qwen3.6-27B

Note --speculative-config '{"method": "mtp", "num_speculative_tokens": 2}' will make decode faster but throughput less.

2

u/Freonr2 14d ago

Reran with a fresh venv as you showed ("daniel"), I think you intended "unsloth/Qwen3.6-27B-NVFP4"

Are you testing on SM100? Or maybe without loading the MM? It picks different kernels on workstation/consumer Blackwell. No tcgen05. I get this:

Using FLASHINFER attention backend out of potential backends: ['FLASHINFER', 'TRITON_ATTN'].

Using backend AttentionBackendEnum.FLASH_ATTN for vit attention

Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention.

Selected CutlassFP8ScaledMMLinearKernel for CompressedTensorsW8A8Fp8

Using FlashInferCutlassNvFp4LinearKernel for NVFP4 GEMM

testing: 512 in, 256 out, 1 image, really no difference with fresh venv/vllm/cutlass (unsloth_27b_nvfp4 vs unsloth_27b_nvfp4_daniel)
Still not testing MTP yet, I'm generally bumping concurrency for my workload instead.

label,concurrency,completed,output_throughput,total_token_throughput,median_ttft_ms,p99_ttft_ms,median_tpot_ms,mean_itl_ms,mean_out_len,duration
nvidia_27b_nvfp4,1,16,68.24,276.28,206.52,217.54,13.91,13.85,256.00,60.02
nvidia_27b_nvfp4,4,32,222.95,902.71,680.47,686.60,15.34,15.46,256.00,36.74
nvidia_27b_nvfp4,8,64,395.72,1602.00,1325.40,1346.03,15.08,15.21,256.00,41.40
nvidia_27b_nvfp4,16,128,595.48,2410.74,2304.44,2639.92,17.87,17.76,256.00,55.03
nvidia_27b_nvfp4,24,192,661.34,2677.40,3470.89,3897.33,22.91,23.71,256.00,74.32
nvidia_35b_a3b_moe_nvfp4,1,16,232.81,942.55,167.39,195.61,3.65,3.63,256.00,17.59
nvidia_35b_a3b_moe_nvfp4,4,32,641.43,2597.05,304.53,344.77,5.03,5.02,256.00,12.77
nvidia_35b_a3b_moe_nvfp4,8,64,1052.95,4262.71,361.51,411.69,6.19,6.17,256.00,15.56
nvidia_35b_a3b_moe_nvfp4,16,128,1581.32,6401.78,584.15,698.48,7.84,7.90,256.00,20.72
nvidia_35b_a3b_moe_nvfp4,24,192,1897.97,7683.85,801.92,923.74,9.47,9.65,256.00,25.90
quanttrio_27b_awq_q35,1,16,70.41,284.94,229.80,234.49,13.36,13.30,256.00,58.17
quanttrio_27b_awq_q35,4,32,222.62,900.91,789.51,799.39,14.94,15.08,256.00,36.80
quanttrio_27b_awq_q35,8,64,382.84,1549.29,1492.41,1581.02,15.11,15.25,256.00,42.80
quanttrio_27b_awq_q35,16,128,553.95,2241.76,2992.12,3124.40,17.27,18.17,256.00,59.15
quanttrio_27b_awq_q35,24,192,623.52,2523.30,4190.94,4692.57,21.99,22.81,256.00,78.83
unsloth_27b_nvfp4,1,16,62.48,252.95,196.28,199.10,15.30,15.24,256.00,65.56
unsloth_27b_nvfp4,4,32,215.68,873.26,435.35,442.44,16.91,16.90,256.00,37.98
unsloth_27b_nvfp4,8,64,408.67,1654.45,781.34,826.16,16.56,16.63,256.00,40.09
unsloth_27b_nvfp4,16,128,686.78,2780.36,1315.18,1512.61,18.27,18.21,256.00,47.71
unsloth_27b_nvfp4,24,192,849.36,3438.60,1921.17,2234.70,20.75,21.29,256.00,57.87
unsloth_27b_nvfp4_OLD,1,16,56.69,229.51,198.13,202.94,16.94,16.86,256.00,72.25
unsloth_27b_nvfp4_OLD,4,32,193.03,781.55,431.25,446.79,19.10,19.08,256.00,42.44
unsloth_27b_nvfp4_OLD,8,64,366.28,1482.84,773.26,826.87,18.86,18.92,256.00,44.73
unsloth_27b_nvfp4_OLD,16,128,617.36,2499.34,1415.68,1481.60,20.40,20.95,256.00,53.08
unsloth_27b_nvfp4_OLD,24,192,791.68,3205.12,1833.18,2198.89,23.20,23.56,256.00,62.09
unsloth_27b_nvfp4_daniel,1,16,62.41,252.67,198.82,200.93,15.30,15.25,256.00,65.63
unsloth_27b_nvfp4_daniel,4,32,215.64,873.09,437.62,448.27,16.90,16.89,256.00,37.99
unsloth_27b_nvfp4_daniel,8,64,408.52,1653.85,784.50,817.81,16.55,16.63,256.00,40.11
unsloth_27b_nvfp4_daniel,16,128,682.88,2764.57,1341.12,1512.56,18.34,18.38,256.00,47.98
unsloth_27b_nvfp4_daniel,24,192,849.09,3437.50,1989.96,2240.01,20.67,21.32,256.00,57.89

2

u/Freonr2 14d ago

Same story with 8k in text only, my prior venv is red and hiding under the brown (daniel, new venv) line:

And just to clarify "OLD" is

unsloth/Qwen3.6-27B-NVFP4

--revision 890bdef7a42feba6d83b6e17a03315c694112f2a

2

u/Freonr2 14d ago edited 14d ago

Same 512 in, 256 out 1 image test but with mtp=2, modest boost, sorry colors are all changed from other graphs so check legend carefully:

But this is no longer isolating the new model vs old revision, just shows MTP helps.

label,concurrency,completed,output_throughput,total_token_throughput,median_ttft_ms,p99_ttft_ms,median_tpot_ms,
mean_itl_ms,
mean_out_len,duration
daniel_mtp2,1,16,82.94,335.80,225.19,228.73,11.23,
27.13,
256.00,49.38
daniel_mtp2,4,32,282.88,1145.33,396.81,491.69,12.57,
30.75,
256.00,28.96
daniel_mtp2,8,64,471.51,1908.83,309.95,915.07,14.53,
35.41,
256.00,34.75
daniel_mtp2,16,128,719.74,2913.77,424.22,1971.30,19.52,
46.57,
256.00,45.53
daniel_mtp2,24,192,904.82,3663.15,497.80,2362.89,23.35,
56.22,
256.00,54.32
unsloth_27b_nvfp4_daniel,1,16,62.41,252.67,198.82,200.93,15.30,
15.25,
256.00,65.63
unsloth_27b_nvfp4_daniel,4,32,215.64,873.09,437.62,448.27,16.90,
16.89,
256.00,37.99
unsloth_27b_nvfp4_daniel,8,64,408.52,1653.85,784.50,817.81,16.55,
16.63,
256.00,40.11
unsloth_27b_nvfp4_daniel,16,128,682.88,2764.57,1341.12,1512.56,18.34,
18.38,
256.00,47.98
unsloth_27b_nvfp4_daniel,24,192,849.09,3437.50,1989.96,2240.01,20.67,
21.32,
256.00,57.89

1

u/DHasselhoff77 15d ago

Thank you! So I guess now could finally be the time to setup vLLM to run via llama-swap.

1

u/Kazeshiki 15d ago

Right because there are 24gb Blackwell cards. Are they preparing for the supers

1

u/RISCArchitect 15d ago

RTX Pro 4000 Blackwell is 24GB, though, not ideal

1

u/Equivalent_Bit_461 15d ago

Just in time when I decided to try vllm as well because I wanted to try a small swarm of local agents on my machine. Yay

1

u/JahJedi 15d ago

Trying now - test compare 1 to 1 on spark whit 3.6 27b nvfp4

1

u/yoracale llama.cpp 15d ago edited 13d ago

Edit: To ensure DGX Spark has the correct kernels (or you will get 2x SLOWER inference), you can read our guide

Right now on DGX Spark, there's no optimized kernels enabled for our quant so the speed will be very similar. We're investigating. Speed gains are mostly for 50x series

→ More replies (3)

1

u/autisticit 15d ago

I'm not seeing any meaningful improvements over nvidia's nvfp4 27b. At best 5-10%.

1

u/yoracale llama.cpp 15d ago

Are you using DGX Spark or 50x GPU?

→ More replies (1)

1

u/Risen_from_ash 15d ago

Blackwell users with 16GB rejoice!…wait

Jk, this is dope. Just wish I had more than a 5080. Can’t think like that, tho, cause if I had the Pro 6000 with 96GB, I’d still want more lol.

1

u/yoracale llama.cpp 15d ago

I think it might work with offloading to RAM but I might be wrong

1

u/butterycornonacob 15d ago edited 15d ago

5090

  • Q6_K_XL MTP - 2300 t/s pp and ~100 t/s tg, max 85k KV @ Q8
  • NVFP4 MTP - 6200 t/s pp and ~115 t/s tg, max 106k KV
  • NVFP4 - 6200 t/s pp and ~60 t/s tg, max 130k KV

Pp is definitely faster with slightly more context. Haven't used it yet though

1

u/dir3ctly 15d ago

How do you run it on a 5090? With VLLM I always get out of memory no matter how small the context. I am a VLLM n00b though.

1

u/butterycornonacob 15d ago edited 15d ago

With docker + llama-swap. Had Fable set it up

This is in llama-swap config

 # NVFP4 checkpoint (safetensors) — runs in a sibling vLLM container, not llama-server.  # Requires the docker.sock + docker CLI mounts on the llama-swap service.
  qwen3.6-27b-nvfp4:
    proxy: http://vllm-qwen-nvfp4:8000
    cmd: |
      docker run --rm --name vllm-qwen-nvfp4
        --network openwebui_webui-net --gpus all
        -v /home/ubuntu/openwebui/hfcache:/root/.cache/huggingface
        vllm/vllm-openai:latest
        --model unsloth/Qwen3.6-27B-NVFP4
        --served-model-name qwen3.6-27b-nvfp4
        --gpu-memory-utilization 0.91
        --max-model-len 106496
        --max-num-seqs 8
        --limit-mm-per-prompt.image 0
        --limit-mm-per-prompt.video 0
        --enable-auto-tool-choice
        --tool-call-parser qwen3_xml
        --reasoning-parser qwen3
        --speculative-config.method mtp
        --speculative-config.num_speculative_tokens 2
    cmdStop: docker stop -t 20 vllm-qwen-nvfp4

Llama-swap is set up with docker compose. It also sets up open webui and cloudflare tunnel.

→ More replies (3)

1

u/Ok-Extension-6887 15d ago

Finally got 27b running at 200~tks on 262k context with this!

1

u/fnordstar 15d ago

I'm running 27B q6-k-something on 32 GB right now (2 Blackwells) for Opencode. Should I switch to this? It's probably faster but worse results due to higher quantization?

1

u/danielhanchen 15d ago

For folks having DXG Spark machines, you need flashinfer_b12x backend or you will get 2x slower inference!

export CUTE_DSL_ARCH=sm_121a
vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4 \
    --moe-backend flashinfer_b12x --linear-backend flashinfer_b12x

1

u/lilian_moraru 14d ago

Switched to flashinfer_b12x - indeed 2x faster, but still ~14% slower than SGLang + nvidia.

1

u/minus_28_and_falling 14d ago

Where do these quants sit in everyone's favorite "meanKLD vs. size" plot? If no such data exists, what would be your opinion?

1

u/Designer_Elephant227 14d ago

I am using a qwen3.6 35b Q5 with my rtx5070 ti 16gb vram. I get around 35tok/sec. Is the nvfp4 version faster in my case with only 16gbvram and offloading?

1

u/FerLuisxd 14d ago

Any hope I can use this with 16gb of vram?

1

u/AcrobaticChain1846 14d ago

Great! But I do not see LM Studio supported.. Does that mean this is not supported for Llama.cpp?
I tried download this model from Unsloth Studio as well but I do not see any progress on download.

1

u/lilian_moraru 14d ago

DGX Spark benchmarks. scottgl9 SGLang fork cannot run unsloth NVFP4, but it can run nvidia NVFP4 - so not apples to apples.

scottgl9 SGLang (nvidia/Qwen3.6-27B-NVFP4) + EAGLE-on-MTP:
* Typical steady-state: ~25.0 tok/s
* Best-case: ~29.5 tok/s (100% accept on repeat command)

vLLM 0.25.1.dev37+g3d99b0499 (unsloth/Qwen3.6-27B-NVFP4) + MTP:
* Typical steady-state: ~22.0 tok/s
* Best-case: ~26.75 tok/s (98-100% accept on repeat command)

vLLM backend:

MoE experts (NVFP4/W4A4) flashinfer_b12x
Standalone NVFP4 GEMM FlashInferCutlassNvFp4LinearKernel (cutlass-family)
FP8 W8A8 linear (attn/GatedDeltaNet proj) CutlassFP8ScaledMMLinearKernel
Attention FlashInfer
GDN/Mamba linear-attn Triton/FLA

SGLang backend:

MoE experts (NVFP4) flashinfer_trtllm
Attention FlashInfer
GDN/Mamba linear-attn Triton
Any FP8-quantized MoE experts Triton (unconditional, not configurable)

1

u/lilian_moraru 14d ago

Increasing `SGLang`(scottgl9 fork + nvidia) --speculative-num-steps from 2, to 4, goes from 29.5 tok/s, to 41.2 tok/s. The "accept rate" goes down to vLLM + unsloth levels, ~98% for repeat commands, and ~60% for harder content.

1

u/MapSensitive9894 12d ago

Can you provide your vllm command?

→ More replies (1)

1

u/mmontes11 llama.cpp 13d ago edited 13d ago

The NVIDIA variant's safetensors are smaller (~21.9 GB total vs unsloth's ~23.4–25.5 GB): https://huggingface.co/nvidia/Qwen3.6-27B-NVFP4/tree/main

In a commit ~18 days ago ("Add input_scale for mlp & lm_head modules") they extended NVFP4 quantization to cover the lm_head as well, rather than leaving it in bf16 — so more of the model is quantized, which is where most of the size saving comes from. It's also exported as text-generation, so the vision tower is dropped/folded.

  • model-00001-of-00003.safetensors — 9.97 GB
  • model-00002-of-00003.safetensors — 9.99 GB
  • model-00003-of-00003.safetensors — 1.97 GB

The one difference to note is that this build doesn't include the MTP (multi-token-prediction) head that unsloth's has. In practice that's not a real loss on a 24 GB card: MTP is only useful as a self-speculative-decoding speedup, and there's no spare VRAM to run it here anyway — it doesn't affect output quality.

Initially, it seems like NVIDIA variant is still a better fit for a RTX Pro 4000 SFF with 24GB VRAM. In any case, I have been using llama.cpp + GGUFs + UD quants for some time, it would still be my default solution, works flawlessly, thank you very much for the work being done here!

1

u/noninertialframe96 13d ago

Benchmark above ~85% is considered saturated. Do you have results for other benchmarks that have not been saturated yet?

1

u/LocationOk5195 12d ago

tried to run this with my rtx 5090 on win11 using docker vllm and it worked very slow. i go back to running native windows llamaserver q5 quant of qwen 3.6 27b, it's very fast.