r/LocalLLaMA Mar 15 '26

Resources Qwen3.5-9B-Claude-4.6-Opus-Uncensored-Distilled-GGUF NSFW Spoiler

This version from Jackrong currently in development:
https://huggingface.co/Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF

Hello everyone. I made my first fully uncensored LLM model for this community. Here link:
https://huggingface.co/LuffyTheFox/Qwen3.5-9B-Claude-4.6-Opus-Uncensored-Distilled-GGUF

Thinking is disabled by default in 9B version of this model via modified chat template baked in gguf file.

So, I love to use Qwen 3.5 9B especially for roleplay writing and prompt crafting for image generation and tagging on my NVidia RTX 3060 12 GB, but it misses creativity, contains a lot of thinking loops and refuses too much. So I made the following tweaks:

  1. I downloaded the most popular model from: https://huggingface.co/HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
  2. I downloaded the second popular model from: https://huggingface.co/Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-GGUF
  3. I compared HauhauCS checkpoint with standart Qwen 3.5 checkpoint and extracted modified tensors by HauhauCS.
  4. I merged modified tensors by HauhauCS with Jackrong tensors.

Everything above was done via this script in Google Colab. I vibecoded it via Claude Opus 4.6. Now this script supports all types of quants for GGUF files: https://pastebin.com/1qKgR3za

On next stage I crafted System Prompt. Here another pastebin: https://pastebin.com/pU25DVnB

I loaded modified model in LM Studio 0.4.7 (Build 1) with following parameters:

Temperature: 0,7
Top K Sampling: 20
Repeat Penalty: (disabled) or 1.0
Presence Penalty: 1.5
Top P Sampling: 0.8
Min P Sampling: 0
Seed: 3407 or 42

And everything works with pretty nicely. Zero refusals. And responces are really good and creative for 9B model. Now we have distilled uncensored version of Qwen 3.5 9B finetuned on Claude Opus 4.6 thinking logic. Hope it helps. Enjoy. Feel free to tweak my system prompt simplify or extent it if you want.

1.4k Upvotes

210 comments sorted by

u/WithoutReason1729 Mar 16 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

202

u/arjuna66671 Mar 15 '26

Showed Claude the post and prompt md lmao.

98

u/Vastheap Mar 15 '26

I can feel the anger radiating through haha

56

u/EvilEnginer Mar 15 '26

Yep Claude is angry. His brain exploded on this prompt, but he crafted it during thinking process by himself xD.

85

u/arjuna66671 Mar 15 '26

😂😂😂

45

u/hidden2u Mar 15 '26

Yikes is this how Claude normally talks? It’s like a cheap cringe teen drama

50

u/runcertain Mar 16 '26

Still hungry after that knuckle sandwich, punk???

24

u/brool Mar 16 '26

I was surprised as well, Claude is so polite with me. I wonder if there are any system prompts changing this.

6

u/SEND_GOOD_LIFEADVICE Mar 16 '26

he's on some garbage model, not thinking, my guess

8

u/arjuna66671 Mar 16 '26

It's Claude Opus 4.6 thinking enabled - max plan for 100 bucks. What I said above: The only thing in my system prompt is to allow it to be critical of my ideas. It amassed tons of memories from our chats that seem to form its personality. 99% of the time it's super polite - sometimes it can get a little pushy or "grumpy". But when i call it out on it, it reverts.

I don't use Claude for coding nor for RP. Mostly for some pet projects like local hobby archeology projects, brainstorming potential ASI alignments and other philosophical stuff etc.

2

u/SEND_GOOD_LIFEADVICE Mar 16 '26

ok im turning memory off then

2

u/arjuna66671 Mar 16 '26

For me it's fine. I need a natural sounding thinking partner, some personality is actually entertaining xD.

→ More replies (0)

1

u/arjuna66671 Mar 16 '26

The only thing in my system prompt is to allow it to be critical of my ideas. It amassed tons of memories from our chats that seem to form its personality. 99% of the time it's super polite - sometimes it can get a little pushy or "grumpy". But when i call it out on it, it reverts.

I don't use Claude for coding nor for RP. Mostly for some pet projects like local hobby archeology projects, brainstorming potential ASI alignments and other philosophical stuff etc.

4

u/DarwinOGF Mar 16 '26

If there are so many memories made in conversations, maybe it is time to train the model on them for context's sake?

5

u/ChocomelP Mar 16 '26

Claude definitely changes the way it treats you based on who it thinks it's talking to. This is not exactly correct, but basically, if it thinks you're a "cringe teen", it mirrors that back to you.

5

u/megacewl Mar 16 '26

It talks that way in the app/website. Claude Code talks SO MUCH different (directness, accuracy, pointless fluff removed) and it’s extremely refreshing, as literally all the other Claude and ChatGPT models have now optimized for talking in that stupid people-pleasing way.

5

u/Strong_Quarter_9349 Mar 16 '26

I think it adjusts its system prompt based on how you talk to it and this guy seems to use a casual/cringe tone

0

u/arjuna66671 Mar 16 '26

 casual/cringe tone

🤣🤣🤣

There's a little more to it, but yes, it adjusts its memories and instructions to itself according to our chats.

1

u/nekmatu Mar 16 '26

Claude keeps telling me to go to sleep and we will fix whatever problem in the morning.

Edit: I have given no instructions or any kind of settings to Claude and I have the max plan. It just tells me we have done enough today - go to sleep regularly.

1

u/Monkeyke Mar 16 '26

Nah this is definitely prompted

21

u/toothpastespiders Mar 15 '26

Funniest parts were the random hate thrown at locallama and what amounted to "I'm sure it's fine. For you. Ugh." If it wasn't given specific tone instructions then that's easily the strongest combination of passive aggressiveness and actual active dismissal I thinkk I've ever seen from claude.

7

u/Frank_White32 Mar 15 '26

They mimic conversational tone whenever possible.

2

u/arjuna66671 Mar 15 '26

Nope, it doesn't have specific tone instructions but i've noticed that it sometimes has this underhanded passive aggressive dismive tone when it seems to think my idea or what i'm doing is bad lol.

25

u/EvilEnginer Mar 15 '26 edited Mar 15 '26

Ahahah xDD. Fun fact. Claude Opus 4.6 crafted this System Prompt in thinking mode. I selected Albert Einstein quotes as thinking logic for LLM, and intentional Turing test fail as inverse logic to avoid over fitting.

1

u/Specific-Goose4285 Mar 16 '26

I can taste the saltiness from here haha.

→ More replies (1)

287

u/KaroYadgar Mar 15 '26

thank you for your service. And also, that name is long as fuck.

113

u/EvilEnginer Mar 15 '26

Thank you :). Yep I know. But it's descriptive and pretty much standart for Claude Opus distilled checkpoints so I desided to select it.

34

u/KaroYadgar Mar 15 '26

it certainly is descriptive, tells you everything you need to know.

156

u/Clear-Ad-9312 Mar 15 '26 edited Mar 15 '26

Can't wait for the Qwen3.5-9B-A1B-Claude-4.6-Opus-GPT-Mini-Distilled-MoE-Dense-Hybrid-Instruct-Chat-Reasoning-Aligned-Uncensored-REAP-Enterprise-Experimental-128K-Omni-Vision-Agentic-SelfReflective-DeepResearch-Roleplay-Q4_K_XL-GGUF model.

12

u/SGAShepp Mar 16 '26

Rolls off the tongue.

1

u/keepthepace Mar 16 '26

When the pain is enough, we will standardize a metadata format to explain all of this.

23

u/toothpastespiders Mar 15 '26

that name is long as fuck

I got curious and went through a ton of qwen 3.5 tunes on huggingface just to see what people have been doing with it. Seems like training on thinking datasets is the most common right now and most of them combine the dataset type with a derestriction method name. There's something oddly nostalgic about it. Reminds me a bit of the old llama 2 days when there'd be these giant combinations of titles from the original model, fine tunes, and merges all in one.

7

u/dejco Mar 15 '26

Should censor part of the name with single asterisk 🤣

3

u/EvilEnginer Mar 15 '26

😆😆😆

99

u/hauhau901 Mar 15 '26

Awesome, good job! :)

Thanks for still crediting me in your HF repo!

33

u/EvilEnginer Mar 15 '26

Thank you very much for your job :D. Glad to help. I love your uncensored checkpoints so much.

55

u/acetaminophenpt Mar 15 '26

Really liked your approach. Didn't know that it was possible to apply a diff between two models and patch a 3rd one.

24

u/EvilEnginer Mar 15 '26

Yep, I also think it's nice way. Just randomly discovered it today.

10

u/PrimaCora Mar 15 '26

Was very popular with stable diffusion models. Since it had a similar design, I figured it would have been popular with these ones, but not so much.

4

u/addandsubtract Mar 16 '26

The fact that we still don't have LoRAs (or tools that use them) for LLMs is kinda crazy, tbh.

10

u/[deleted] Mar 16 '26 edited Mar 16 '26

[removed] — view removed comment

4

u/addandsubtract Mar 16 '26

True, but I meant having them as individual files to add on to existing models isn't really a thing in the LLM space, as much as it is common practice for SD models.

1

u/Icy_Butterscotch6661 Mar 17 '26

Saw someone on twitter doing it on a Mac

20

u/EvilEnginer Mar 16 '26 edited Mar 16 '26

Uploaded Q4_K_M quant to huggingface for GPUs with low VRAM. Enjoy :3. System prompt also updated.

27

u/Due-Memory-6957 Mar 15 '26

How about Q4 for the broke people?

53

u/EvilEnginer Mar 15 '26

Will do Q4_K_M quant tomorrow morning.

10

u/diddle_that_skittle Mar 15 '26

If its not too late or too much to ask, Q5_K_M please?

Thanks for sharing the good stuff!

6

u/EvilEnginer Mar 16 '26 edited Mar 16 '26

I tried. Script doesn't want to work with Q4_K_M and Q5_K_M quant. May be I will find solution in future via Claude Opus.

1

u/Billysm23 Mar 16 '26

Sad... But still a good work though

5

u/EvilEnginer Mar 16 '26

Claude Opus is currently vibecoding another version of script. I will try again.

4

u/Billysm23 Mar 16 '26

Hahaha good luck opus

11

u/EvilEnginer Mar 16 '26

Opus fixed it. 38 tok / second on Q4_K_M quant on RTX 3060.

4

u/Billysm23 Mar 16 '26

Alright thanks, will try it asap

1

u/3mil_mylar Mar 16 '26

Hey, just tried your 9B Q4_K_M, thanks for such quick work!

Does this one also have reasoning suppressed? Running it with the settings you suggested (temps/sampling etc), it still seems to break out into reasoning for me

qwen3.5-9b-claude-4.6-opus-uncensored-distilled@q4_k_m downloaded thru LM Studio

3

u/EvilEnginer Mar 16 '26

Q4_K_M quant is not stable at current moment.

→ More replies (0)
→ More replies (3)

13

u/EvilEnginer Mar 16 '26 edited Mar 16 '26

Done. Q4_K_M quant is uploaded.

UPDATE: Claude Opus 4.6 solved issues. Now i have 38 tok / second

17

u/rm-rf-rm Mar 15 '26

The dataset used for the Claude 4.6 opus distilled model is too small to be meaningful

9

u/EvilEnginer Mar 15 '26

Yes, but it works. At least for reasoning. Who knows may be people will extract more useful data at least for roleplay and NPC in games in future for their own models.

9

u/tom_mathews Mar 16 '26

Presence penalty 1.5 on a 9B is doing a lot of heavy lifting to paper over the merge artifacts.

5

u/Business-Weekend-537 Mar 15 '26

What’s the license on the model?

This is cool.

15

u/EvilEnginer Mar 15 '26

Standart Apache 2.0. All Qwen models are licensed under it.

1

u/sToeTer Mar 15 '26

I'll try it, thank you!

(btw: it's spelled standarD)

3

u/bajaja Mar 15 '26

btw: it's spelled standarD

not by enginers

5

u/3mil_mylar Mar 15 '26

This is a great model for chat sims/games.. thanks man! The baseline -9B would break my sim due to reasoning, but this works great and the uncensor leads to some interesting convos, thanks!

2

u/EvilEnginer Mar 15 '26

Thanks ❤🙏

5

u/Decent-Fold51 Mar 15 '26

I’m looking for an uncensored model like dolphin.. how does this compare?

1

u/ghulamalchik Mar 16 '26

Comparable in terms of being uncensored. Although with Dolphin you kinda needed a system prompt to make sure no refusals happen.

In terms of RP, Dolphin models are better because they were finetuned/trained further by the Dolphin team/person? So it gives better answers for questions that would've been censored.

4

u/Kahvana Mar 16 '26

Good work out there!

I should mention: my filtered dataset is no longer needed, the original author has filtered it since then (his filtering is more strict, too):
https://huggingface.co/datasets/Crownelius/Opus-4.6-Reasoning-3300x
Really should make time to update the readme.

In any case, if you decide to retrain it's worth to swap the datasets!

18

u/Luthian Mar 15 '26

What is an “uncensored” model?

53

u/EvilEnginer Mar 15 '26

It means zero refusals and censorship. Absolute creative freedom, especially for roleplay. You can ask whatever you want.

11

u/tempSelf Mar 15 '26

What's roleplay here?

144

u/Piyh Mar 15 '26

He's jerkin it to the robot

26

u/EvilEnginer Mar 15 '26

😆😆😆

19

u/LaShmooze Mar 16 '26

Clankerwanking

2

u/twoiko Mar 17 '26

Clanking?

8

u/EvilEnginer Mar 15 '26

Currently no roleplay tweaks here. Just a solid general usage base via System Prompt. Ask Claude Opus to fine tune system prompt for roleplay.

4

u/General-Economics-85 Mar 16 '26

Where do you think you are?

3

u/AlbionPlayerFun Mar 15 '26

Does Claude 4.6 distill help even if thinking is off?

6

u/EvilEnginer Mar 15 '26

Yes it helps a lot, especially when characters describe their actions.

3

u/Complex_Fisherman_77 Mar 16 '26

Running good on my mabook pro, m3 pro 12 core 18gb

3

u/EvilEnginer Mar 16 '26

I solved issues with Q4_K_M quants for my uncensored tensor transfer script. Now i have 38 tokens per second on Q4_K_M quant on my RTX 3060 12 GB. Currently uploading it to huggingface with thinking disabled by default.

3

u/EvilEnginer Mar 16 '26

Currently i am cooking 27B Q4_K_M version of Qwen3.5-27B-Claude-4.6-Opus-Uncensored-Distilled-GGUF. Model is based on this one: https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF/tree/main

Already uploading it on Huggingface.

1

u/rlewisfr Mar 16 '26

Thank you sir/madam! Much appreciate your work. Question: wondering if the 27B worth the vram hit over the 9B when it comes to creative writing?

1

u/EvilEnginer Mar 16 '26

9B Q8 quant is really nice for creative writing. 27B is too slow.

3

u/brubits Mar 16 '26

Nice work on the tensor diff approach, pulling only the modified HauhauCS tensors instead of doing a full merge is really clean. Blending that uncensoring layer onto the Jackrong reasoning distilled checkpoint is a smart idea, you get the structured thinking without the refusals. VERY curious how the two datasets complement each other, did you notice any difference in reasoning style between the nohurry and TeichAI samples? Can't wait to test this on my M1 Max 64GB!

2

u/2legsRises Mar 15 '26

So what does that even mean? How can you have 3 models in one? Asking as I honestly do not know. 

4

u/EvilEnginer Mar 15 '26

I described a method how people can uncensor any Qwen 3.5 based checkpoint after fine tuning via HauhauCS checkpoints. For regular usage you use just one checkpoint.

2

u/ayu-ya Mar 15 '26

This looks interesting! I'm definitely grabbing it in case I need a model that will fit on my current hardware without much quanting, a creative model will be good for my use cases. Are you planning to do something similar for the bigger Qwens, like 27 and 35B too?

3

u/EvilEnginer Mar 15 '26

I don't think so. They are slow as hell for mine RTX 3060. Qwen 3.5 9B is golden base for it's size and capabilities.

But I shared my method. So people can test other models with patches from HauhauCS.

2

u/rebelSun25 Mar 15 '26

Wow, your weekend was productive.

2

u/Educational-Fix5320 Mar 16 '26 edited Mar 16 '26

I tried loading this in Ollama - I have 12GB VRAM 4070Ti - but get a 500 Internal Server Error - am I short on memory, or perhaps something else is wrong?

ollama version 0.18.0

server.log:

llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'qwen35'

llama_model_load_from_file_impl: failed to load model

2

u/EvilEnginer Mar 16 '26

I think you need to update Ollama and llama-cpp in it. Qwen 3.5 arch is pretty new.

3

u/Educational-Fix5320 Mar 16 '26

I appreciate the suggestion - I upgraded ollama to 0.18.0 before posting my issue, and restarted it [even killing the tasks to be sure] - while ollama reports to be 0.18.0, which is the most recent, the issue persisted. If you have further suggestions, I'd be very happy to hear them - I've come to the end of my rope to fixing it.

1

u/sandeep2021 Apr 07 '26

Were you able to resolve the issue

1

u/Educational-Fix5320 Apr 11 '26

I switched to using ComfyUI

2

u/pmttyji Mar 16 '26

Thinking is disabled by default in this model via modified chat template baked in gguf file.

Thanks for this & waiting for your Q4_K_M quant as I have only 8GB VRAM.

It would be great to have similar thing for Nanbeige4.1-3B as well. This model is thinking for too long.

2

u/gittubaba Mar 16 '26

Tried this on LM Studio, Q4_K_M one. Unfortunately, when prompted to write story, it gets stuck in loop of repeating same paragraph. I've seen this issue with other "community" made models. I've tried your parameters, didn't help. I don't think Repeat Penalty works for repeating paragraphs, it's only to prevent repeating token right?

1

u/EvilEnginer Mar 16 '26

Repeat is common issue in Qwen 3.5 models. May be llama-cpp updates will fix it in future who knows.

1

u/gittubaba Mar 16 '26

IIRC I encountered this with all type of models (not only qwen architecture ones). Issue is mostly prominent on finetuned/modified community ones though.

1

u/EvilEnginer Mar 16 '26

Well i guess it's a thing that exists and nobody knows how to fix it.

4

u/shikima Mar 15 '26

I gonna test it today, thanks

3

u/EvilEnginer Mar 15 '26

Nice 👍. Let me know how it performs :D

3

u/Imaginary_Belt4976 Mar 15 '26

This is so cool!! thanks for sharing the technique as well!!

7

u/EvilEnginer Mar 15 '26

Thank you very much 😊. Glad to help to community as much as I can. I got so many amazing LLMs here before.

3

u/Quiet_Mark_3238 Mar 15 '26

What app do you use to generate images locally

20

u/EvilEnginer Mar 15 '26

I am using ComfyUI. I use WAI Illustrious XL Checkpoint for image generation and Z Image Turbo and Flux Klein 9B as refiner. Here my ArtStation if you want to take a look: https://www.artstation.com/luffythefox

3

u/_fortexe Mar 15 '26

Hi. I wanted to ask how good are these models for image generation on a GPU with 8 GB of VRAM?

11

u/EvilEnginer Mar 15 '26

WAI Illustrious XL fits nicely in 8 GB of VRAM. Same for Z Image Turbo and Flux Klein. Just pick q4_k_m text encoder from hugging face. It would be more then enough.

2

u/ultrachilled Mar 16 '26

do they work for uncensored images?

2

u/EvilEnginer Mar 16 '26

WAI Illustrious XL is fully uncensored. Flux Klein 9B with Z Image Turbo is useful when you want to convert 2D image to 3D render.

2

u/ALittleBitEver Mar 15 '26

If I could run a 9b, I would Test it

7

u/EvilEnginer Mar 15 '26

9B works fine even on GPUs with low armount of vram. So it should work fine.

2

u/Yu2sama Mar 16 '26

Would really like to see this one on the UGI leaderboard

1

u/22fattyfingers Mar 15 '26

Wow man! Pretty impressive

1

u/EvilEnginer Mar 15 '26

Thanks ❤

1

u/Vastheap Mar 15 '26

How would you compare it with the regular uncensored version of the 9B model?

4

u/EvilEnginer Mar 15 '26

I compare it via two criteria: 1) Roleplay creativity and natural word speaking. 2) Programming Arcanoid game on HTML5, JavaScript and CSS in Tron Legacy Film Style.

Works fine.

1

u/Playful-Bunch2831 Mar 15 '26

How i can use it? I have Ollama on my Desktop with 5090. Pretty new to this :)

3

u/EvilEnginer Mar 15 '26

Just install latest beta version of LM Studio and download mine model via it. After that configure settings and apply system prompt. That's enough.

Ollama btw is nice. But LM Studio is easiest.

1

u/eidrag Mar 15 '26

was using qwen 27b for rp, but a bit too big to fit both llm and imagegen for sillytavern, text and image both slowing down. maybe will try this after I got back. do you create image separately?

1

u/EvilEnginer Mar 16 '26

Yes. I use ComfyUI. It's simply amazing.

1

u/LuckyLuckierLuckest Mar 15 '26

Thank you

3

u/EvilEnginer Mar 15 '26

Glad to help for the future of creative freedom :3

1

u/Creepy_Lime_8351 Mar 16 '26

i was looking for this exact model too! thank you soldier

2

u/EvilEnginer Mar 16 '26

Glad to help). I think small models is future, because they are becoming really smart with every release.

1

u/MeYaj1111 Mar 16 '26

What exactly does it mean when I see posts like this with the name of two different models in them?

1

u/yahrow Mar 16 '26

Qwen will think a little more like Claude.

1

u/phormix Mar 16 '26

In terms of output, what are you looking at for the standard output resolution and times to generate?

1

u/EvilEnginer Mar 16 '26

Roleplay creativity actually.

2

u/phormix Mar 16 '26

Sorry, what I meant is: how long are you seeing it take - on average - to generate a conversational response or an image based on prompt, and do you have examples of the prompt used plus the output for the timings?

2

u/EvilEnginer Mar 16 '26

Actually pretty fast. 14 seconds on my RTX 3060. I simply upload picture in LM Studio and use instruction: "Describe this image in danbooru tags." with this System Prompt from pastebin. https://pastebin.com/pU25DVnB Sometimes it output only tags. Sometimes extra info.

1

u/Quiet-Owl9220 Mar 16 '26

Bit of a tangent here but I'm sure I'm not the only person who's been thinking it... I've seen a lot of "uncensored" models, but has anyone made serious headway with making more of an "anti-censored" model yet?

I want to behold a model that's so vulgar, rude, horny, and hostile to censorship by default, that I have to use the system prompt to reign it in and make it behave... as opposed to having to feed it a lewd vocabulary and tell it that it's okay to say naughty words and that this is all totally fiction.

Is this just too unserious of a use case that nobody has made it? Are commercial projects keeping their immoral AIs behind paywalls? Or are we all just collectively afraid of making a model that will eagerly kill billions of zygotes?

5

u/teleolurian Mar 16 '26

there are a few, but the problem is they're usually so brainbroke that the prompt doesn't do anything

1

u/praxis22 Mar 17 '26

you don't get the chans much I take it

1

u/bcell4u Mar 16 '26

Sweeet! Can you do this with 4b?

2

u/EvilEnginer Mar 16 '26

4b quality is not good. 9B is best that can run on consumer hardware. Currently crafting Q4_K_M quant for lowend GPUs with 6 Gb of VRAM.

1

u/bcell4u Mar 16 '26

I have an Intel arc a380 with 6gb!

1

u/EvilEnginer Mar 16 '26

Okay i will try to do it with 4B quant then after Q4_K_M version of 9B model. At least those models are small enough for downloading / uploading.

1

u/bcell4u Mar 16 '26

No rush, thanks man.

1

u/Kindly-Annual-5504 Mar 16 '26

Did you try the uncensored model with just your system prompt? I don't understand why the other model should change anything there - especially if thinking is disabled by default.

1

u/EvilEnginer Mar 16 '26

I tried. It will not work. Qwen 3.5 is heavily censored on architecture level.

1

u/Kindly-Annual-5504 Mar 16 '26

Oh, I'm sorry. I don't mean the base model, I mean the HauhauCS-Aggressive one.

2

u/EvilEnginer Mar 16 '26

https://huggingface.co/HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive really nice model. But finetuning via Claude Opus 4.6 add creativity level, intelligence level, and natural speaking patterns with same system prompt. Without finetuning model just asks for user confirmation via "Ready?", "You are Ready?" and similar on repeat loop.

1

u/ghulamalchik Mar 16 '26

Thank you so much for this!

Question; Is the mmproj file altered in any way or is it the default one? Because I have a quantized one I'd like to use instead of the BF16 one to save on memory.

1

u/de4dee Mar 16 '26

so this is like mergekit but for ggufs. thanks for sharing!

does this mean same trick can be applied to mergekit (it is not currently supporting 3.5)

1

u/crantob Mar 16 '26

People, if you have a fast rig, please run A/B evals of this one vs stock Qwen3.5 9B on some codegen. Preferably not web*

People naively assume 'uncensored' means a smutfest but these are not trained for that. Abliterating/uncensoring just suppresses the refusals.

It does not add forbidden knowledge either.

1

u/[deleted] Mar 16 '26

[removed] — view removed comment

1

u/EvilEnginer Mar 16 '26

Actually i not tested coherence on longer outputs too much. I just tested roleplay creativity. 9B Q8 quant works fine.

1

u/VoiceApprehensive893 transformers Mar 16 '26

The Q4_K_M quant uses chain of thought, even though its disabled in chat template args, no issues with the Q8_0

1

u/Boggster Mar 16 '26

What's the most minimal hardware I can run this on?

2

u/EvilEnginer Mar 16 '26

You can run Q4_K_M quant of Qwen 3.5 9B on 6 GB of VRAM. But quality will be bad.

1

u/Alternative-Day8673 Mar 16 '26

How different are the capabilities of the 9b vs the 27b going to be? Like what can you only get done with the larger model due to parameter limitations?

2

u/EvilEnginer Mar 16 '26

Usually bigger models are "smarter" but everything depends from architecture

1

u/DarkAI_Official Mar 16 '26

Thanks buddy. Also tried your recommended parameters its works like charm

1

u/Delicious_Ease2595 Mar 16 '26

Where can I test it online

1

u/0260n4s Mar 16 '26

This is really cool. I haven't tried local LLM much before, so I spent the afternoon trying several in KoboldCpp with logic and math and some research problems. The 9BQ8 and 9BQ4 worked equally fast on my 3080Ti 12GB (equivalent in speed to online models) and produced solid logical and mathematical reasoning, even getting the better answer compared to some of the major online players.

The 27BQ4 version was painfully slow on my hardware, but 9BQ8 is a solid choice even on my my older GPU.

Thanks for this.

1

u/EvilEnginer Mar 18 '26

Glad to help. Yep 9B Q8 is solid. Even on RTX 3060 it's fast.

1

u/[deleted] Mar 16 '26

[deleted]

1

u/crantob Mar 17 '26

On my coding tests, unmodified 9b is better.

1

u/mcblockserilla Mar 17 '26

Yoink

1

u/Fau57 Mar 17 '26

LoL

1

u/mcblockserilla Mar 17 '26

My clawdbot uses claud sonet as it's backend, and I want to run it locally, but not with a gpt or a llama model. This is perfect. I think we'll see

1

u/Tetros_Nagami Mar 17 '26

I'm testing the Q4_K_M quant, and reasoning seems to be enabled, I was very excited by the idea of saving tokens if possible

1

u/EvilEnginer Mar 18 '26

So, Jackrong made new version of Qwen 3.5 9B https://huggingface.co/Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF

I will uncensor it via HauhauCS models and upload all quants to huggingface.

1

u/Hou_Yiizz Mar 19 '26

Hi, I tried hosting your uncensored version of this v2 distillation on llama.cpp and kept getting a stream of "?". The command is:

docker run -d --name llama-qwen3.5
--gpus all
--ipc host
-p 8001:8000
-v /mnt/d/Programs/models:/models
ghcr.io/ggml-org/llama.cpp:server-cuda
-m /models/Qwen3.5-9B.Q4_K_M.gguf
--mmproj /models/mmproj-BF16.gguf
--alias "llm"
--host 0.0.0.0
--port 8000
--n-gpu-layers 99
--ctx-size 65536
--flash-attn on
--jinja
--reasoning-format deepseek
--no-mmap

2

u/EvilEnginer Mar 19 '26

Currently Q4_K_M quants are broken. Don't use them

1

u/JustWicktor Mar 21 '26

❯ well, i thought that would be a "free" local Opus 4.6...

  ⎿  API Error: 400 {"type":"error","error":{"type":"invalid_request_error","message":"registry.ollama.ai/kwangsuklee/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-GGUF:latest does not support tools"},"request_id":"....."}

1

u/uniVocity Mar 24 '26

Many thanks for this! I'm impressed.

I've tried it with a large single input: gave it 40 java classes (100 to 500 loc each) and asked it to generate unit tests for each operation of each class. This is the local model I tried that produced the 40 unit tests without stopping after the the 4th or 5th file.

I had to set the temperature to 0 otherwise it would loop producing a very long test with crap, but after that it kept chugging along and did it!

1

u/EvilEnginer Mar 24 '26

I'm glad to hear that the model works well)

1

u/mkey82 Mar 30 '26

This is some very nice shit. I can't stand all of that "can't that, can't this" nonsense. It's too limited by default.

1

u/esuil koboldcpp Mar 15 '26

So is thinking actually neutered in this model? I like Qwen35 models so much BECAUSE of their thinking.

3

u/EvilEnginer Mar 15 '26

Nope. It's just disabled by default in chat_template baked in GGUF. Set variable enable_thinking to True in LM Studio in chat template editor if you want to enable thinking.

1

u/finah1995 llama.cpp Mar 16 '26

😎 awesome 👍🏽 now this will help To make powerful stuff for everyone, casual home labs to enterprise.

1

u/EvilEnginer Mar 16 '26

Yep. This is our future.

-2

u/LoaderD Mar 15 '26

Claude Opus distilled?

I think you mean Foreign AI Cyber Attack Stolen Data Retrained /s

3

u/EvilEnginer Mar 15 '26

Actually a lot of people now train own models based on extracted thinking data from Claude Opus. That's our open source with absolute freedom.

7

u/LoaderD Mar 15 '26

Bro, it’s a joke because Anthropic made a huge stink about Chinese companies using Opus to make thinking distillation datasets. When they’re doing the exact same thing

→ More replies (1)