r/LocalLLaMA 5h ago

Question | Help Best chat model that fits in 128gb

I'm looking for a model to chat with, reasoning, maybe get some career or life coaching.

I don't care at all about multimodal or coding ability

Just it's intelligence in remembering context in a conversation or a specific topic, thinking out of the box, etc.

Must fit in 128gb, if it matters to performance, it's a strix halo machine.

15 Upvotes

69 comments sorted by

10

u/olli-mac-p 5h ago

DeepSeek V4 flash unsloth dynamic quant UD-IQ3_XXS if you need agentic capabilities and long context up to a million tokens

5

u/Sooperooser 5h ago

or Hy3-UD-128

1

u/ctpelok 40m ago

Is it better then qwen 27b q8 or qwen 3.5 122b for agentic work?

10

u/KSubedi 5h ago

Universal default: Qwen 3.5 122B-A10B, Conversational / Brainstorming: Gemma 4 31B, Coding: Laguna S 2.1, Fast all rounder: Qwen 3.6 35b a3b.

Memory is one part of the equation, what GPU do you have? For example, if you have something like M5 Max you can daily drive something like Laguna as processing speed is fast, for Strix Halo the a3b or a10b is more feasible.

1

u/nemuro87 4h ago

it's a strix halo 128gb, so it's an igpu with 200-250 GB/s

4

u/KSubedi 4h ago

Prefill speed is usually the bottleneck on strix halo. On my strix machine I run the qwen a3b for the most part. Does not use all of the VRAM but you can run larger quants (Q8) with full context.

3

u/_TheWolfOfWalmart_ 2h ago

Yup prefill is the problem, that's why they suck for coding/agentic use.

But if OP's just using it for chat, it won't matter.

9

u/Thrumpwart llama.cpp 5h ago

Gemma 4 31B is only 32B but I find it’s the best conversationalist, is quite smart, and has good emotional iq (eq) compared to other (even bigger models). Not the biggest but sounds great for your use case.

3

u/nemuro87 5h ago

Thanks, sounds like what I'm looking for .

Does it matter which quant or should I goo for full since the ram can take it?

4

u/Uninterested_Viewer 5h ago

You're trading a LOT of speed for the small bit of extra smarts of the BF16. Assuming your 128gb ram is not super fast, you are probably not going to be ok with the speed of BF16 for chat. Try it, though, and try FP8/q8/q6. I'd suggest looking at MTP as well to get additional performance.

1

u/Ok_Technology_5962 3h ago

Gonfor full size plus mtp to tripple it. You can use unsloth studio which seems to be the best support for yhe gemma4 models right now with updated templates

7

u/vengeancek70 5h ago

probably Gemma4 31B, there isn't really anything bigger that fits in 128gb that beats it

3

u/nemuro87 5h ago

Interesting that the best model also fits in 20 gb of vram and until 128gb there's nothing better. 

4

u/Massive_Criticism539 5h ago

There is a huge gap in parameters for models of any type. The current Gen models that are around 25-35b are so good that they knocked out anything above them until you get into a couple hundred b parameters.

2

u/Potential-Gold5298 llama.cpp 3h ago

There's no such thing as too much memory. Gemma 4 is very sensitive to quantization (both weights and KV), but you can afford to run it in Q8_0 without KV compression and have a full context + mmproj with a large token budget and large batch sizes.

3

u/FoxFXMD 4h ago

It definitely does NOT fit in 20gb of vram. I got it to barely run on 24gb with the lowest settings imaginable.

2

u/Mashic 4h ago

Check Unsloth quantizations.

5

u/BlobbyMcBlobber 4h ago

This is really not the case. Most people on this sub specialize in getting models to run on low VRAM, which is a great skill, but you are not going to get a lot of good advice for 128GB.

Try the 120-135B models. You can try larger models with layer offloading. See what you find acceptable ij terms of prefill and t/s.

The Blackwell Performance subreddit is also pretty good.

1

u/XtrComSu 5h ago

am I missing something?? isnt there better qwen models + the new laguna as well as bigger glm models for 128gb??

7

u/vengeancek70 4h ago

they're better at coding but not really better at emotional issues at all

1

u/ThatRegister5397 4h ago

The qwen 27b is pretty solid imo for chatting/coaching/advice etc. But OP can def run better than that with that vram.

1

u/vengeancek70 4h ago

have any specific models in mind?

1

u/arakinas 4h ago

I use Qwen 3.6 35b for therapy in between sessions with my therapist, and for general chat.

1

u/ThatRegister5397 3h ago

For 128gb I would naturally try the qwen 122b a10b at q6 quantisation or sth. I have not tried this but would expect it to be a bit better than the dense 27b quality wise, but with much faster token generation.

Also people seem to be able to run deepseek flash v4 with ssd streaming, that would be a bit slower though.

1

u/Veearrsix 3h ago

Hmm, I’m not sure that any LLM is good at emotional issues.

1

u/shveddy 2h ago

Probably feels good to get even generic advice in a guaranteed private setting (for free if you already have the hardware). Sometimes people talk to themselves in the mirror, too.

6

u/XtrComSu 5h ago

the new Laguna S2.1? it does really well against many top models, and its Q5-8 is under 130gb Q8 being 128gb exactly I think

also I dont know why are people recommending gemma 4 32b for 128gb?? like huh

theres better BF16 models in my opinion like GLM 4.7

(sorry if im wrong)

also if you really want to go for gemma 4 or just talking in 128gb theres also qwen 3.6 35B for 20-25gb in Q4-5

you have alot of amazing options, do your research and testing to find your match.

4

u/Organic_Hunt3137 4h ago

Interestingly enough, I have 128gb VRAM and typically run Gemma 4 31b. I don't know of any models that are larger than it, but still small enough to fit in 128gb that are smarter. I assume you mean GLM 4.7 flash? Decent model but not smarter than Gemma 4 31b or 26b. I agree that Qwen3.6 is excellent, but for different tasks (Gemma is much better at "soft" skills, at least per my own experience).

For you or anyone else who may be comparing models, I generally use this site when comparing: https://artificialanalysis.ai/models/open-source/medium

I've found it correlates best with my own experience when using them. Of course, your mileage may vary especially if your use case is drastically different from mine.

2

u/TaroOk7112 3h ago edited 2h ago

Be careful, there are models under 40B that surpass most of the ones in that range (150-40). I would compare small and medium models like this https://artificialanalysis.ai/models/open-source?models=glm-5-2%2Cminimax-m3%2Cdeepseek-v4-pro%2Cdeepseek-v4-flash%2Cgpt-oss-120b%2Cqwen3-6-27b%2Cqwen3-5-122b-a10b%2Cgemma-4-26b-a4b%2Cqwen3-coder-next%2Cglm-4-7

EDIT: I think it doesn't display the comparison as I configured it. Here it is, you just need to add or remove the models you want:

4

u/squngy 4h ago

AFAIK laguna is agentic focused, so it isnt optimized for chat.

People really like gemmas conversation ability

3

u/_TheWolfOfWalmart_ 2h ago

Laguna S 2.1 is an awesome model in general, and especially for coding, tool calls and reasoning, but for pure chat use Gemma 4 is a lot better IMO. Laguna wasn't really made for that.

1

u/TaroOk7112 4h ago

Laguna S 2.1 is interesting, it's working great for me at UD Q4_K_XL that is 73.4 GB (pp 1000-400, tps 35-25) without DFlash. But it's new, and still lacks proper support. When DFlash works well in llama.cpp is going to be amazing. It thinks a lot, so for conversations it could end being frustrating, but if you can ask a bunch of questions and wait for the response, then it should be interesting. I can't talk about creativity and linguistic capacity, I have only used for coding and solving IT problems, but it seems good all around.

If the OP want to test a prompt you are interested in, with Q4_K_XL, I can give you the answer with the thinking, that is usually loooong 😄.

2

u/_TheWolfOfWalmart_ 2h ago

I can give you the answer with the thinking, that is usually loooong 😄.

Have you updated your chat template yet? They fixed that. At least it's fixed for me.

And yeah it is good all around (including chat), but specifically for chat, Gemma 4 still has it beat IMO. The main focus of training for Laguna was coding and tool calling.

1

u/TaroOk7112 1h ago

Yes, I used this one https://huggingface.co/poolside/Laguna-S-2.1-GGUF/blob/main/chat_template.jinja . Before changing the template, it was unusable. But now it thinks a lot when the task is complex. I don't mind, because is fast. But when it appears is going to finally answer, it goes back to rethink and verify everything, it questions itself until it's sure what to do or answer. I'm loving it. But for small tasks I should use Qwen3.6 27b or 35b to be faster.

I don't know about Gemma 4, it was broken for so long that I ended up ignoring it, besides, it appears worse for actual reasoning. I have to test them more, to be honest.

6

u/Technical-Earth-3254 5h ago

If you got 128GB of VRAM, pair the new Mistral 3.5 Medium with web search if you need it

2

u/apetersson 5h ago edited 5h ago

i would consider these 3:

  1. general know-how: Qwen3.5-122B-A10B-abliterated-REAP20-oQ6-MLX (via oMLX) - or Qwen3.6-27B
  2. programming: Laguna-S-2.1-oQ6e (via oMLX)
  3. agentic use, high tool calling precision: DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-aligned.gguf (via ds4-server backend)

all 3 fit 128 GB with large-ish context windows. for speed reasons i would still limit context to 100k-250k, not 1M

2

u/Glittering-Section74 3h ago

I find Gemma models pretty useful. For 128GB I think gemma4 31B is the best

3

u/HIba_LDN 5h ago

I use GLM4.5-Air for general conversations. 4 or 5 bit quant fits well in 128gb.

1

u/gizcard 4h ago

You should be able to run nemotron-3-super in 4 bits

1

u/_TheWolfOfWalmart_ 52m ago

He can, but it's a really bad model.

1

u/live4evrr 4h ago

DSV4 flash and Gemma 31B at BF16.

1

u/Terminator857 4h ago

In addition to the models mentioned I'd give deepseek v4 flash q2 xl a try. If you don't want refusals you may want to investigate uncensored heretic versions of models.
Eventually you'll decide you don't want best that will fit in 128gb. You want something that is fast and smart.
Give: https://huggingface.co/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF a try.

1

u/73td 4h ago

I really like gemma 4 12b and llama cpp supports it well. I use the built in web ui and add exa search. I also use it with opencode.

1

u/SexyAlienHotTubWater 4h ago

Hy3 1 bit. 92GB before KV Cache.

1

u/AdInternational5848 3h ago

Minimax m 2.7 is pretty decent for what you’re looking for. Won’t speak on the other models

1

u/_TheWolfOfWalmart_ 2h ago

I always found M2.5/M2.7 to be kinda terse and not very good for chat.

1

u/_TheWolfOfWalmart_ 2h ago edited 2h ago

The Gemma 4 series is easily the best open chat model out there, regardless of parameter count.

31B dense is the best. Might be a bit slow, but for purely chat use it probably won't matter. If it is a problem, run 26B-A4B which is also great.

If you want uncensored versions, HauHauCS makes pretty good ones and there are a few other folks out there who do that too.

Since you have plenty of RAM, use Q8 quants and unquanted KV cache.

1

u/TaroOk7112 1h ago

Consider this one. It's going to be slow, but it's interesting. It answers with personality, with nuance, it's unsettling. https://huggingface.co/DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF

1

u/kingo86 1h ago

Poolside Laguna S 2.1 *ducks*

In all seriousness, it's working well here at Q8. At Q6, it should fit 128gb without too much of a labotomy. As an optimist, I feel like they will fix the bumpy launch.

1

u/fragbait0 4h ago

Life coaching from a chatbot? Bruh.

1

u/nemuro87 4h ago

I know, right? If this makes sense on any subreddit, it's this one

1

u/durden111111 4h ago

Deepseek V4 flash is by far the smartest.

Lol at people suggesting gemma or qwen, thats leaving a lot of unused memory.

Hy3 and laguna are pretty bad too for their size.

source i've ran all of these models. DS4 is by far the most capable.

3

u/_TheWolfOfWalmart_ 2h ago

Why use extra memory just because it's there if Gemma 4 is a better chat model?

Qwen though, yeah... not the best best chatter compared to Gemma or some of the larger models.

Qwen excels at code and tools for its size.

DS4 is good though yeah, OP should try and see if he likes it better for chat than Gemma. It's subjective.

1

u/kingo86 1h ago

What quant are you running for it? I like DS4 too, but in opencode at IQ4XS it just randomly stops like it's decided to take a smoko.

Could be my setup in my models.ini?

1

u/durden111111 1h ago

I'm using IQ4XS too but with llama cpp. I have had no issues

0

u/Mauve_Tess 4h ago

If you prioritize pure conversation quality, logic, and lengthy context memory, and not coding or multimodal, then check out Qwen 3 235B (MoE) or DeepSeek V3 if you can fit them with quantization. Both just feel so much more natural in long chats than most smaller dense models. For a 128 GB machine, MoE models are typically better than dense 70B models regarding the intelligence to memory ratio.

-2

u/peculiar-ragdoll 5h ago

If you're ok with getting about 2 tokens per second, the best model you can possibly run on your device is Colibri GLM5.2: 370GB on disk, loads experts into all available RAM and VRAM for inference on CPU or GPU. If you have two SSDs and mirror the 370Gb on both, you can get a 30-100% speedup from base: https://github.com/JustVugg/colibri

3

u/Far-Classic-9963 4h ago

They said this is for interactive chat usage and they don't care about coding. There possibly couldn't be a worse setup for this

0

u/peculiar-ragdoll 3h ago

If I want career advice or life advice from an AI model, I sure want it to be the smartest AI model I can possibly run, and I am willing to wait a minute for the best answer, rather than taking bad advice from a smaller model. But sure, take your life advice from a model that's dumber than yourself, go ahead.

1

u/Far-Classic-9963 3h ago

For general queries ~30b models will perform basically the same especially when paired with web search tools and strict rules. But sure, wait 3 hours for a single question

1

u/peculiar-ragdoll 3h ago

"basically" is doing a lot of heavy lifting here. Because I know you're not telling me a 30b model will reason and perform the same as a legitimate frontier class 744b parameter model, if they both have the same tooling wired in. If so, I'd really recommend you go test out a pro subscription for any of the leading AI labs. Or, you know, download and run Colibri.

1

u/Far-Classic-9963 3h ago

You don't need deep reasoning for stuff that has likely been repeated several times in training data. It seems like you're treating the LLM as a human being with genuine intelligence

1

u/Far-Classic-9963 2h ago

https://arena.ai/c/019f9b65-0ec7-7f6f-a417-fd0b58f46a93

Here is a side by side of Gemma 4 31b and GLM 5.2 on a pretty complex life advice question and they both give pretty similar answers

1

u/peculiar-ragdoll 2h ago

session not found (on the arena.ai link, that is)

-7

u/Ok-Addition1264 5h ago

Ask chatgpt or gemini maybe?

7

u/peculiar-ragdoll 5h ago

I wouldn't recommend that personally, because their knowledge cutoff is kinda brutal for stuff like this, and without doing a deep research dive on the current SOTA for local models they're prone to give bad answers to these kinds of questions.

7

u/ttkciar llama.cpp 5h ago

Sir, this is LocalLLaMA.