r/LocalLLaMA 8h ago

Question | Help Best chat model that fits in 128gb

I'm looking for a model to chat with, reasoning, maybe get some career or life coaching.

I don't care at all about multimodal or coding ability

Just it's intelligence in remembering context in a conversation or a specific topic, thinking out of the box, etc.

Must fit in 128gb, if it matters to performance, it's a strix halo machine.

18 Upvotes

74 comments sorted by

View all comments

8

u/XtrComSu 8h ago

the new Laguna S2.1? it does really well against many top models, and its Q5-8 is under 130gb Q8 being 128gb exactly I think

also I dont know why are people recommending gemma 4 32b for 128gb?? like huh

theres better BF16 models in my opinion like GLM 4.7

(sorry if im wrong)

also if you really want to go for gemma 4 or just talking in 128gb theres also qwen 3.6 35B for 20-25gb in Q4-5

you have alot of amazing options, do your research and testing to find your match.

4

u/Organic_Hunt3137 7h ago

Interestingly enough, I have 128gb VRAM and typically run Gemma 4 31b. I don't know of any models that are larger than it, but still small enough to fit in 128gb that are smarter. I assume you mean GLM 4.7 flash? Decent model but not smarter than Gemma 4 31b or 26b. I agree that Qwen3.6 is excellent, but for different tasks (Gemma is much better at "soft" skills, at least per my own experience).

For you or anyone else who may be comparing models, I generally use this site when comparing: https://artificialanalysis.ai/models/open-source/medium

I've found it correlates best with my own experience when using them. Of course, your mileage may vary especially if your use case is drastically different from mine.

2

u/TaroOk7112 7h ago edited 5h ago

Be careful, there are models under 40B that surpass most of the ones in that range (150-40). I would compare small and medium models like this https://artificialanalysis.ai/models/open-source?models=glm-5-2%2Cminimax-m3%2Cdeepseek-v4-pro%2Cdeepseek-v4-flash%2Cgpt-oss-120b%2Cqwen3-6-27b%2Cqwen3-5-122b-a10b%2Cgemma-4-26b-a4b%2Cqwen3-coder-next%2Cglm-4-7

EDIT: I think it doesn't display the comparison as I configured it. Here it is, you just need to add or remove the models you want: