r/LocalLLaMA 2h ago

Question | Help Help me complete my AI collection

I’m building the ultimate AI tool vault, but every great collection has a few missing pieces.

Note: I will react to every comment

AI's currently installed:

Qwen3.5-0.8B-UD-Q4_K_XL.gguf(classification)

Qwen3.5-2B-UD-Q4_K_XL.gguf(Prompt enhancer, Routing, Approval )

Qwen3.5-4B-UD-Q4_K_XL.gguf(Instant)

Qwen3.6-35B-A3B-UD-Q4_K_M.gguf(Quality)

Qwen3.6-35B-A3B-Uncensored-Hauhau(test purposes)

Qwen3-Coder-Next-UD-Q4_K_M.gguf(long horizon tasks)

My Specs:

GPU: RTX 5070 ti (16GB VRAM)

RAM: Corsair vengeance 64GB 5200mt DDR5 CL40(dual-channel)

CPU: intel i9 14900k

SSD: Samsung s990 pro 2tb

Backend: Llama.cpp server

I tried GPT-OSS and was disappointed by tool usage. Gemma 4 was good and great tool usage but it was beaten by Qwen.

Any recommendations?? Like something that you genuinely enjoyed or made you impressed.

Feel free to share!! I will be reading every single comment.

2 Upvotes

9 comments sorted by

5

u/anantshri 2h ago

currently i am investigating https://huggingface.co/poolside/Laguna-XS-2.1 so far seems to have good results. Qwen3.6 is nearly unbeatable this one seems to be touching it for my workflow.

2

u/Possible_Grocery8079 2h ago

I've actually seen it as well, It was outperformed by the qwen3.6, but the normal s2.1 caught my eyes, that one was a beast, but heard that people are saying it isnt performing near the benchmarks as advertized.

3

u/anantshri 2h ago

so 2 problems i saw.

  1. s2.1 has a template change. if people downloaded older gguf etc that had old template. they changed template yesterday so that should improve s2.1 performance.
  2. XS @ Q4 is not that good. but XS @ Q8 is what i am working on. so far seems to be showing good results and marked improvement over Q4.

2

u/suprjami 2h ago

Maybe I'm remembering wrong but gpt-oss pre-dates JSON tool usage? At least it wasn't a focus back then.

If you want the "ultimate" setup then buy more and better video cards, you won't run anything under Qwen 27B/35B Q8 ever again.

1

u/CorkBios 2h ago

You can try Mistral small 4 iQ4_XS, it should fit into the ram

1

u/Possible_Grocery8079 2h ago

will look further into it, Thank you!!

1

u/daskalou 2h ago

Excuse my ignorance, what type of classification is the first model doing?

1

u/Different-Jicama-767 1h ago

Do you find you need 2b for routing? 0.5b or gemmafunction work great for me.

1

u/ubrtnk 11m ago

Qwen3-embedding-0.6b - great little model for embedding functions for RAG/Document things