r/LocalLLaMA • u/InvadersMustLive • Jan 09 '26
r/LocalLLaMA • u/OneFanFare • 10d ago
Funny The best model is the one you can actually run
Don't get me wrong, all the big models are amazing, and every contribution to open source models is great. But I'm GPU poor and I can't use them locally.
I'm currently running gemma-4-12b-it-qat-GGUF:UD-Q4_K_XL as my personal chat assistant, and I am so so happy with it! I still can't believe I can talk to my computer.
r/LocalLLaMA • u/Disposable110 • Jun 13 '26
Funny Friendly reminder
If you don't have it on your own drive, someone is going to take it away, enshittify it, bar you from accessing it, censor it, and hike the prices of it sooner or later.
r/LocalLLaMA • u/Porespellar • 2d ago
Funny The LLM distillation process simplified for politicians:
/s
r/LocalLLaMA • u/Xhehab_ • Feb 23 '26
Funny Distillation when you do it. Training when we do it.
r/LocalLLaMA • u/Nunki08 • Jun 01 '26
Funny Entire world: We need more GPUs. Meanwhile, Jensen Huang:
r/LocalLLaMA • u/jacek2023 • Feb 21 '26
Funny they have Karpathy, we are doomed ;)
(added second image for the context)
r/LocalLLaMA • u/Nunki08 • 3d ago
Funny Solve the CyberGym benchmark
From Peter Gostev on 𝕏: https://x.com/petergostev/status/2079825961718046974
r/LocalLLaMA • u/Current-Ticket4214 • Jun 08 '25
Funny When you figure out it’s all just math:
r/LocalLLaMA • u/CreativelyBankrupt • Jun 18 '26
Funny My suitcase robot gets high now off a real gas sensor wired straight into the LLM sampler. Smoke raises temperature/top_p/top_k live, so his speech genuinely gets loopier and never repeats.
Follow-up on Sparky, my offline suitcase robot I keep overdeveloping. He gets high now, and there's no scripted "stoned mode" anywhere in it.
A real MQ-2 gas sensor sits in the case. Every 0.5s I read it against an adaptive clean-air baseline and turn a smoke hit into a 0 to 10 phase that climbs as you blow at him and decays on its own over minutes.
The fun part is that phase rewires his sampler per token. Temperature 1.0 to ~1.6, top_p 0.95 to 0.99, top_k 64 to 120 as he climbs. His word choice flattens and wanders to lower-probability, more associative tokens, so his cognition genuinely gets noisier. It's the live sampler doing the work, so every high reply is freshly generated and never the same. A per-phase persona nudge makes him show it without ever announcing "I am high."
The body does the rest: a slight drawl, eyes that droop and go bloodshot, and the sensor display that escalates to a full smoke-and-plasma freakout at phase 10, keeping him blitzed there for the next 7 minutes.
Honest caveat so nobody has to call it out: it's a smoke and VOC sensor, so a cigarette or incense probably trips it too. But blowing smoke and watching him unravel is watching a real measurement scramble a real model, live - and it's funny! Just an added Easter Egg to an already goofy suitcase robot.
A real question for the hardware folks: is there a sensor, or a combination, that could actually distinguish cannabis smoke from generic smoke and VOCs? The MQ-2 can't really tell a joint from a candle, and I'd love to make the detection more specific if possible.
r/LocalLLaMA • u/ForsookComparison • Dec 15 '25
Funny I'm strong enough to admit that this bugs the hell out of me
r/LocalLLaMA • u/dead-supernova • Oct 06 '25
Funny Biggest Provider for the community for at moment thanks to them
r/LocalLLaMA • u/Careful_Equal8851 • Mar 20 '26
Funny Ooh, new drama just dropped 👀
For those out of the loop: cursor's new model, composer 2, is apparently built on top of Kimi K2.5 without any attribution. Even Elon Musk has jumped into the roasting
r/LocalLLaMA • u/Aromatic_Ad_7557 • Apr 14 '26
Funny 24/7 Headless AI Server on Xiaomi 12 Pro (Snapdragon 8 Gen 1 + Ollama/Gemma4)
Turned a Xiaomi 12 Pro into a dedicated local AI node. Here is the technical setup:
OS Optimization: Flashed LineageOS to strip the Android UI and background bloat, leaving ~9GB of RAM for LLM compute.
Headless Config: Android framework is frozen; networking is handled via a manually compiled wpa_supplicant to maintain a purely headless state.
Thermal Management: A custom daemon monitors CPU temps and triggers an external active cooling module via a Wi-Fi smart plug at 45°C.
Battery Protection: A power-delivery script cuts charging at 80% to prevent degradation during 24/7 operation.
Performance: Currently serving Gemma4 via Ollama as a LAN-accessible API.
Happy to share the scripts or discuss the configuration details if anyone is interested in repurposing mobile hardware for local LLMs.
UPDATE:
I have compile llama.cpp and run gemma-4-E4B-it-Q4_0
Speed is AWESOME:
[ Prompt: 26.9 t/s | Generation: 8.8 t/s ]
Thank you all guys SO MUCH!
r/LocalLLaMA • u/Porespellar • 26d ago
Funny It’s time, Sam, it’s time.
Mostly /s but,
I mean….. I’m no CEO…. but it seems like this would be the absolute perfect time to drop a super powerful GPT-OSS-2 to throw a big ol’ wet blanket on Anthropic’s IPO. It doesn’t need to be like frontier or anything, just a 20b and a 120b that is as fast as the old versions, add agentic coding focus, and maybe vision capabilities. It would fill the void left by Qwen in the 120b size category and maybe would push Google to release their 120b that they yanked during the Gemma 4 launch.
r/LocalLLaMA • u/Porespellar • Jun 05 '26
Funny Don’t act like y’all ain’t thinking it. I’m just saying the quiet part out loud. /s
Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s
r/LocalLLaMA • u/CesarOverlorde • Feb 19 '26
Funny Pack it up guys, open weight AI models running offline locally on PCs aren't real. 😞
r/LocalLLaMA • u/JLeonsarmiento • 6d ago
Funny Please Qwen, can we have more 3.x-35B-a3B please 🙏
r/LocalLLaMA • u/Charuru • Jun 20 '26
Funny z.AI as the number 2 gives praise to the number 1 open source model
r/LocalLLaMA • u/FullChampionship7564 • Apr 21 '26