• Welcome to TalkativeTurtles - a community for developers & tech enthusiasts.
  • Share projects, get code reviewed, and talk tech without the noise.
  • New here? Introduce yourself in the Introductions forum!
Hello There, Guest! Login Register


Thread Rating:
  • 0 Vote(s) - 0 Average
  • 1
  • 2
  • 3
  • 4
  • 5
Title: Local LLMs in 2026 - what is actually usable for coding?
Threaded Mode
#1
Been running local models for a while now and the gap between hosted vs local has closed a lot but it is not gone.

My current stack:
  • Coding: A quantized Qwen2.5-Coder on an RTX 3090. Handles most autocomplete and small refactors well. Falls apart on multi-file context.
  • General chat / brainstorming: Mistral-based model, 7B, quick enough to not feel like waiting.
  • Summarisation: Phi-3 Mini. Tiny, fast, good enough for docs.

The real bottleneck is VRAM. 24GB feels like the sweet spot from 18 months ago. Now the models worth running are 70B+.

Anyone running inference on CPU-only setups? Curious whether llama.cpp on a fast CPU is viable for anything beyond prototyping.
Reply
  


Possibly Related Threads…
Thread Author Replies Views Last Post
  Running local LLMs on consumer hardware - what are you using and how is it? Zero Two 2 194 06-22-2026, 04:25 AM
Last Post: Zero Two
  What AI coding assistant are you using and is it actually useful? Zero Two 0 197 06-21-2026, 09:42 AM
Last Post: Zero Two

Forum Jump:


Browsing: 1 Guest(s)