VRAM for local LLMs: why memory bandwidth sets your tokens per second

Running local Large Language Models (LLMs) effectively requires more than just sufficient VRAM capacity; memory bandwidth is the critical factor determining performance. Because LLMs must read their entire weight set from memory for every generated token, bandwidth bottlenecks often dictate the speed of inference. When models exceed VRAM capacity and spill into system RAM via PCIe, performance drops by up to 20x. The author emphasizes that for local inference, hardware choices should prioritize high memory bandwidth over raw compute power. A used RTX 3090 with 24GB of VRAM is highlighted as a cost-effective solution for running 32B models with sufficient context. The article also provides a guide on estimating VRAM needs based on model size and KV cache requirements, warning that 'almost fitting' a model into VRAM results in severe performance degradation. Ultimately, the author suggests balancing local hardware for daily tasks with cloud APIs for complex, multi-step reasoning.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
English-language SEO tools penalize Russian-language texts
Researchers have discovered that popular SEO tools for evaluating citation and content quality in the AI-search era exhibit a bias against Russian-lan…
Dots, GPT-6.1 Sol, and a $500 Plan: Key Highlights from OpenAI DevDay 2026
At the DevDay 2026 conference in San Francisco, OpenAI announced over twenty new products and updates. The highlight is 'Dots'—autonomous agents with…
14 Personal AI Agents in 2026: A Technical Guide to Architecture, Memory, Tools & Autonomy
The landscape of personal AI is undergoing a fundamental shift, moving from simple chatbot interfaces to autonomous agents capable of delegation, stat…



