Technologies
Back
Artificial Intelligence & Machine Learning

VRAM for local LLMs: why memory bandwidth sets your tokens per second

Dev.to
Advertisement468 × 90
VRAM for local LLMs: why memory bandwidth sets your tokens per second

Running local Large Language Models (LLMs) effectively requires more than just sufficient VRAM capacity; memory bandwidth is the critical factor determining performance. Because LLMs must read their entire weight set from memory for every generated token, bandwidth bottlenecks often dictate the speed of inference. When models exceed VRAM capacity and spill into system RAM via PCIe, performance drops by up to 20x. The author emphasizes that for local inference, hardware choices should prioritize high memory bandwidth over raw compute power. A used RTX 3090 with 24GB of VRAM is highlighted as a cost-effective solution for running 32B models with sufficient context. The article also provides a guide on estimating VRAM needs based on model size and KV cache requirements, warning that 'almost fitting' a model into VRAM results in severe performance degradation. Ultimately, the author suggests balancing local hardware for daily tasks with cloud APIs for complex, multi-step reasoning.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250