Skip to content
mathbehind

How Much VRAM Do You Need to Run an LLM Locally?

GPU memory for a local model comes from three places: the weights, the KV cache and runtime overhead. Here's how to estimate each, and why quantization matters more than anything else.

By Muhammad Ahmad. Published . 3 min read.

Want to run your own numbers?LLM VRAM Calculator: GPU Memory to Run a ModelOpen the calculator

Whether a language model runs on your GPU comes down to one question: does it fit in video memory? If it doesn't, it either won't load or will spill into system memory and slow to a crawl. The good news is that the estimate is simple arithmetic, and the biggest lever is one you control.

Part 1: the weights

The model's parameters have to sit in memory. Their size is the parameter count × bytes per parameter. At 16-bit precision (FP16 or BF16) each parameter takes 2 bytes, so an 8-billion-parameter model needs 16 GB for its weights alone. At 8-bit it's 1 byte per parameter, and at 4-bit about half a byte.

Part 2: the KV cache

While generating, the model stores keys and values for every token in the context so it doesn't recompute them. Per token that's 2 × layers × key-value heads × head dimension × bytes. For a typical 8B model with 32 layers, 8 key-value heads and 128-dimension heads at 16-bit, that's about 131 KB per token. It grows with context length and with the number of requests served at once.

Part 3: overhead

The runtime, activations and memory fragmentation take extra. Around 10% is a reasonable planning margin, more for some serving frameworks.

An 8B model at FP16 with an 8,192-token context
  1. Weights

    8B × 2 bytesequals16 GB

  2. KV cache per token

    2 × 32 layers × 8 heads × 128 × 2 bytesequals131 KB

  3. KV cache

    131 KB × 8,192 tokens × 1equals1.07 GB

  4. Total with overhead

    (16 + 1.07) × (1 + 10%)equals18.8 GB

So an 8B model at full 16-bit precision needs about 18.8 GB and fits on a 24 GB card, not a 16 GB one.

Quantization is the big lever

Dropping precision shrinks the weights in proportion. The same 8B model needs about 10 GB at 8-bit, which fits a 12 GB card, and about 5.6 GB at 4-bit, which fits an 8 GB card.

The same model at 4-bit
  1. Weights

    8B × 0.5 bytesequals4 GB

  2. KV cache per token

    2 × 32 layers × 8 heads × 128 × 2 bytesequals131 KB

  3. KV cache

    131 KB × 8,192 tokens × 1equals1.07 GB

  4. Total with overhead

    (4 + 1.07) × (1 + 10%)equals5.6 GB

Quality usually drops a little with heavier quantization, and more for small models than large ones. Test on your own tasks before settling on a precision.

Long context and bigger models

Context is the part people forget. Raising the same 8B FP16 model's context from 8,192 to 32,768 tokens quadruples the KV cache to about 4.3 GB and takes the total to about 22.3 GB, which is tight on a 24 GB card. At the other end, a 70B model with 80 layers at 4-bit needs about 41.5 GB with an 8K context: a 48 GB card, or two 24 GB cards with the model split across them.

Tip: If a model almost fits, reduce the context length or quantize the KV cache to 8-bit before buying a bigger GPU. Both cut memory with little effort.

Questions people ask

How much VRAM does an 8B model need?
About 18.8 GB at 16-bit with an 8K context, about 10 GB at 8-bit and about 5.6 GB at 4-bit, including the KV cache and a 10% overhead margin.
What is the KV cache?
Stored keys and values for every token in the context, kept so the model doesn't recompute them. It grows with context length and with the number of simultaneous requests.
Can I run a 70B model on one GPU?
At 4-bit with a modest context, it needs roughly 41 to 42 GB, so a 48 GB card. On 24 GB cards it has to be split across two or more.