Quick Answer
At Q4_K_M, 4K context and batch size 1, expect about 6.3 GB for an 8B model, 10.8 GB for 14B, 22.7 GB for 32B and 47 GB for 70B. Longer context, bigger batches and higher-precision quantization all add to that. For a 7B to 8B model, 8 GB is the practical minimum.
How to Use the Calculator
- Pick a model. Choose a preset (it loads the real layer and KV-head counts) or type a custom size in billions of parameters.
- Choose quantization. Q4_K_M is the usual balance of size and quality. Use Q5 or Q6 if you have spare memory.
- Set context length and batch size. Context is the biggest hidden cost, see below.
- Read the breakdown and GPU fit. The tool shows the smallest card that fits and the first card with comfortable headroom.
How This Calculator Works
total VRAM = weights + KV cache + runtime overhead
weights = parameters (billions) × bytes per parameter
KV cache = 2 × layers × KV heads × head dim × bytes × context × batch
overhead = 0.6 GB + 5% of (weights + KV cache)
Worked example: Llama 3.1 8B, Q4_K_M, 4,096 context, batch 1
| Part | Calculation | GB |
|---|---|---|
| Weights | 8.03B × 0.61 | 4.90 |
| KV cache | 2 × 32 layers × 8 KV heads × 128 × 2 bytes × 4,096 | 0.54 |
| Overhead | 0.6 + 5% of 5.44 | 0.87 |
| Total | 6.31 |
Bytes per Parameter by Quantization
Bytes per parameter used (approximate, based on llama.cpp GGUF sizes; real files include some higher-precision layers, so they are not simply bits divided by eight):
| Format | Bytes per parameter |
|---|---|
| FP32 | 4.00 |
| FP16 / BF16 | 2.00 |
| Q8_0 | 1.06 |
| Q6_K | 0.82 |
| Q5_K_M | 0.71 |
| Q4_K_M | 0.61 |
| Q3_K_M | 0.49 |
| Q2_K | 0.42 |
Why the KV cache depends on the model. Modern models use grouped-query attention (GQA), which shares key/value heads and keeps the cache small. An older model without GQA, such as Llama 2 13B (40 layers, 40 KV heads), needs about 3.4 GB of KV cache at 4K context, roughly six times more than Llama 3.1 8B. That is why the calculator asks for the model instead of using one flat number.
Context Length: The Cost People Miss
The KV cache grows linearly with context and batch size. For Llama 3.1 8B at FP16 KV precision:
| Context | KV cache |
|---|---|
| 4,096 | 0.54 GB |
| 32,768 | 4.3 GB |
| 131,072 | 17.2 GB |
A model that loads fine at 4K can run out of memory at 32K. Quantizing the KV cache to 8-bit halves these numbers.
VRAM Needed for Popular Models
Estimates at 4,096 context, batch size 1, FP16 KV cache, using the formula above.
| Model | FP16 | Q8_0 | Q4_K_M |
|---|---|---|---|
| Mistral 7B v0.3 | 16.4 GB | 9.2 GB | 5.8 GB |
| Llama 3.1 8B | 18.0 GB | 10.1 GB | 6.3 GB |
| Qwen3 14B | 32.4 GB | 17.8 GB | 10.8 GB |
| Qwen3 32B | 70.6 GB | 38.2 GB | 22.7 GB |
| Llama 3.1 70B | 150.3 GB | 80.6 GB | 47.2 GB |
| Llama 3.1 405B | 853 GB | 454 GB | 262 GB |
These are estimates, not guaranteed minimums. Check the download size of the exact file you plan to use.
What Size Model Fits on Your GPU?
Usable memory is about 90% of the card's VRAM, because the driver and display take a share. Figures assume 4K context and batch size 1.
| VRAM | Example cards | Realistic fit |
|---|---|---|
| 6 GB | RTX 2060, laptop RTX 3060 | 3B to 4B at Q4; 7B only at Q3 or short context |
| 8 GB | RTX 4060, RTX 3070 | 7B to 8B at Q4 |
| 12 GB | RTX 3060 12GB, RTX 4070, RTX 5070 | 8B up to Q8; 14B at Q4 (tight) |
| 16 GB | RTX 4060 Ti 16GB, RTX 4080, RTX 5070 Ti, RTX 5080 | 14B up to Q6 |
| 24 GB | RTX 3090, RTX 4090, RX 7900 XTX | 14B at Q8; 32B at Q4 with short context |
| 32 GB | RTX 5090 | 32B up to Q6 |
| 48 GB | RTX A6000, L40S, 2x 24 GB | 32B at Q8; 70B at Q4 is borderline (47 GB) |
| 80 GB | A100 80GB, H100 80GB | 70B at Q4 with room for context |
| 96 GB | RTX PRO 6000 | 70B at Q4 with long context |
If the Model Does Not Fit
- Use a lower quantization (Q4_K_M instead of Q6).
- Shorten the context window.
- Quantize the KV cache to 8-bit.
- Lower the batch size.
- Offload some layers to system RAM. It works, but token speed drops sharply.
Mixture-of-Experts (MoE) Models
For a normal GPU setup, an MoE model needs enough VRAM for all its weights, not just the active parameters. Active parameters make generation faster, but they do not shrink the memory footprint unless you offload experts to system RAM. Enter the total parameter count.
VRAM vs System RAM
VRAM sits on the graphics card and feeds the GPU at very high bandwidth. System RAM can hold model data for CPU inference or offloading, but moving data between the two is slower. VRAM size decides whether a model fits; memory bandwidth largely decides how fast it generates.
What This Calculator Does Not Cover
- Training and fine-tuning. They need extra memory for gradients, optimizer state and activations, so this tool is for inference only.
- Engine differences. llama.cpp, Ollama and vLLM allocate memory differently. vLLM, for example, reserves a fixed share of GPU memory up front, so your monitoring tool may show more than the estimate.
- Multi-GPU and tensor-parallel overhead.
Leave about 10% headroom on top of the estimate.
Frequently Asked Questions
Is 6 GB of VRAM enough for a local LLM in 2026?
For 3B to 4B models at Q4, yes; a 3B model at 4K context needs about 3.2 GB. A 7B to 8B model at Q4_K_M needs about 5.8 to 6.3 GB, which is more than the roughly 5.4 GB usable on a 6 GB card. Q3 or a shorter context can squeeze it in, but 8 GB is the practical minimum for regular 7B use.
How much VRAM do 7B, 14B and 70B models need?
At Q4_K_M and 4K context: about 5.8 GB for 7B, 10.8 GB for 14B and 47 GB for 70B. Longer context adds more.
Does context length change VRAM use?
Yes. The KV cache grows with every token kept in context. For an 8B model it goes from 0.5 GB at 4K to about 4.3 GB at 32K.
Do MoE models need VRAM for all parameters?
Yes, for standard GPU inference. Active parameters affect speed, not the memory needed to hold the model.
Can I run an LLM on system RAM only?
Yes, with CPU inference or offloading. It is usually much slower than running from VRAM.
Which quantization should I choose?
Q4_K_M is the usual default. Move to Q5 or Q6 when you have spare VRAM, and Q8 when quality matters most. Below Q4, quality loss becomes more noticeable.
How accurate is this calculator?
It is an estimate built from published model architectures and llama.cpp file sizes. Real usage varies with the inference engine, so leave about 10% headroom.
Does VRAM bandwidth matter?
Yes. Two cards with the same VRAM can generate at very different speeds because token generation depends heavily on memory bandwidth.