VRAM Calculator for Local LLMs & Gaming

Calculate how much VRAM you need for local AI models or gaming before downloading a model or upgrading your GPU. Estimate model weights, KV cache, runtime overhead, and gaming memory requirements in seconds.

LLM Inference

Model Configuration

Enter your model size, quantization, context length, and batch size.

B

Enter the model size in billions of parameters.

Lower-bit quantization reduces model weight memory but can slightly affect output quality.

Longer context windows increase KV cache memory usage.

Larger batches increase cache and runtime memory requirements.

Estimated VRAM Required

6.5GB

Includes model weights, estimated KV cache, and runtime overhead.

Minimum GPU tier

8 GB

Comfortable target

8 GB

Memory Breakdown

Where your VRAM goes

Model Weights

4.00 GB

Memory used to store the model parameters.

KV Cache

1.50 GB

Estimated memory used by the context window and batch.

Runtime Overhead

0.99 GB

18% allowance for runtime buffers and framework memory.

GPU Compatibility

Which VRAM tiers fit?

Requirement: 6.5 GB

6 GB
Not Enough

RTX 3060 6GB

8 GB
Recommended

RTX 4060

1.5 GB headroom

12 GB
Recommended

RTX 4070 • RTX 5070

5.5 GB headroom

16 GB
Recommended

RTX 4080 • RTX 5080 • RTX 5070 Ti

9.5 GB headroom

24 GB
Recommended

RTX 4090

17.5 GB headroom

32 GB
Recommended

RTX 5090

25.5 GB headroom

48 GB+
Recommended

Dual 24GB GPUs • Multi-GPU setup

41.5 GB headroom

96 GB+
Recommended

High-memory multi-GPU setup

89.5 GB headroom

“Fits” means the estimated workload is within the card's capacity. “Recommended” includes additional headroom for a more comfortable setup.

VRAM results are estimates. Actual memory usage can vary by model architecture, inference engine, game, driver, and workload.

Quick Answer

At Q4_K_M, 4K context and batch size 1, expect about 6.3 GB for an 8B model, 10.8 GB for 14B, 22.7 GB for 32B and 47 GB for 70B. Longer context, bigger batches and higher-precision quantization all add to that. For a 7B to 8B model, 8 GB is the practical minimum.

How to Use the Calculator

  1. Pick a model. Choose a preset (it loads the real layer and KV-head counts) or type a custom size in billions of parameters.
  2. Choose quantization. Q4_K_M is the usual balance of size and quality. Use Q5 or Q6 if you have spare memory.
  3. Set context length and batch size. Context is the biggest hidden cost, see below.
  4. Read the breakdown and GPU fit. The tool shows the smallest card that fits and the first card with comfortable headroom.

How This Calculator Works

total VRAM = weights + KV cache + runtime overhead

weights = parameters (billions) × bytes per parameter

KV cache = 2 × layers × KV heads × head dim × bytes × context × batch

overhead = 0.6 GB + 5% of (weights + KV cache)

Worked example: Llama 3.1 8B, Q4_K_M, 4,096 context, batch 1

PartCalculationGB
Weights8.03B × 0.614.90
KV cache2 × 32 layers × 8 KV heads × 128 × 2 bytes × 4,0960.54
Overhead0.6 + 5% of 5.440.87
Total6.31

Bytes per Parameter by Quantization

Bytes per parameter used (approximate, based on llama.cpp GGUF sizes; real files include some higher-precision layers, so they are not simply bits divided by eight):

FormatBytes per parameter
FP324.00
FP16 / BF162.00
Q8_01.06
Q6_K0.82
Q5_K_M0.71
Q4_K_M0.61
Q3_K_M0.49
Q2_K0.42

Why the KV cache depends on the model. Modern models use grouped-query attention (GQA), which shares key/value heads and keeps the cache small. An older model without GQA, such as Llama 2 13B (40 layers, 40 KV heads), needs about 3.4 GB of KV cache at 4K context, roughly six times more than Llama 3.1 8B. That is why the calculator asks for the model instead of using one flat number.

Context Length: The Cost People Miss

The KV cache grows linearly with context and batch size. For Llama 3.1 8B at FP16 KV precision:

ContextKV cache
4,0960.54 GB
32,7684.3 GB
131,07217.2 GB

A model that loads fine at 4K can run out of memory at 32K. Quantizing the KV cache to 8-bit halves these numbers.

What Size Model Fits on Your GPU?

Usable memory is about 90% of the card's VRAM, because the driver and display take a share. Figures assume 4K context and batch size 1.

VRAMExample cardsRealistic fit
6 GBRTX 2060, laptop RTX 30603B to 4B at Q4; 7B only at Q3 or short context
8 GBRTX 4060, RTX 30707B to 8B at Q4
12 GBRTX 3060 12GB, RTX 4070, RTX 50708B up to Q8; 14B at Q4 (tight)
16 GBRTX 4060 Ti 16GB, RTX 4080, RTX 5070 Ti, RTX 508014B up to Q6
24 GBRTX 3090, RTX 4090, RX 7900 XTX14B at Q8; 32B at Q4 with short context
32 GBRTX 509032B up to Q6
48 GBRTX A6000, L40S, 2x 24 GB32B at Q8; 70B at Q4 is borderline (47 GB)
80 GBA100 80GB, H100 80GB70B at Q4 with room for context
96 GBRTX PRO 600070B at Q4 with long context

If the Model Does Not Fit

  1. Use a lower quantization (Q4_K_M instead of Q6).
  2. Shorten the context window.
  3. Quantize the KV cache to 8-bit.
  4. Lower the batch size.
  5. Offload some layers to system RAM. It works, but token speed drops sharply.

Mixture-of-Experts (MoE) Models

For a normal GPU setup, an MoE model needs enough VRAM for all its weights, not just the active parameters. Active parameters make generation faster, but they do not shrink the memory footprint unless you offload experts to system RAM. Enter the total parameter count.

VRAM vs System RAM

VRAM sits on the graphics card and feeds the GPU at very high bandwidth. System RAM can hold model data for CPU inference or offloading, but moving data between the two is slower. VRAM size decides whether a model fits; memory bandwidth largely decides how fast it generates.

What This Calculator Does Not Cover

  • Training and fine-tuning. They need extra memory for gradients, optimizer state and activations, so this tool is for inference only.
  • Engine differences. llama.cpp, Ollama and vLLM allocate memory differently. vLLM, for example, reserves a fixed share of GPU memory up front, so your monitoring tool may show more than the estimate.
  • Multi-GPU and tensor-parallel overhead.

Leave about 10% headroom on top of the estimate.

Frequently Asked Questions

Is 6 GB of VRAM enough for a local LLM in 2026?

For 3B to 4B models at Q4, yes; a 3B model at 4K context needs about 3.2 GB. A 7B to 8B model at Q4_K_M needs about 5.8 to 6.3 GB, which is more than the roughly 5.4 GB usable on a 6 GB card. Q3 or a shorter context can squeeze it in, but 8 GB is the practical minimum for regular 7B use.

How much VRAM do 7B, 14B and 70B models need?

At Q4_K_M and 4K context: about 5.8 GB for 7B, 10.8 GB for 14B and 47 GB for 70B. Longer context adds more.

Does context length change VRAM use?

Yes. The KV cache grows with every token kept in context. For an 8B model it goes from 0.5 GB at 4K to about 4.3 GB at 32K.

Do MoE models need VRAM for all parameters?

Yes, for standard GPU inference. Active parameters affect speed, not the memory needed to hold the model.

Can I run an LLM on system RAM only?

Yes, with CPU inference or offloading. It is usually much slower than running from VRAM.

Which quantization should I choose?

Q4_K_M is the usual default. Move to Q5 or Q6 when you have spare VRAM, and Q8 when quality matters most. Below Q4, quality loss becomes more noticeable.

How accurate is this calculator?

It is an estimate built from published model architectures and llama.cpp file sizes. Real usage varies with the inference engine, so leave about 10% headroom.

Does VRAM bandwidth matter?

Yes. Two cards with the same VRAM can generate at very different speeds because token generation depends heavily on memory bandwidth.