Determining the VRAM Requirements for Running LLMs Locally

When executing a local LLM, the fundamental query is whether the model can reside within your GPU's memory. The answer hinges on model size, quantization methods, and context length. This guide provides a practical baseline for selecting an appropriate VRAM capacity.

The Impact of Quantization on VRAM

Quantization lowers the precision used for storing model weights. While lower-bit quantization reduces the model footprint and VRAM consumption, it introduces a trade-off in quality.

Quant Bits per weight Typical use
Q8_0 8 Very high quality
Q6_K ~6.6 Very good quality
Q5_K_M ~5.5 Good quality and size
Q4_K_M ~4.5 Good balance of size and quality
Q3_K_M ~3.5 Lower VRAM, more quality loss

Q4_K_M is frequently chosen when VRAM is constrained. If greater VRAM capacity is available, opting for Q5 or Q6 allows the model to run with reduced quantization effects.

Estimated VRAM Requirements by Model Size

The following are approximate figures for model weights alone. Actual VRAM demand is typically higher due to overhead from the runtime, KV cache, and context storage.

Model size Q8_0 Q6_K Q4_K_M Q3_K_M
4B ~5 GB ~4 GB ~3 GB ~2.5 GB
8B ~9 GB ~7 GB ~5.5 GB ~4.5 GB
12B ~13 GB ~10 GB ~8 GB ~6.5 GB
14B ~16 GB ~12 GB ~9 GB ~7.5 GB
27B ~30 GB ~22 GB ~17 GB ~13 GB
32B ~36 GB ~27 GB ~20 GB ~16 GB
70B ~80 GB ~60 GB ~42 GB ~34 GB

These figures serve as estimates rather than strict limits. Variations in model architecture and quantization formats may alter the final size.

VRAM Capacities and Supported Model Ranges

VRAM Practical range Current examples
8 GB Small models around 4B to 9B Gemma 4 E4B, Qwen3.5 9B
12 GB Small to mid-sized models around 9B to 14B Gemma 4 12B, Qwen3.5 9B
16 GB 12B to 27B with lower quantization Gemma 4 26B-A4B, Qwen3.6 27B at Q4
24 GB 27B to 35B at Q4 to Q6 Qwen3.8 27B, Gemma 4 31B
32 GB 27B to 35B at higher quantization Qwen3.8 27B, Gemma 4 31B
48 GB Large dense models at lower quantization 70B-class models at Q3 to Q4
80 GB Large dense models at higher quantization 70B-class models at Q4 to Q6

These ranges assume the model weights can fully reside on the GPU. Large MoE models behave differently: although only a subset of parameters is active per token, the entire weight set must still be stored in memory. Consequently, a model with 100B or more total parameters cannot fit within a 100B-sized VRAM budget solely based on its active parameter count.

Mixture-of-Experts (MoE) Models

MoE architectures comprise multiple parameter groups known as experts. Since only specific experts process each token, inference can be more efficient than a dense model with the same total parameter count.

However, inactive experts remain part of the model structure. Therefore, large MoE models may demand significantly more memory than their active parameter count implies. Very large models may necessitate multiple GPUs or offloading to system RAM.

Context Length Consumes VRAM

Model weights represent only a portion of the total memory requirement. The KV cache expands as context length increases; thus, executing the same model at 64K context may require substantially more VRAM than at 4K.

  • Longer context windows demand increased VRAM.
  • KV-cache precision influences memory usage.
  • Batch size and concurrent user counts elevate memory consumption.
  • Reserve some VRAM for the runtime rather than fully saturating the GPU with model weights.

Practical Recommendations

  • Verify the actual size of the specific quantized model intended for execution.
  • Avoid using the model file size as a precise VRAM requirement; allow for KV cache and runtime overhead.
  • If the model does not fit entirely in VRAM, partial offloading to system RAM is possible, though inference speed may decrease.
  • For long-context or agentic workloads, budget more VRAM than what is required for the model weights alone.
  • Multiple GPUs can be utilized to split model execution when a single GPU lacks sufficient VRAM.

Execute on DaDesktop

There is no need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop equipped with the necessary VRAM, enabling direct model execution without hardware ownership.

Select the VRAM tier that accommodates your model, load it, and begin usage. This eliminates the need for setup, hardware procurement, or driver troubleshooting. Refer to available GPUs to view options.