Approximate total VRAM (model weights + KV cache + overhead) to run a model at 8K context, computed with the same formula as the calculator above. Adjust the calculator for your exact model, quantization and context length.
| Model size | FP16 | Q8_0 | Q6_K | Q4_K_M ★ | Smallest GPU (Q4_K_M) |
|---|---|---|---|---|---|
| 3B | 7.8 GB | 5.0 GB | 4.3 GB | 3.6 GB | 8 GB |
| 7B | 16.1 GB | 9.5 GB | 7.9 GB | 6.4 GB | 8 GB |
| 8B | 18.2 GB | 10.7 GB | 8.7 GB | 7.1 GB | 8 GB |
| 13B | 28.7 GB | 16.2 GB | 13.1 GB | 10.4 GB | 12 GB |
| 14B | 30.9 GB | 17.3 GB | 14.0 GB | 11.0 GB | 12 GB |
| 32B | 69.2 GB | 37.7 GB | 29.6 GB | 22.6 GB | 24 GB (RTX 4090) |
| 70B | 149.8 GB | 80.7 GB | 63.1 GB | 47.6 GB | 48 GB |
| 405B | 856 GB | 456 GB | 354 GB | 265 GB | multi-GPU |
Q4_K_M (~4.85-bit) is the recommended quality/size sweet spot. Figures include a depth-calibrated KV cache and a small runtime buffer; longer context raises the KV cache. Open the calculator ↑ for exact numbers.
Exact per-model requirement pages — browse all models →
Llama 3.3 70B · Llama 3.1 8B · Qwen2.5 72B · Qwen2.5 Coder 32B · Qwen2.5 32B · DeepSeek-R1 671B · DeepSeek-R1 70B · Mixtral 8x7B · Gemma 2 27B · Phi-4 14B · Mistral 7B
About 48 GB in Q4_K_M (4-bit), 81 GB in Q8_0, or 150 GB in FP16 at 8K context. A 70B model fits on one 48 GB GPU (RTX 6000 Ada / Radeon PRO W7900) in 4-bit, or an 80 GB A100 in 8-bit.
Yes — at Q4_K_M a 32B model needs ~23 GB, which just fits a 24 GB RTX 4090 / RX 7900 XTX at 8K context. Q6_K (~30 GB) and Q8_0 (~38 GB) need a larger card.
About 6.4 GB in Q4_K_M, 9.5 GB in Q8_0, and 16 GB in FP16 — a 7B model runs comfortably on an 8 GB GPU in 4-bit, so most modern laptops and entry GPUs can handle it.
Drop the VRAM calculator into your own blog or docs — paste this snippet: