LLM VRAM Requirements by Model

Estimates for all 30 models. Each total is weights, cache, and 1 GB of overhead at 8,192 tokens, or at the model's maximum when that maximum is lower.

Models that fit 8 GB

These models need at most 8 GB at Q4_K_M with 8K context.

VRAM in GB at 8,192 tokens, or the model's maximum when that is lower. The FP16 column is the full estimate. Weights alone are the parameter count times 2 bytes.
Model Parameters Q4_K_M Q8_0 FP16
Llama 3.1 8BParameters 8.03B · Q8_0 9.95 8.03B 6.21 9.95 16.96
Llama 3 8BParameters 8.03B · Q8_0 9.95 8.03B 6.21 9.95 16.96
Mistral 7BParameters 7.25B · Q8_0 9.17 7.25B 5.80 9.17 15.50
Phi-3 miniParameters 3.82B · Q8_0 5.53 3.82B 3.75 5.53 8.87
Qwen3 8BParameters 8.19B · Q8_0 10.23 8.19B 6.42 10.23 17.38

Models that fit 16 GB

These models need more than 8 GB and at most 16 GB at Q4_K_M with 8K context.

VRAM in GB at 8,192 tokens, or the model's maximum when that is lower. The FP16 column is the full estimate. Weights alone are the parameter count times 2 bytes.
Model Parameters Q4_K_M Q8_0 FP16
gpt-oss-20bParameters 20.91B · Q8_0 21.89 20.91B 12.15 21.89 40.15
Mistral Small 3.2Parameters 24.01B · Q8_0 26.01 24.01B 14.83 26.01 46.97
Gemma 3 12BParameters 12.19B · Q8_0 13.87 12.19B 8.20 13.87 24.51
Phi-4Parameters 14.66B · Q8_0 17.07 14.66B 10.24 17.07 29.87
Qwen3 14BParameters 14.77B · Q8_0 16.86 14.77B 9.99 16.86 29.76

Models that fit 24 GB

These models need more than 16 GB and at most 24 GB at Q4_K_M with 8K context.

VRAM in GB at 8,192 tokens, or the model's maximum when that is lower. The FP16 column is the full estimate. Weights alone are the parameter count times 2 bytes.
Model Parameters Q4_K_M Q8_0 FP16
Gemma 4Parameters 31.27B · Q8_0 33.04 31.27B 18.48 33.04 60.34
Gemma 3 27BParameters 27.43B · Q8_0 29.18 27.43B 16.40 29.18 53.13
DeepSeek R1 Distill Qwen 32BParameters 32.76B · Q8_0 35.42 32.76B 20.16 35.42 64.03
QwQ-32BParameters 32.76B · Q8_0 35.42 32.76B 20.16 35.42 64.03
Qwen3 CoderParameters 30.53B · Q8_0 31.96 30.53B 17.74 31.96 58.62
Qwen2.5 Coder 32BParameters 32.76B · Q8_0 35.42 32.76B 20.16 35.42 64.03
Qwen3 32BParameters 32.76B · Q8_0 35.42 32.76B 20.16 35.42 64.02

Models that fit 48 GB+

These models need more than 24 GB at Q4_K_M with 8K context.

VRAM in GB at 8,192 tokens, or the model's maximum when that is lower. The FP16 column is the full estimate. Weights alone are the parameter count times 2 bytes.
Model Parameters Q4_K_M Q8_0 FP16
gpt-oss-120bParameters 116.8B · Q8_0 116.9 116.8B 62.49 116.9 218.9
Kimi K2Parameters 1026.4B · Q8_0 1017.2 1026.4B 539.2 1017.2 1913.4
DeepSeek R1Parameters 684.5B · Q8_0 678.9 684.5B 360.1 678.9 1276.5
Llama 3.3 70BParameters 70.55B · Q8_0 73.32 70.55B 40.46 73.32 134.9
Llama 3.1 70BParameters 70.55B · Q8_0 73.32 70.55B 40.46 73.32 134.9
DeepSeek V3Parameters 684.5B · Q8_0 678.9 684.5B 360.1 678.9 1276.6
Llama 3 70BParameters 70.55B · Q8_0 73.32 70.55B 40.46 73.32 134.9
Llama 3.1 405BParameters 405.9B · Q8_0 406.5 405.9B 217.6 406.5 760.9
Llama 4 ScoutParameters 108.6B · Q8_0 110.0 108.6B 59.41 110.0 204.9
Llama 4 MaverickParameters 401.6B · Q8_0 399.9 401.6B 212.9 399.9 750.5
Mixtral 8x7BParameters 46.70B · Q8_0 48.21 46.70B 26.47 48.21 88.99
DeepSeek R1 Distill Llama 70BParameters 70.55B · Q8_0 73.32 70.55B 40.46 73.32 134.9
Qwen3 235B-A22BParameters 235.1B · Q8_0 235.1 235.1B 125.6 235.1 440.4