Llama 3.1 405B VRAM and GPU Requirements

meta-llama/Llama-3.1-405B-Instruct · llama3.1 · Updated 2024-09-25

Parameters
405.9Bmodel record
Default context
8,192tokens
At Q4_K_M
217.6 GB8,192 tokens
At FP16
760.9 GBfull estimate
Smallest fit
No single cardNo single GPU in our list fits this; see multi-GPU or CPU offload

No single GPU in our list fits this; see multi-GPU or CPU offload

217.6 GB

Llama 3.1 405B · Q4_K_M · 8,192 tokens

Weights

212.6 GB

params × bits ÷ 8

KV cache

3.94 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Open in calculator

Llama 3.1 405B needs about 217.6 GB at Q4_K_M with 8K context.

Llama 3.1 405B is the dense long-context giant, and it does not fit any single GPU in this list at Q4_K_M. The estimate is 217.6 GB at 8,192 tokens. No single GPU in our list fits this; see multi-GPU or CPU offload.

FP16 weights alone are 756.0 GB, the parameter count times 2 bytes, before cache and overhead. The official file is gated. The only public architecture copy used here writes 16 key heads and a 4-bit packing block.

Both are set aside: the official parameter total matches a dense reconstruction only at 8 key heads, and the weights are not treated as pre-quantized. Depth is 126 layers. Context scales to 131,072 tokens.

The 8B and 70B Llama 3.1 pages are the same generation with far fewer layers.

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 748.9 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 752.9 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 750.9 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 736.9 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 752.9 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 748.9 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 736.9 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 748.9 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 728.9 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 736.9 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 744.9 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 680.9 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 680.9 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 724.9 GB short Won't fit
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 394.5 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 398.5 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 396.5 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 382.5 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 398.5 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 394.5 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 382.5 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 394.5 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 374.5 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 382.5 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 390.5 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 326.5 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 326.5 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 370.5 GB short Won't fit
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 303.0 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 307.0 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 305.0 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 291.0 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 307.0 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 303.0 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 291.0 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 303.0 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 283.0 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 291.0 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 299.0 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 235.0 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 235.0 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 279.0 GB short Won't fit
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 252.8 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 256.8 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 254.8 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 240.8 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 256.8 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 252.8 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 240.8 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 252.8 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 232.8 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 240.8 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 248.8 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 184.8 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 184.8 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 228.8 GB short Won't fit
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 205.6 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 209.6 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 207.6 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 193.6 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 209.6 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 205.6 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 193.6 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 205.6 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 185.6 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 193.6 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 201.6 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 137.6 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 137.6 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 181.6 GB short Won't fit
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 155.4 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 159.4 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 157.4 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 143.4 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 159.4 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 155.4 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 143.4 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 155.4 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 135.4 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 143.4 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 151.4 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 87.35 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 87.35 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 131.4 GB short Won't fit
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 117.0 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 121.0 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 119.0 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 105.0 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 121.0 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 117.0 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 105.0 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 117.0 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 96.96 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 105.0 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 113.0 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 48.96 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 48.96 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 92.96 GB short CPU offload

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 3.94 · Overhead 1.00 16 756.0 3.94 1.00 760.9
Q8_0Bits 8.5 · Cache 3.94 · Overhead 1.00 8.5 401.6 3.94 1.00 406.5
Q6_KBits 6.5625 · Cache 3.94 · Overhead 1.00 6.5625 310.1 3.94 1.00 315.0
Q5_K_MBits 5.5 · Cache 3.94 · Overhead 1.00 5.5 259.9 3.94 1.00 264.8
Q4_K_MBits 4.5 · Cache 3.94 · Overhead 1.00 4.5 212.6 3.94 1.00 217.6
Q3_K_MBits 3.4375 · Cache 3.94 · Overhead 1.00 3.4375 162.4 3.94 1.00 167.4
Q2_KBits 2.625 · Cache 3.94 · Overhead 1.00 2.625 124.0 3.94 1.00 129.0

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 214.6 GB
  • 8,192 217.6 GB
  • 32,768 229.4 GB
  • 131,072 276.6 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total
2048FP16 total 758.0 0.98 214.6 758.0
8192FP16 total 760.9 3.94 217.6 760.9
32768FP16 total 772.7 15.75 229.4 772.7
131072FP16 total 820.0 63.00 276.6 820.0

Llama 3.1 405B hardware notes

  • Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
  • Last updated 2024-09-25.
  • Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
  • Architecture numbers come from hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4 because the official download is gated. The parameter count is the official Hugging Face record.

Llama 3.1 405B VRAM FAQ

How much VRAM does Llama 3.1 405B need at Q4_K_M?

217.6 GB at 8,192 tokens. Weights are 212.6 GB, the cache is 3.94 GB, and overhead is 1 GB.

Does a 24 GB GPU fit Llama 3.1 405B at Q4_K_M?

GeForce RTX 3090 24GB is Won't fit, 193.6 GB over. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 131,072.

What context length is configured for Llama 3.1 405B?

The configured context length is 131,072 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 131,072.

Which source supplies the parameter count for Llama 3.1 405B?

405.9B parameters, from the Hugging Face model record for meta-llama/Llama-3.1-405B-Instruct. Architecture numbers come from the public copy at hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4, because the official download is gated. The parameter count is the official record.

Models in the same family, then the closest Q4_K_M totals.
Model Q4_K_M GB
Llama 3.1 8B 6.21
Llama 3.1 70B 40.46
DeepSeek R1 360.1
DeepSeek V3 360.1
gpt-oss-120b 62.49
Llama 3.3 70B 40.46