Llama 4 Maverick needs about 212.9 GB at Q4_K_M with 8K context.
Llama 4 Maverick is the wide mixture: 128 local experts and 1 active per token, across 48 layers. The parameter total includes the vision tower. Unlike Scout, this file does not scale context.
The maximum is 1,048,576 tokens, and the table itself stops at 131,072. Chunked layers still use an 8,192-token chunk. Q4_K_M at 8,192 tokens is 212.9 GB.
No single GPU in our list fits this; see multi-GPU or CPU offload. The download is gated and the public architecture copy is Unsloth. There is no packed 4-bit expert format and no compressed latent cache.
Query and key head counts match Scout. The expert fan-out does not.
Which GPUs fit
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 738.5 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 742.5 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 740.5 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 726.5 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 742.5 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 738.5 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 726.5 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 738.5 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · Won't fit | 32 GB | 718.5 GB short | Won't fit |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 726.5 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 734.5 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · Won't fit | 80 GB | 670.5 GB short | Won't fit |
| NVIDIA H100 SXM 80GB80 GB · Won't fit | 80 GB | 670.5 GB short | Won't fit |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit | 48 GB | 714.5 GB short | Won't fit |
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 387.9 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 391.9 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 389.9 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 375.9 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 391.9 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 387.9 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 375.9 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 387.9 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · Won't fit | 32 GB | 367.9 GB short | Won't fit |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 375.9 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 383.9 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · Won't fit | 80 GB | 319.9 GB short | Won't fit |
| NVIDIA H100 SXM 80GB80 GB · Won't fit | 80 GB | 319.9 GB short | Won't fit |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit | 48 GB | 363.9 GB short | Won't fit |
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 297.3 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 301.3 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 299.3 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 285.3 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 301.3 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 297.3 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 285.3 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 297.3 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · Won't fit | 32 GB | 277.3 GB short | Won't fit |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 285.3 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 293.3 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · CPU offload | 80 GB | 229.3 GB short | CPU offload |
| NVIDIA H100 SXM 80GB80 GB · CPU offload | 80 GB | 229.3 GB short | CPU offload |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit | 48 GB | 273.3 GB short | Won't fit |
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 247.6 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 251.6 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 249.6 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 235.6 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 251.6 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 247.6 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 235.6 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 247.6 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · Won't fit | 32 GB | 227.6 GB short | Won't fit |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 235.6 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 243.6 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · CPU offload | 80 GB | 179.6 GB short | CPU offload |
| NVIDIA H100 SXM 80GB80 GB · CPU offload | 80 GB | 179.6 GB short | CPU offload |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit | 48 GB | 223.6 GB short | Won't fit |
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 200.9 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 204.9 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 202.9 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 188.9 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 204.9 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 200.9 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 188.9 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 200.9 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · Won't fit | 32 GB | 180.9 GB short | Won't fit |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 188.9 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 196.9 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · CPU offload | 80 GB | 132.9 GB short | CPU offload |
| NVIDIA H100 SXM 80GB80 GB · CPU offload | 80 GB | 132.9 GB short | CPU offload |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit | 48 GB | 176.9 GB short | Won't fit |
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 151.2 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 155.2 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 153.2 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 139.2 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 155.2 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 151.2 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 139.2 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 151.2 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · Won't fit | 32 GB | 131.2 GB short | Won't fit |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 139.2 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 147.2 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · CPU offload | 80 GB | 83.20 GB short | CPU offload |
| NVIDIA H100 SXM 80GB80 GB · CPU offload | 80 GB | 83.20 GB short | CPU offload |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit | 48 GB | 127.2 GB short | Won't fit |
| GPU | Memory | Status | Headroom |
|---|---|---|---|
| GeForce RTX 3060 12GB12 GB · Won't fit | 12 GB | 113.2 GB short | Won't fit |
| GeForce RTX 3070 8GB8 GB · Won't fit | 8 GB | 117.2 GB short | Won't fit |
| GeForce RTX 3080 10GB10 GB · Won't fit | 10 GB | 115.2 GB short | Won't fit |
| GeForce RTX 3090 24GB24 GB · Won't fit | 24 GB | 101.2 GB short | Won't fit |
| GeForce RTX 4060 8GB8 GB · Won't fit | 8 GB | 117.2 GB short | Won't fit |
| GeForce RTX 4060 Ti 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| GeForce RTX 4070 12GB12 GB · Won't fit | 12 GB | 113.2 GB short | Won't fit |
| GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| GeForce RTX 4080 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| GeForce RTX 4090 24GB24 GB · Won't fit | 24 GB | 101.2 GB short | Won't fit |
| GeForce RTX 5060 Ti 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| GeForce RTX 5070 12GB12 GB · Won't fit | 12 GB | 113.2 GB short | Won't fit |
| GeForce RTX 5070 Ti 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| GeForce RTX 5080 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| GeForce RTX 5090 32GB32 GB · CPU offload | 32 GB | 93.22 GB short | CPU offload |
| Radeon RX 7900 XTX 24GB24 GB · Won't fit | 24 GB | 101.2 GB short | Won't fit |
| Radeon RX 9070 XT 16GB16 GB · Won't fit | 16 GB | 109.2 GB short | Won't fit |
| NVIDIA A100 80GB80 GB · CPU offload | 80 GB | 45.22 GB short | CPU offload |
| NVIDIA H100 SXM 80GB80 GB · CPU offload | 80 GB | 45.22 GB short | CPU offload |
| Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload | 48 GB | 89.22 GB short | CPU offload |
VRAM by quantization
Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.
| Quantization | Bits | Weights | Cache | Overhead | Total |
|---|---|---|---|---|---|
| FP16/BF16Bits 16 · Cache 1.50 · Overhead 1.00 | 16 | 748.0 | 1.50 | 1.00 | 750.5 |
| Q8_0Bits 8.5 · Cache 1.50 · Overhead 1.00 | 8.5 | 397.4 | 1.50 | 1.00 | 399.9 |
| Q6_KBits 6.5625 · Cache 1.50 · Overhead 1.00 | 6.5625 | 306.8 | 1.50 | 1.00 | 309.3 |
| Q5_K_MBits 5.5 · Cache 1.50 · Overhead 1.00 | 5.5 | 257.1 | 1.50 | 1.00 | 259.6 |
| Q4_K_MBits 4.5 · Cache 1.50 · Overhead 1.00 | 4.5 | 210.4 | 1.50 | 1.00 | 212.9 |
| Q3_K_MBits 3.4375 · Cache 1.50 · Overhead 1.00 | 3.4375 | 160.7 | 1.50 | 1.00 | 163.2 |
| Q2_KBits 2.625 · Cache 1.50 · Overhead 1.00 | 2.625 | 122.7 | 1.50 | 1.00 | 125.2 |
Context length
Q4_K_M total as the context changes. Cache is the part that grows.
| Context tokens | Cache GB | Q4_K_M total | FP16 total |
|---|---|---|---|
| 2048FP16 total 749.4 | 0.38 | 211.8 | 749.4 |
| 8192FP16 total 750.5 | 1.50 | 212.9 | 750.5 |
| 32768FP16 total 751.6 | 2.63 | 214.0 | 751.6 |
| 131072FP16 total 756.1 | 7.13 | 218.5 | 756.1 |
The configured context length is 1,048,576 tokens. The context table stops at 131,072 tokens.
Llama 4 Maverick hardware notes
- Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
- Last updated 2025-05-22.
- Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
- Architecture numbers come from unsloth/Llama-4-Maverick-17B-128E-Instruct because the official download is gated. The parameter count is the official Hugging Face record.
Llama 4 Maverick VRAM FAQ
How much VRAM does Llama 4 Maverick need at Q4_K_M?
212.9 GB at 8,192 tokens. Weights are 210.4 GB, the cache is 1.50 GB, and overhead is 1 GB.
Does a 24 GB GPU fit Llama 4 Maverick at Q4_K_M?
GeForce RTX 3090 24GB is Won't fit, 188.9 GB over. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 1,048,576.
What context length is configured for Llama 4 Maverick?
The configured context length is 1,048,576 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 1,048,576.
Which source supplies the parameter count for Llama 4 Maverick?
401.6B parameters, from the Hugging Face model record for meta-llama/Llama-4-Maverick-17B-128E-Instruct. Architecture numbers come from the public copy at unsloth/Llama-4-Maverick-17B-128E-Instruct, because the official download is gated. The parameter count is the official record.
Related models
| Model | Q4_K_M GB |
|---|---|
| DeepSeek R1 | 360.1 |
| DeepSeek V3 | 360.1 |
| gpt-oss-120b | 62.49 |
| Llama 3.1 70B | 40.46 |