DeepSeek R1 VRAM and GPU Requirements

deepseek-ai/DeepSeek-R1 · mit · Updated 2025-03-27

Parameters
684.5Bmodel record
Default context
8,192tokens
At Q4_K_M
360.1 GB8,192 tokens
At FP16
1276.5 GBfull estimate
Smallest fit
No single cardNo single GPU in our list fits this; see multi-GPU or CPU offload

No single GPU in our list fits this; see multi-GPU or CPU offload

360.1 GB

DeepSeek R1 · Q4_K_M · 8,192 tokens

Weights

358.6 GB

params × bits ÷ 8

KV cache

0.54 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Open in calculator

DeepSeek R1 needs about 360.1 GB at Q4_K_M with 8K context.

DeepSeek R1 is the reasoning mixture under an MIT license, created January. Routing keeps 256 experts, activates 8 per token, and adds 1 shared expert. The checkpoint is stored in FP8.

The cache estimate bills the compressed latent across 61 layers, not the expanded key and value tensors a runtime may build later. Q4_K_M at 8,192 tokens is 360.1 GB. No single GPU in our list fits this; see multi-GPU or CPU offload.

The measured file row matches the base V3 checkpoint's byte sum even though the parameter buckets are not identical. Configured context is 163,840 tokens. The base pretrain is a different URL.

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 1264.5 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 1268.5 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 1266.5 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 1252.5 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 1268.5 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 1264.5 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 1252.5 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 1264.5 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 1244.5 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 1252.5 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 1260.5 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 1196.5 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 1196.5 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 1240.5 GB short Won't fit
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 666.9 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 670.9 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 668.9 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 654.9 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 670.9 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 666.9 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 654.9 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 666.9 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 646.9 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 654.9 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 662.9 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 598.9 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 598.9 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 642.9 GB short Won't fit
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 512.5 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 516.5 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 514.5 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 500.5 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 516.5 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 512.5 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 500.5 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 512.5 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 492.5 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 500.5 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 508.5 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 444.5 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 444.5 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 488.5 GB short Won't fit
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 427.8 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 431.8 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 429.8 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 415.8 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 431.8 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 427.8 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 415.8 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 427.8 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 407.8 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 415.8 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 423.8 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 359.8 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 359.8 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 403.8 GB short Won't fit
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 348.1 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 352.1 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 350.1 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 336.1 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 352.1 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 348.1 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 336.1 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 348.1 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 328.1 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 336.1 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 344.1 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 280.1 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 280.1 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 324.1 GB short Won't fit
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 263.5 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 267.5 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 265.5 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 251.5 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 267.5 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 263.5 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 251.5 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 263.5 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 243.5 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 251.5 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 259.5 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 195.5 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 195.5 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 239.5 GB short Won't fit
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 198.7 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 202.7 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 200.7 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 186.7 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 202.7 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 198.7 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 186.7 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 198.7 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 178.7 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 186.7 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 194.7 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 130.7 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 130.7 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 174.7 GB short Won't fit
Fit at Published FP8 files and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 630.8 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 634.8 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 632.8 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 618.8 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 634.8 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 630.8 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 618.8 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 630.8 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 610.8 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 618.8 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 626.8 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 562.8 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 562.8 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 606.8 GB short Won't fit

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 0.54 · Overhead 1.00 16 1275.0 0.54 1.00 1276.5
Q8_0Bits 8.5 · Cache 0.54 · Overhead 1.00 8.5 677.3 0.54 1.00 678.9
Q6_KBits 6.5625 · Cache 0.54 · Overhead 1.00 6.5625 522.9 0.54 1.00 524.5
Q5_K_MBits 5.5 · Cache 0.54 · Overhead 1.00 5.5 438.3 0.54 1.00 439.8
Q4_K_MBits 4.5 · Cache 0.54 · Overhead 1.00 4.5 358.6 0.54 1.00 360.1
Q3_K_MBits 3.4375 · Cache 0.54 · Overhead 1.00 3.4375 273.9 0.54 1.00 275.5
Q2_KBits 2.625 · Cache 0.54 · Overhead 1.00 2.625 209.2 0.54 1.00 210.7
Published FP8 filesBits 8.05 · Cache 0.54 · Overhead 1.00 8.05 641.3 0.54 1.00 642.8

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 359.7 GB
  • 8,192 360.1 GB
  • 32,768 361.7 GB
  • 131,072 368.2 GB
  • 163,840 370.3 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total Published-file total
2048FP16 total 1276.1 · Published-file total 642.4 0.13 359.7 1276.1 642.4
8192FP16 total 1276.5 · Published-file total 642.8 0.54 360.1 1276.5 642.8
32768FP16 total 1278.1 · Published-file total 644.4 2.14 361.7 1278.1 644.4
131072FP16 total 1284.5 · Published-file total 650.9 8.58 368.2 1284.5 650.9
163840FP16 total 1286.7 · Published-file total 653.0 10.72 370.3 1286.7 653.0

DeepSeek R1 hardware notes

  • Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
  • Last updated 2025-03-27.
  • Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
  • Architecture numbers and the parameter count come from the public model record, checked 2026-10-10.
  • The published-file row uses a measured byte sum for the weights. Cache and overhead on that row are still estimates.

DeepSeek R1 VRAM FAQ

How much VRAM does DeepSeek R1 need at Q4_K_M?

360.1 GB at 8,192 tokens. Weights are 358.6 GB, the cache is 0.54 GB, and overhead is 1 GB.

Does a 24 GB GPU fit DeepSeek R1 at Q4_K_M?

GeForce RTX 3090 24GB is Won't fit, 336.1 GB over. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 163,840.

What context length is configured for DeepSeek R1?

The configured context length is 163,840 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 163,840.

Which source supplies the parameter count for DeepSeek R1?

684.5B parameters, from the Hugging Face model record for deepseek-ai/DeepSeek-R1.

Models in the same family, then the closest Q4_K_M totals.
Model Q4_K_M GB
DeepSeek V3 360.1
Kimi K2 539.2
gpt-oss-120b 62.49
Llama 3.1 70B 40.46
Llama 3.3 70B 40.46