Kimi K2 VRAM and GPU Requirements

moonshotai/Kimi-K2-Instruct · other · Updated 2026-04-23

Parameters
1026.4Bmodel record
Default context
8,192tokens
At Q4_K_M
539.2 GB8,192 tokens
At FP16
1913.4 GBfull estimate
Smallest fit
No single cardNo single GPU in our list fits this; see multi-GPU or CPU offload

No single GPU in our list fits this; see multi-GPU or CPU offload

539.2 GB

Kimi K2 · Q4_K_M · 8,192 tokens

Weights

537.7 GB

params × bits ÷ 8

KV cache

0.54 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Open in calculator

Kimi K2 needs about 539.2 GB at Q4_K_M with 8K context.

These tables use Kimi K2 Instruct, the verified checkpoint at 1026.4B parameters. FP16 weights are 1911.8 GB, the parameter count times 2 bytes. The full FP16 estimate is 1913.4 GB because it also adds the cache and 1 GB of overhead.

It is a mixture of 384 routed experts plus 1 shared expert, and the cache stores a compressed latent rather than a full key and value pair, which is why the cache stays small. Q4_K_M at 8,192 tokens is 539.2 GB. No single GPU in our list fits this; see multi-GPU or CPU offload.

The license on the Hub record is a modified MIT. Later Kimi checkpoints, including K2.6, stay in the family table and are not given a second VRAM table. Configured context is 131,072 tokens.

Other Kimi K2 repos

The quant tables use the lead repo. The rows below are parameter totals, license tags, creation dates, and context from the model record. They are not VRAM estimates.

Family repos from the model record. No VRAM total is computed for these rows.
Repo Parameters License tag Created Context tokens
moonshotai/Kimi-K2-InstructLicense tag other · Created 2025-07-11 1026.4B other 2025-07-11 131072
moonshotai/Kimi-K2-Instruct-0905License tag other · Created 2025-09-03 1026.5B other 2025-09-03 262144
moonshotai/Kimi-K2-ThinkingLicense tag other · Created 2025-11-04 1026.4B other 2025-11-04 262144
moonshotai/Kimi-K2.5License tag other · Created 2026-01-01 1026.9B other 2026-01-01 262144
moonshotai/Kimi-K2.6License tag other · Created 2026-04-14 1026.9B other 2026-04-14 262144
moonshotai/Kimi-K2.7-CodeLicense tag other · Created 2026-06-11 1026.9B other 2026-06-11 262144

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 1901.4 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 1905.4 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 1903.4 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 1889.4 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 1905.4 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 1901.4 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 1889.4 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 1901.4 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 1881.4 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 1889.4 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 1897.4 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 1833.4 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 1833.4 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 1877.4 GB short Won't fit
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 1005.2 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 1009.2 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 1007.2 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 993.2 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 1009.2 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 1005.2 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 993.2 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 1005.2 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 985.2 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 993.2 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 1001.2 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 937.2 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 937.2 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 981.2 GB short Won't fit
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 773.7 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 777.7 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 775.7 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 761.7 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 777.7 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 773.7 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 761.7 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 773.7 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 753.7 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 761.7 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 769.7 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 705.7 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 705.7 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 749.7 GB short Won't fit
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 646.7 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 650.7 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 648.7 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 634.7 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 650.7 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 646.7 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 634.7 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 646.7 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 626.7 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 634.7 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 642.7 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 578.7 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 578.7 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 622.7 GB short Won't fit
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 527.2 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 531.2 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 529.2 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 515.2 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 531.2 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 527.2 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 515.2 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 527.2 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 507.2 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 515.2 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 523.2 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 459.2 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 459.2 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 503.2 GB short Won't fit
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 400.3 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 404.3 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 402.3 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 388.3 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 404.3 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 400.3 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 388.3 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 400.3 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 380.3 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 388.3 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 396.3 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 332.3 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 332.3 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 376.3 GB short Won't fit
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 303.2 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 307.2 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 305.2 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 291.2 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 307.2 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 303.2 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 291.2 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 303.2 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 283.2 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 291.2 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 299.2 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 235.2 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 235.2 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 279.2 GB short Won't fit

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 0.54 · Overhead 1.00 16 1911.8 0.54 1.00 1913.4
Q8_0Bits 8.5 · Cache 0.54 · Overhead 1.00 8.5 1015.7 0.54 1.00 1017.2
Q6_KBits 6.5625 · Cache 0.54 · Overhead 1.00 6.5625 784.2 0.54 1.00 785.7
Q5_K_MBits 5.5 · Cache 0.54 · Overhead 1.00 5.5 657.2 0.54 1.00 658.7
Q4_K_MBits 4.5 · Cache 0.54 · Overhead 1.00 4.5 537.7 0.54 1.00 539.2
Q3_K_MBits 3.4375 · Cache 0.54 · Overhead 1.00 3.4375 410.8 0.54 1.00 412.3
Q2_KBits 2.625 · Cache 0.54 · Overhead 1.00 2.625 313.7 0.54 1.00 315.2

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 538.8 GB
  • 8,192 539.2 GB
  • 32,768 540.9 GB
  • 131,072 547.3 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total
2048FP16 total 1913.0 0.13 538.8 1913.0
8192FP16 total 1913.4 0.54 539.2 1913.4
32768FP16 total 1915.0 2.14 540.9 1915.0
131072FP16 total 1921.4 8.58 547.3 1921.4

Kimi K2 hardware notes

  • Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
  • Last updated 2026-04-23.
  • Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
  • Architecture numbers and the parameter count come from the public model record, checked 2026-10-10.

Kimi K2 VRAM FAQ

How much VRAM does Kimi K2 need at Q4_K_M?

539.2 GB at 8,192 tokens. Weights are 537.7 GB, the cache is 0.54 GB, and overhead is 1 GB.

Does a 24 GB GPU fit Kimi K2 at Q4_K_M?

GeForce RTX 3090 24GB is Won't fit, 515.2 GB over. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 131,072.

What context length is configured for Kimi K2?

The configured context length is 131,072 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 131,072.

Which source supplies the parameter count for Kimi K2?

1026.4B parameters, from the Hugging Face model record for moonshotai/Kimi-K2-Instruct.

Models in the same family, then the closest Q4_K_M totals.
Model Q4_K_M GB
DeepSeek V3 360.1
DeepSeek R1 360.1
gpt-oss-120b 62.49
Llama 3.1 70B 40.46