Qwen3 Coder VRAM and GPU Requirements

Qwen/Qwen3-Coder-30B-A3B-Instruct · apache-2.0 · Updated 2025-12-03

Parameters
30.53Bmodel record
Default context
8,192tokens
At Q4_K_M
17.74 GB8,192 tokens
At FP16
58.62 GBfull estimate
Smallest fit
GeForce RTX 3090 24GB10 percent free

Runs on the RTX 3090 with 6.26 GB to spare

Fits at Q4_K_M with 8,192 tokens of context, with room left for your desktop and browser.

17.74 / 24 GB

Qwen3 Coder 30B-A3B · Q4_K_M · 8,192 tokens

RTX 3090 · 24 GB

Weights

15.99 GB

params × bits ÷ 8

KV cache

0.75 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Enough room to step up to Q5_K_MUses 21.30 GB, 2.70 GB to spare

Open in calculator

Qwen3 Coder 30B-A3B needs about 17.74 GB at Q4_K_M with 8K context.

Qwen3 Coder is 1 URL with 2 mixture sizes. The lead is the 30B-A3B model. The 480B block is a second full table, not a footnote.

Both are Apache 2.0, public, and stored without a native low-bit packing, so there is no measured packed-file row. Both context maxima are 262,144 tokens. 30B-A3B stores 30.53B parameters, 48 layers, 128 experts with 8 active, context 262,144 tokens, Q4_K_M 17.74 GB at 8,192 tokens.

480B-A35B stores 480.2B parameters, 62 layers, 160 experts with 8 active, context 262,144 tokens, Q4_K_M 254.5 GB at 8,192 tokens. For the lead, The smallest listed GPU that still leaves 10 percent free is GeForce RTX 3090 24GB, with 6.26 GB left.

Sizes

Each size links to its own tables. Totals use that size's default context. FP16 is the full estimate. FP16 weights alone are the parameter count times 2 bytes.
Size Parameters Max tokens Q4_K_M FP16 Smallest GPU with 10 percent free
30B-A3BParameters 30.53B · Max tokens 262,144 · FP16 58.62 30.53B 262,144 17.74 58.62 GeForce RTX 3090 24GB
480B-A35BParameters 480.2B · Max tokens 262,144 · FP16 897.3 480.2B 262,144 254.5 897.3 No single GPU in our list fits this; see multi-GPU or CPU offload

30B-A3B

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+21.38 GB +21.38 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+21.38 GB +21.38 GB
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 46.62 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 50.62 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 48.62 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 34.62 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 50.62 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 46.62 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 34.62 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 46.62 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 26.62 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 34.62 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 42.62 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 22.62 GB short CPU offload
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 5090 32GB32 GB · Tight fit 32 GB Tight fit+0.04 GB +0.04 GB
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+48.04 GB +48.04 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+48.04 GB +48.04 GB
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · No CPU offload 48 GB Fits+4.04 GB +4.04 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 19.96 GB short CPU offload
GeForce RTX 3070 8GB8 GB · CPU offload 8 GB 23.96 GB short CPU offload
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 21.96 GB short CPU offload
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 7.96 GB short CPU offload
GeForce RTX 4060 8GB8 GB · CPU offload 8 GB 23.96 GB short CPU offload
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 19.96 GB short CPU offload
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 7.96 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 19.96 GB short CPU offload
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 7.96 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 15.96 GB short CPU offload
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 5090 32GB32 GB · No CPU offload 32 GB Fits+6.92 GB +6.92 GB
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+54.92 GB +54.92 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+54.92 GB +54.92 GB
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · No CPU offload 48 GB Fits+10.92 GB +10.92 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 13.08 GB short CPU offload
GeForce RTX 3070 8GB8 GB · CPU offload 8 GB 17.08 GB short CPU offload
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 15.08 GB short CPU offload
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 1.08 GB short CPU offload
GeForce RTX 4060 8GB8 GB · CPU offload 8 GB 17.08 GB short CPU offload
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 13.08 GB short CPU offload
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 1.08 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 13.08 GB short CPU offload
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 1.08 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 9.08 GB short CPU offload
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3090 24GB24 GB · No CPU offload 24 GB Fits+2.70 GB +2.70 GB
GeForce RTX 4090 24GB24 GB · No CPU offload 24 GB Fits+2.70 GB +2.70 GB
GeForce RTX 5090 32GB32 GB · No CPU offload 32 GB Fits+10.70 GB +10.70 GB
Radeon RX 7900 XTX 24GB24 GB · No CPU offload 24 GB Fits+2.70 GB +2.70 GB
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+58.70 GB +58.70 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+58.70 GB +58.70 GB
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · No CPU offload 48 GB Fits+14.70 GB +14.70 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 9.30 GB short CPU offload
GeForce RTX 3070 8GB8 GB · CPU offload 8 GB 13.30 GB short CPU offload
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 11.30 GB short CPU offload
GeForce RTX 4060 8GB8 GB · CPU offload 8 GB 13.30 GB short CPU offload
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 9.30 GB short CPU offload
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 9.30 GB short CPU offload
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 5.30 GB short CPU offload
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3090 24GB24 GB · No CPU offload 24 GB Fits+6.26 GB +6.26 GB
GeForce RTX 4090 24GB24 GB · No CPU offload 24 GB Fits+6.26 GB +6.26 GB
GeForce RTX 5090 32GB32 GB · No CPU offload 32 GB Fits+14.26 GB +14.26 GB
Radeon RX 7900 XTX 24GB24 GB · No CPU offload 24 GB Fits+6.26 GB +6.26 GB
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+62.26 GB +62.26 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+62.26 GB +62.26 GB
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · No CPU offload 48 GB Fits+18.26 GB +18.26 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 5.74 GB short CPU offload
GeForce RTX 3070 8GB8 GB · CPU offload 8 GB 9.74 GB short CPU offload
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 7.74 GB short CPU offload
GeForce RTX 4060 8GB8 GB · CPU offload 8 GB 9.74 GB short CPU offload
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 5.74 GB short CPU offload
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 5.74 GB short CPU offload
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 1.74 GB short CPU offload
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3090 24GB24 GB · No CPU offload 24 GB Fits+10.03 GB +10.03 GB
GeForce RTX 4060 Ti 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
GeForce RTX 4070 Ti Super 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
GeForce RTX 4080 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
GeForce RTX 4090 24GB24 GB · No CPU offload 24 GB Fits+10.03 GB +10.03 GB
GeForce RTX 5060 Ti 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
GeForce RTX 5070 Ti 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
GeForce RTX 5080 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
GeForce RTX 5090 32GB32 GB · No CPU offload 32 GB Fits+18.03 GB +18.03 GB
Radeon RX 7900 XTX 24GB24 GB · No CPU offload 24 GB Fits+10.03 GB +10.03 GB
Radeon RX 9070 XT 16GB16 GB · No CPU offload 16 GB Fits+2.03 GB +2.03 GB
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+66.03 GB +66.03 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+66.03 GB +66.03 GB
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · No CPU offload 48 GB Fits+22.03 GB +22.03 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 1.97 GB short CPU offload
GeForce RTX 3070 8GB8 GB · CPU offload 8 GB 5.97 GB short CPU offload
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 3.97 GB short CPU offload
GeForce RTX 4060 8GB8 GB · CPU offload 8 GB 5.97 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 1.97 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 1.97 GB short CPU offload
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Tight fit 12 GB Tight fit+0.92 GB +0.92 GB
GeForce RTX 3090 24GB24 GB · No CPU offload 24 GB Fits+12.92 GB +12.92 GB
GeForce RTX 4060 Ti 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
GeForce RTX 4070 12GB12 GB · Tight fit 12 GB Tight fit+0.92 GB +0.92 GB
GeForce RTX 4070 Ti Super 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
GeForce RTX 4080 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
GeForce RTX 4090 24GB24 GB · No CPU offload 24 GB Fits+12.92 GB +12.92 GB
GeForce RTX 5060 Ti 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
GeForce RTX 5070 12GB12 GB · Tight fit 12 GB Tight fit+0.92 GB +0.92 GB
GeForce RTX 5070 Ti 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
GeForce RTX 5080 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
GeForce RTX 5090 32GB32 GB · No CPU offload 32 GB Fits+20.92 GB +20.92 GB
Radeon RX 7900 XTX 24GB24 GB · No CPU offload 24 GB Fits+12.92 GB +12.92 GB
Radeon RX 9070 XT 16GB16 GB · No CPU offload 16 GB Fits+4.92 GB +4.92 GB
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+68.92 GB +68.92 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+68.92 GB +68.92 GB
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · No CPU offload 48 GB Fits+24.92 GB +24.92 GB
GeForce RTX 3070 8GB8 GB · CPU offload 8 GB 3.08 GB short CPU offload
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 1.08 GB short CPU offload
GeForce RTX 4060 8GB8 GB · CPU offload 8 GB 3.08 GB short CPU offload

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 0.75 · Overhead 1.00 16 56.87 0.75 1.00 58.62
Q8_0Bits 8.5 · Cache 0.75 · Overhead 1.00 8.5 30.21 0.75 1.00 31.96
Q6_KBits 6.5625 · Cache 0.75 · Overhead 1.00 6.5625 23.33 0.75 1.00 25.08
Q5_K_MBits 5.5 · Cache 0.75 · Overhead 1.00 5.5 19.55 0.75 1.00 21.30
Q4_K_MBits 4.5 · Cache 0.75 · Overhead 1.00 4.5 15.99 0.75 1.00 17.74
Q3_K_MBits 3.4375 · Cache 0.75 · Overhead 1.00 3.4375 12.22 0.75 1.00 13.97
Q2_KBits 2.625 · Cache 0.75 · Overhead 1.00 2.625 9.33 0.75 1.00 11.08

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 17.18 GB
  • 8,192 17.74 GB
  • 32,768 19.99 GB
  • 131,072 28.99 GB
  • 262,144 40.99 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total
2048FP16 total 58.06 0.19 17.18 58.06
8192FP16 total 58.62 0.75 17.74 58.62
32768FP16 total 60.87 3.00 19.99 60.87
131072FP16 total 69.87 12.00 28.99 69.87
262144FP16 total 81.87 24.00 40.99 81.87

480B-A35B

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 885.3 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 889.3 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 887.3 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 873.3 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 889.3 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 885.3 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 873.3 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 885.3 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 865.3 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 873.3 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 881.3 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 817.3 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 817.3 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 861.3 GB short Won't fit
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 466.1 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 470.1 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 468.1 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 454.1 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 470.1 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 466.1 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 454.1 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 466.1 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 446.1 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 454.1 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 462.1 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 398.1 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 398.1 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 442.1 GB short Won't fit
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 357.8 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 361.8 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 359.8 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 345.8 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 361.8 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 357.8 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 345.8 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 357.8 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 337.8 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 345.8 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 353.8 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 289.8 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 289.8 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 333.8 GB short Won't fit
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 298.4 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 302.4 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 300.4 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 286.4 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 302.4 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 298.4 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 286.4 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 298.4 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 278.4 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 286.4 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 294.4 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 230.4 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 230.4 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 274.4 GB short Won't fit
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 242.5 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 246.5 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 244.5 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 230.5 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 246.5 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 242.5 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 230.5 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 242.5 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 222.5 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 230.5 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 238.5 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 174.5 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 174.5 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 218.5 GB short Won't fit
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 183.1 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 187.1 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 185.1 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 171.1 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 187.1 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 183.1 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 171.1 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 183.1 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 163.1 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 171.1 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 179.1 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 115.1 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 115.1 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 159.1 GB short Won't fit
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 137.7 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 141.7 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 139.7 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 125.7 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 141.7 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 137.7 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 125.7 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 137.7 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 117.7 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 125.7 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 133.7 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 69.67 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 69.67 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 113.7 GB short Won't fit

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 1.94 · Overhead 1.00 16 894.4 1.94 1.00 897.3
Q8_0Bits 8.5 · Cache 1.94 · Overhead 1.00 8.5 475.1 1.94 1.00 478.1
Q6_KBits 6.5625 · Cache 1.94 · Overhead 1.00 6.5625 366.8 1.94 1.00 369.8
Q5_K_MBits 5.5 · Cache 1.94 · Overhead 1.00 5.5 307.4 1.94 1.00 310.4
Q4_K_MBits 4.5 · Cache 1.94 · Overhead 1.00 4.5 251.5 1.94 1.00 254.5
Q3_K_MBits 3.4375 · Cache 1.94 · Overhead 1.00 3.4375 192.2 1.94 1.00 195.1
Q2_KBits 2.625 · Cache 1.94 · Overhead 1.00 2.625 146.7 1.94 1.00 149.7

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 253.0 GB
  • 8,192 254.5 GB
  • 32,768 260.3 GB
  • 131,072 283.5 GB
  • 262,144 314.5 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total
2048FP16 total 895.8 0.48 253.0 895.8
8192FP16 total 897.3 1.94 254.5 897.3
32768FP16 total 903.1 7.75 260.3 903.1
131072FP16 total 926.4 31.00 283.5 926.4
262144FP16 total 957.4 62.00 314.5 957.4

Qwen3 Coder hardware notes

  • Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
  • Last updated 2025-12-03.
  • Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
  • Architecture numbers and the parameter count come from the public model record, checked 2026-10-10.

Qwen3 Coder VRAM FAQ

How much VRAM does Qwen3 Coder 30B-A3B need at Q4_K_M?

17.74 GB at 8,192 tokens. Weights are 15.99 GB, the cache is 0.75 GB, and overhead is 1 GB.

Does a 24 GB GPU fit Qwen3 Coder 30B-A3B at Q4_K_M?

GeForce RTX 3090 24GB is Fits, 6.26 GB left. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 262,144.

What context length is configured for Qwen3 Coder 30B-A3B?

The configured context length is 262,144 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 262,144.

Which source supplies the parameter count for Qwen3 Coder 30B-A3B?

30.53B parameters, from the Hugging Face model record for Qwen/Qwen3-Coder-30B-A3B-Instruct.

Models in the same family, then the closest Q4_K_M totals.
Model Q4_K_M GB
Gemma 4 18.48
gpt-oss-20b 12.15
Llama 3 8B 6.21
Llama 3.1 8B 6.21