Qwen3 235B-A22B VRAM and GPU Requirements

Qwen/Qwen3-235B-A22B · apache-2.0 · Updated 2025-07-26

Parameters
235.1Bmodel record
Default context
8,192tokens
At Q4_K_M
125.6 GB8,192 tokens
At FP16
440.4 GBfull estimate
Smallest fit
No single cardNo single GPU in our list fits this; see multi-GPU or CPU offload

No single GPU in our list fits this; see multi-GPU or CPU offload

125.6 GB

Qwen3 235B-A22B · Q4_K_M · 8,192 tokens

Weights

123.2 GB

params × bits ÷ 8

KV cache

1.47 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Open in calculator

Qwen3 235B-A22B needs about 125.6 GB at Q4_K_M with 8K context.

Qwen3 235B-A22B wears a dense-looking name on a mixture: 128 experts, 8 active, 94 layers, which is the longest layer count among the Qwen3 pages here. Query heads are 64 over 4 key heads, 16 to 1. The expert feed-forward width is much smaller than the dense 32B file's feed-forward.

Apache 2.0, public, no native low-bit packing, context 40,960 tokens. Q4_K_M at 8,192 tokens is 125.6 GB. No single GPU in our list fits this; see multi-GPU or CPU offload.

The coder mixtures are a different page, with a much longer context and different expert counts.

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 428.4 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 432.4 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 430.4 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 416.4 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 432.4 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 428.4 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 416.4 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 428.4 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 408.4 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 416.4 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 424.4 GB short Won't fit
NVIDIA A100 80GB80 GB · Won't fit 80 GB 360.4 GB short Won't fit
NVIDIA H100 SXM 80GB80 GB · Won't fit 80 GB 360.4 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 404.4 GB short Won't fit
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 223.1 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 227.1 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 225.1 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 211.1 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 227.1 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 223.1 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 211.1 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 223.1 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 203.1 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 211.1 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 219.1 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 155.1 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 155.1 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 199.1 GB short Won't fit
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 170.1 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 174.1 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 172.1 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 158.1 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 174.1 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 170.1 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 158.1 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 170.1 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 150.1 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 158.1 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 166.1 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 102.1 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 102.1 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 146.1 GB short Won't fit
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 141.0 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 145.0 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 143.0 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 129.0 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 145.0 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 141.0 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 129.0 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 141.0 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 121.0 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 129.0 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 137.0 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 73.00 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 73.00 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 117.0 GB short Won't fit
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 113.6 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 117.6 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 115.6 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 101.6 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 117.6 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 113.6 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 101.6 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 113.6 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 93.63 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 101.6 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 109.6 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 45.63 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 45.63 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 89.63 GB short CPU offload
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 84.55 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 88.55 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 86.55 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 72.55 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 88.55 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 84.55 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 72.55 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 84.55 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 64.55 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 72.55 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 80.55 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 16.55 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 16.55 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 60.55 GB short CPU offload
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · Tight fit 80 GB Tight fit+5.69 GB +5.69 GB
NVIDIA H100 SXM 80GB80 GB · Tight fit 80 GB Tight fit+5.69 GB +5.69 GB
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 62.31 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 66.31 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 64.31 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 50.31 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 66.31 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 62.31 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 50.31 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 62.31 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 42.31 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 50.31 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 58.31 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 38.31 GB short CPU offload

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 1.47 · Overhead 1.00 16 437.9 1.47 1.00 440.4
Q8_0Bits 8.5 · Cache 1.47 · Overhead 1.00 8.5 232.6 1.47 1.00 235.1
Q6_KBits 6.5625 · Cache 1.47 · Overhead 1.00 6.5625 179.6 1.47 1.00 182.1
Q5_K_MBits 5.5 · Cache 1.47 · Overhead 1.00 5.5 150.5 1.47 1.00 153.0
Q4_K_MBits 4.5 · Cache 1.47 · Overhead 1.00 4.5 123.2 1.47 1.00 125.6
Q3_K_MBits 3.4375 · Cache 1.47 · Overhead 1.00 3.4375 94.08 1.47 1.00 96.55
Q2_KBits 2.625 · Cache 1.47 · Overhead 1.00 2.625 71.84 1.47 1.00 74.31

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 124.5 GB
  • 8,192 125.6 GB
  • 32,768 130.0 GB
  • 40,960 131.5 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total
2048FP16 total 439.3 0.37 124.5 439.3
8192FP16 total 440.4 1.47 125.6 440.4
32768FP16 total 444.8 5.88 130.0 444.8
40960FP16 total 446.2 7.34 131.5 446.2

Qwen3 235B-A22B hardware notes

  • Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
  • Last updated 2025-07-26.
  • Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
  • Architecture numbers and the parameter count come from the public model record, checked 2026-10-10.

Qwen3 235B-A22B VRAM FAQ

How much VRAM does Qwen3 235B-A22B need at Q4_K_M?

125.6 GB at 8,192 tokens. Weights are 123.2 GB, the cache is 1.47 GB, and overhead is 1 GB.

Does a 24 GB GPU fit Qwen3 235B-A22B at Q4_K_M?

GeForce RTX 3090 24GB is Won't fit, 101.6 GB over. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 40,960.

What context length is configured for Qwen3 235B-A22B?

The configured context length is 40,960 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 40,960.

Which source supplies the parameter count for Qwen3 235B-A22B?

235.1B parameters, from the Hugging Face model record for Qwen/Qwen3-235B-A22B.

Models in the same family, then the closest Q4_K_M totals.
Model Q4_K_M GB
gpt-oss-120b 62.49
Llama 3.1 70B 40.46
Llama 3.3 70B 40.46
Gemma 4 18.48