gpt-oss-120b VRAM and GPU Requirements

openai/gpt-oss-120b · apache-2.0 · Updated 2025-08-26

Parameters
116.8Bmodel record
Default context
8,192tokens
At Q4_K_M
62.49 GB8,192 tokens
At FP16
218.9 GBfull estimate
Smallest fit
NVIDIA A100 80GB10 percent free

Runs on the A100 with 17.51 GB to spare

Fits at Q4_K_M with 8,192 tokens of context, with room left for your desktop and browser.

62.49 / 80 GB

gpt-oss-120b · Q4_K_M · 8,192 tokens

A100 · 80 GB

Weights

61.20 GB

params × bits ÷ 8

KV cache

0.29 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Enough room to step up to Q5_K_MUses 76.09 GB, 3.91 GB to spare

Open in calculator

gpt-oss-120b needs about 62.49 GB at Q4_K_M with 8K context.

gpt-oss-120b is a large reasoning mixture: 36 layers and 128 experts, with 4 experts active on each token, so you store every expert and run only a few. The publisher's own card says about 5.1B parameters are active and that the expert weights use a 4-bit packing. This page's Q4_K_M row is a different question: 62.49 GB at 8,192 tokens, applied to the full 116.8B parameter record.

The smallest listed GPU that still leaves 10 percent free is NVIDIA A100 80GB, with 17.51 GB left. A separate row uses the measured checkpoint bytes, plus an estimated cache and 1 GB of overhead. Attention alternates a short sliding window with full layers, and the configured context is 131,072 tokens.

The 20B sibling is a smaller expert inventory on its own page.

Which GPUs fit

Fit at FP16/BF16 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 206.9 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 210.9 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 208.9 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 194.9 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 210.9 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 206.9 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 194.9 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 206.9 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
GeForce RTX 5090 32GB32 GB · Won't fit 32 GB 186.9 GB short Won't fit
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 194.9 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 202.9 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 138.9 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 138.9 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · Won't fit 48 GB 182.9 GB short Won't fit
Fit at Q8_0 and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 104.9 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 108.9 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 106.9 GB short Won't fit
GeForce RTX 3090 24GB24 GB · Won't fit 24 GB 92.89 GB short Won't fit
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 108.9 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 104.9 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
GeForce RTX 4090 24GB24 GB · Won't fit 24 GB 92.89 GB short Won't fit
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 104.9 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 84.89 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · Won't fit 24 GB 92.89 GB short Won't fit
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 100.9 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 36.89 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 36.89 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 80.89 GB short CPU offload
Fit at Q6_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 78.54 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 82.54 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 80.54 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 66.54 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 82.54 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 78.54 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 66.54 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 78.54 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 58.54 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 66.54 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 74.54 GB short Won't fit
NVIDIA A100 80GB80 GB · CPU offload 80 GB 10.54 GB short CPU offload
NVIDIA H100 SXM 80GB80 GB · CPU offload 80 GB 10.54 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 54.54 GB short CPU offload
Fit at Q5_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · Tight fit 80 GB Tight fit+3.91 GB +3.91 GB
NVIDIA H100 SXM 80GB80 GB · Tight fit 80 GB Tight fit+3.91 GB +3.91 GB
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 64.09 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 68.09 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 66.09 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 52.09 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 68.09 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 64.09 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
GeForce RTX 4080 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 52.09 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 64.09 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
GeForce RTX 5080 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 44.09 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 52.09 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · Won't fit 16 GB 60.09 GB short Won't fit
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 40.09 GB short CPU offload
Fit at Q4_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+17.51 GB +17.51 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+17.51 GB +17.51 GB
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 50.49 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 54.49 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 52.49 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 38.49 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 54.49 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 50.49 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 38.49 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 50.49 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 30.49 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 38.49 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 46.49 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 26.49 GB short CPU offload
Fit at Q3_K_M and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+31.96 GB +31.96 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+31.96 GB +31.96 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 36.04 GB short CPU offload
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 40.04 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 38.04 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 24.04 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 40.04 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 36.04 GB short CPU offload
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 24.04 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 36.04 GB short CPU offload
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 16.04 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 24.04 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 32.04 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 12.04 GB short CPU offload
Fit at Q2_K and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+43.01 GB +43.01 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+43.01 GB +43.01 GB
GeForce RTX 3060 12GB12 GB · CPU offload 12 GB 24.99 GB short CPU offload
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 28.99 GB short Won't fit
GeForce RTX 3080 10GB10 GB · CPU offload 10 GB 26.99 GB short CPU offload
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 12.99 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 28.99 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
GeForce RTX 4070 12GB12 GB · CPU offload 12 GB 24.99 GB short CPU offload
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 12.99 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
GeForce RTX 5070 12GB12 GB · CPU offload 12 GB 24.99 GB short CPU offload
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 4.99 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 12.99 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 20.99 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 0.99 GB short CPU offload
Fit at Published MXFP4 files and 8,192 tokens. Cards that fit in usable memory are listed first. "Short" shows how far the model goes over usable memory. Headroom uses usable memory. A unified row uses 75 percent of the listed pool.
GPU Memory Status Headroom
NVIDIA A100 80GB80 GB · No CPU offload 80 GB Fits+17.95 GB +17.95 GB
NVIDIA H100 SXM 80GB80 GB · No CPU offload 80 GB Fits+17.95 GB +17.95 GB
GeForce RTX 3060 12GB12 GB · Won't fit 12 GB 50.05 GB short Won't fit
GeForce RTX 3070 8GB8 GB · Won't fit 8 GB 54.05 GB short Won't fit
GeForce RTX 3080 10GB10 GB · Won't fit 10 GB 52.05 GB short Won't fit
GeForce RTX 3090 24GB24 GB · CPU offload 24 GB 38.05 GB short CPU offload
GeForce RTX 4060 8GB8 GB · Won't fit 8 GB 54.05 GB short Won't fit
GeForce RTX 4060 Ti 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
GeForce RTX 4070 12GB12 GB · Won't fit 12 GB 50.05 GB short Won't fit
GeForce RTX 4070 Ti Super 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
GeForce RTX 4080 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
GeForce RTX 4090 24GB24 GB · CPU offload 24 GB 38.05 GB short CPU offload
GeForce RTX 5060 Ti 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
GeForce RTX 5070 12GB12 GB · Won't fit 12 GB 50.05 GB short Won't fit
GeForce RTX 5070 Ti 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
GeForce RTX 5080 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
GeForce RTX 5090 32GB32 GB · CPU offload 32 GB 30.05 GB short CPU offload
Radeon RX 7900 XTX 24GB24 GB · CPU offload 24 GB 38.05 GB short CPU offload
Radeon RX 9070 XT 16GB16 GB · CPU offload 16 GB 46.05 GB short CPU offload
Apple M5 Max 48GB unified48 GB · 36.00 GB usable · CPU offload 48 GB 26.05 GB short CPU offload

VRAM by quantization

Totals at 8,192 tokens. Q4_K_M is the usual balance of size and quality. FP16 weights are the parameter count times 2 bytes. The FP16 column in the table is the full estimate: weights, cache, and 1 GB.

Estimates at 8192 tokens, in GB. Formula rows apply one block size to every parameter. A published-file row uses measured weight bytes. FP16 weights equal the parameter count times 2 bytes.
Quantization Bits Weights Cache Overhead Total
FP16/BF16Bits 16 · Cache 0.29 · Overhead 1.00 16 217.6 0.29 1.00 218.9
Q8_0Bits 8.5 · Cache 0.29 · Overhead 1.00 8.5 115.6 0.29 1.00 116.9
Q6_KBits 6.5625 · Cache 0.29 · Overhead 1.00 6.5625 89.25 0.29 1.00 90.54
Q5_K_MBits 5.5 · Cache 0.29 · Overhead 1.00 5.5 74.80 0.29 1.00 76.09
Q4_K_MBits 4.5 · Cache 0.29 · Overhead 1.00 4.5 61.20 0.29 1.00 62.49
Q3_K_MBits 3.4375 · Cache 0.29 · Overhead 1.00 3.4375 46.75 0.29 1.00 48.04
Q2_KBits 2.625 · Cache 0.29 · Overhead 1.00 2.625 35.70 0.29 1.00 36.99
Published MXFP4 filesBits 4.47 · Cache 0.29 · Overhead 1.00 4.47 60.77 0.29 1.00 62.05

Context length

Q4_K_M total as the context changes. Cache is the part that grows.

  • 2,048 62.28 GB
  • 8,192 62.49 GB
  • 32,768 63.33 GB
  • 131,072 66.71 GB
Cache and totals as the context changes. Tokens above the configured maximum are not listed.
Context tokens Cache GB Q4_K_M total FP16 total Published-file total
2048FP16 total 218.7 · Published-file total 61.84 0.07 62.28 218.7 61.84
8192FP16 total 218.9 · Published-file total 62.05 0.29 62.49 218.9 62.05
32768FP16 total 219.7 · Published-file total 62.90 1.13 63.33 219.7 62.90
131072FP16 total 223.1 · Published-file total 66.27 4.50 66.71 223.1 66.27

gpt-oss-120b hardware notes

  • Unified memory counts 75 percent of the listed pool as usable. Discrete cards use the full listed memory. The memory column still shows the listed figure.
  • Last updated 2025-08-26.
  • Weights, cache, and the 1 GB overhead are estimates. How this is calculated.
  • Architecture numbers and the parameter count come from the public model record, checked 2026-10-10.
  • The published-file row uses a measured byte sum for the weights. Cache and overhead on that row are still estimates.

gpt-oss-120b VRAM FAQ

How much VRAM does gpt-oss-120b need at Q4_K_M?

62.49 GB at 8,192 tokens. Weights are 61.20 GB, the cache is 0.29 GB, and overhead is 1 GB.

Does a 24 GB GPU fit gpt-oss-120b at Q4_K_M?

GeForce RTX 3090 24GB is CPU offload, 38.49 GB over. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 131,072.

What context length is configured for gpt-oss-120b?

The configured context length is 131,072 tokens. The quant table uses 8,192 tokens, the smaller of 8,192 and the configured maximum of 131,072.

Which source supplies the parameter count for gpt-oss-120b?

116.8B parameters, from the Hugging Face model record for openai/gpt-oss-120b.

Models in the same family, then the closest Q4_K_M totals.
Model Q4_K_M GB
gpt-oss-20b 12.15
Llama 3.1 70B 40.46
Llama 3.3 70B 40.46
Gemma 4 18.48
Llama 3 8B 6.21