LLM VRAM Calculator

Estimate how much VRAM a local model needs at each quantization, then see which of 20 GPUs can hold it.

31.27B

30 models.

Quantization 4.5 bits

Q4_K_M is the usual starting point for 8 to 24 GB cards.

Optional. Pick your card to get a verdict and see it highlighted in the list.

Estimated VRAM

Smallest card that fits: RTX 3090, with 5.52 GB to spare

Fits at Q4_K_M with 8,192 tokens of context, with room left for your desktop and browser.

18.48 GB

Gemma 4 31B · Q4_K_M · 8,192 tokens

Weights

16.38 GB

params × bits ÷ 8

KV cache

1.09 GB

2 × layers × KV heads × head dim × tokens × 2 B

Overhead

1.00 GB

fixed buffers and context

Enough room to step up to Q5_K_MUses 22.12 GB, 1.88 GB to spare

Which GPUs fit

Cards that fit first 8 of 20 shown

Estimate, not a measurement. Usage varies with runtime, batch size, and drivers. How we calculate.