Short answer
Here's what fits on 8 GB, 16 GB and 24 GB+ cards at Q4_K_M, the most common 4-bit setting.
A model is listed once, under the smallest tier that holds its Q4_K_M total. Anything above 16 GB is in the last section, including models that fit in 24 GB and models that fit in no single card here.
The total is weights plus cache at the default context plus 1 GB of overhead. It is an estimate.
8 GB
These models have a Q4_K_M total at or under 8 GB.
| Model | Parameters | Q4_K_M GB | Context used |
|---|---|---|---|
| Llama 3.1 8BParameters 8.03B | 8.03B | 6.21 | 8192 |
| Llama 3 8BParameters 8.03B | 8.03B | 6.21 | 8192 |
16 GB
These models need more than 8 GB and no more than 16 GB at Q4_K_M.
| Model | Parameters | Q4_K_M GB | Context used |
|---|---|---|---|
| gpt-oss-20bParameters 20.91B | 20.91B | 12.15 | 8192 |
24 GB and above
These models need more than 16 GB at Q4_K_M. The last column is the smallest memory in the 20-GPU table that can hold the total, with no extra 10 percent rule. Some totals exceed every card.
| Model | Parameters | Q4_K_M GB | Context used | Smallest memory that holds the total |
|---|---|---|---|---|
| Gemma 4Parameters 31.27B · Context used 8192 | 31.27B | 18.48 | 8192 | GeForce RTX 3090 24GB |
| Llama 3.3 70BParameters 70.55B · Context used 8192 | 70.55B | 40.46 | 8192 | NVIDIA A100 80GB |
| Llama 3.1 70BParameters 70.55B · Context used 8192 | 70.55B | 40.46 | 8192 | NVIDIA A100 80GB |
| gpt-oss-120bParameters 116.8B · Context used 8192 | 116.8B | 62.49 | 8192 | NVIDIA A100 80GB |
| DeepSeek R1Parameters 684.5B · Context used 8192 | 684.5B | 360.1 | 8192 | No single GPU in our list fits this; see multi-GPU or CPU offload |
| DeepSeek V3Parameters 684.5B · Context used 8192 | 684.5B | 360.1 | 8192 | No single GPU in our list fits this; see multi-GPU or CPU offload |
| Kimi K2Parameters 1026.4B · Context used 8192 | 1026.4B | 539.2 | 8192 | No single GPU in our list fits this; see multi-GPU or CPU offload |
The GPU guide uses a stricter rule: 10 percent of the card left unused. A model can appear in a tier here and still fail that headroom test on the same card.