Best Local LLMs for Your GPU

By GPUFit AI · Updated 2026-10-10

Short answer

Here's what fits on 8 GB, 16 GB and 24 GB+ cards at Q4_K_M, the most common 4-bit setting.

A model is listed once, under the smallest tier that holds its Q4_K_M total. Anything above 16 GB is in the last section, including models that fit in 24 GB and models that fit in no single card here.

The total is weights plus cache at the default context plus 1 GB of overhead. It is an estimate.

8 GB

These models have a Q4_K_M total at or under 8 GB.

One row per model in this tier. GB figures are estimates.
Model Parameters Q4_K_M GB Context used
Llama 3.1 8BParameters 8.03B 8.03B 6.21 8192
Llama 3 8BParameters 8.03B 8.03B 6.21 8192

16 GB

These models need more than 8 GB and no more than 16 GB at Q4_K_M.

One row per model in this tier. GB figures are estimates.
Model Parameters Q4_K_M GB Context used
gpt-oss-20bParameters 20.91B 20.91B 12.15 8192

24 GB and above

These models need more than 16 GB at Q4_K_M. The last column is the smallest memory in the 20-GPU table that can hold the total, with no extra 10 percent rule. Some totals exceed every card.

One row per model in this tier. GB figures are estimates.
Model Parameters Q4_K_M GB Context used Smallest memory that holds the total
Gemma 4Parameters 31.27B · Context used 8192 31.27B 18.48 8192 GeForce RTX 3090 24GB
Llama 3.3 70BParameters 70.55B · Context used 8192 70.55B 40.46 8192 NVIDIA A100 80GB
Llama 3.1 70BParameters 70.55B · Context used 8192 70.55B 40.46 8192 NVIDIA A100 80GB
gpt-oss-120bParameters 116.8B · Context used 8192 116.8B 62.49 8192 NVIDIA A100 80GB
DeepSeek R1Parameters 684.5B · Context used 8192 684.5B 360.1 8192 No single GPU in our list fits this; see multi-GPU or CPU offload
DeepSeek V3Parameters 684.5B · Context used 8192 684.5B 360.1 8192 No single GPU in our list fits this; see multi-GPU or CPU offload
Kimi K2Parameters 1026.4B · Context used 8192 1026.4B 539.2 8192 No single GPU in our list fits this; see multi-GPU or CPU offload

The GPU guide uses a stricter rule: 10 percent of the card left unused. A model can appear in a tier here and still fail that headroom test on the same card.