Short answer
Memory decides which local model you can run. This page lists the models at Q4_K_M, using the same estimate as the calculator: weights, cache, and 1 GB of overhead.
For one GPU, VRAM is the usual limit. System RAM matters when layers move off the card. A unified-memory Mac counts 75 percent of the pool as usable, because the operating system shares the rest.
Fits means the estimate is at or under usable memory and at least 10 percent of that memory remains. Tight fit means it fits with less than 10 percent free. CPU offload means the card clears 6 GB and a quarter of the weights, but not the full total. Won't fit means it clears neither. The words carry the meaning. Color only repeats them.
Models at the default context
The default context in the table is 8,192 tokens, or the model's maximum when that maximum is lower. Q4_K_M here is a nominal block size applied to the parameter count, not a measured file.
| Model | Parameters | Q4_K_M | FP16 | Context | Smallest GPU with 10 percent free |
|---|---|---|---|---|---|
| Gemma 4Parameters 31.27B · FP16 60.34 · Context 8192 | 31.27B | 18.48 | 60.34 | 8192 | GeForce RTX 3090 24GB |
| gpt-oss-120bParameters 116.8B · FP16 218.9 · Context 8192 | 116.8B | 62.49 | 218.9 | 8192 | NVIDIA A100 80GB |
| gpt-oss-20bParameters 20.91B · FP16 40.15 · Context 8192 | 20.91B | 12.15 | 40.15 | 8192 | GeForce RTX 4060 Ti 16GB |
| Kimi K2Parameters 1026.4B · FP16 1913.4 · Context 8192 | 1026.4B | 539.2 | 1913.4 | 8192 | No single GPU in our list fits this; see multi-GPU or CPU offload |
| DeepSeek R1Parameters 684.5B · FP16 1276.5 · Context 8192 | 684.5B | 360.1 | 1276.5 | 8192 | No single GPU in our list fits this; see multi-GPU or CPU offload |
| Llama 3.3 70BParameters 70.55B · FP16 134.9 · Context 8192 | 70.55B | 40.46 | 134.9 | 8192 | NVIDIA A100 80GB |
| Llama 3.1 8BParameters 8.03B · FP16 16.96 · Context 8192 | 8.03B | 6.21 | 16.96 | 8192 | GeForce RTX 3070 8GB |
| Llama 3.1 70BParameters 70.55B · FP16 134.9 · Context 8192 | 70.55B | 40.46 | 134.9 | 8192 | NVIDIA A100 80GB |
| DeepSeek V3Parameters 684.5B · FP16 1276.6 · Context 8192 | 684.5B | 360.1 | 1276.6 | 8192 | No single GPU in our list fits this; see multi-GPU or CPU offload |
| Llama 3 8BParameters 8.03B · FP16 16.96 · Context 8192 | 8.03B | 6.21 | 16.96 | 8192 | GeForce RTX 3070 8GB |
Runners
Ollama, llama.cpp, and LM Studio are local runners. This page does not measure their speed and does not document llama.cpp or LM Studio defaults.
Ollama's own FAQ is quoted on the Ollama system requirements page: a 4096-token default context, a 5-minute keep-alive, and a rule that a new GPU-loaded model has to fit in VRAM for concurrent loads.
Pick a model in the calculator, then open its page for the quant table and the sources.