Local LLMs: What You Need to Run AI at Home

By GPUFit AI · Updated 2026-10-10

Short answer

Memory decides which local model you can run. This page lists the models at Q4_K_M, using the same estimate as the calculator: weights, cache, and 1 GB of overhead.

For one GPU, VRAM is the usual limit. System RAM matters when layers move off the card. A unified-memory Mac counts 75 percent of the pool as usable, because the operating system shares the rest.

Fits means the estimate is at or under usable memory and at least 10 percent of that memory remains. Tight fit means it fits with less than 10 percent free. CPU offload means the card clears 6 GB and a quarter of the weights, but not the full total. Won't fit means it clears neither. The words carry the meaning. Color only repeats them.

Models at the default context

The default context in the table is 8,192 tokens, or the model's maximum when that maximum is lower. Q4_K_M here is a nominal block size applied to the parameter count, not a measured file.

Totals are estimates at the default context, in GB.
Model Parameters Q4_K_M FP16 Context Smallest GPU with 10 percent free
Gemma 4Parameters 31.27B · FP16 60.34 · Context 8192 31.27B 18.48 60.34 8192 GeForce RTX 3090 24GB
gpt-oss-120bParameters 116.8B · FP16 218.9 · Context 8192 116.8B 62.49 218.9 8192 NVIDIA A100 80GB
gpt-oss-20bParameters 20.91B · FP16 40.15 · Context 8192 20.91B 12.15 40.15 8192 GeForce RTX 4060 Ti 16GB
Kimi K2Parameters 1026.4B · FP16 1913.4 · Context 8192 1026.4B 539.2 1913.4 8192 No single GPU in our list fits this; see multi-GPU or CPU offload
DeepSeek R1Parameters 684.5B · FP16 1276.5 · Context 8192 684.5B 360.1 1276.5 8192 No single GPU in our list fits this; see multi-GPU or CPU offload
Llama 3.3 70BParameters 70.55B · FP16 134.9 · Context 8192 70.55B 40.46 134.9 8192 NVIDIA A100 80GB
Llama 3.1 8BParameters 8.03B · FP16 16.96 · Context 8192 8.03B 6.21 16.96 8192 GeForce RTX 3070 8GB
Llama 3.1 70BParameters 70.55B · FP16 134.9 · Context 8192 70.55B 40.46 134.9 8192 NVIDIA A100 80GB
DeepSeek V3Parameters 684.5B · FP16 1276.6 · Context 8192 684.5B 360.1 1276.6 8192 No single GPU in our list fits this; see multi-GPU or CPU offload
Llama 3 8BParameters 8.03B · FP16 16.96 · Context 8192 8.03B 6.21 16.96 8192 GeForce RTX 3070 8GB

Runners

Ollama, llama.cpp, and LM Studio are local runners. This page does not measure their speed and does not document llama.cpp or LM Studio defaults.

Ollama's own FAQ is quoted on the Ollama system requirements page: a 4096-token default context, a 5-minute keep-alive, and a rule that a new GPU-loaded model has to fit in VRAM for concurrent loads.

Pick a model in the calculator, then open its page for the quant table and the sources.