Ollama System Requirements

By GPUFit AI · Updated 2026-10-10

Short answer

Ollama does not publish a per-model RAM table that this site can copy, so this page does not invent one. The numbers below are this site's weight-plus-cache estimates. The sentences in quotation marks are from Ollama's FAQ, fetched 2026-10-10.

What the Ollama FAQ states

The 4096-token default is why the table has a column at 4096 as well as at this site's 8192-token default. A parallel setting can multiply the context Ollama allocates. That multiplier is not included in the estimate.

System RAM is the pool Ollama uses for a CPU load. The FAQ points at system memory for CPU inference and VRAM for GPU inference. It does not give a RAM formula, and neither does this page.

By default, Ollama uses a context window size of 4096 tokens.

By default models are kept in memory for 5 minutes before being unloaded.

When using GPU inference new models must be able to completely fit in VRAM to allow concurrent model loads.

Parallel request processing for a given model results in increasing the context size by the number of parallel requests. For example, a 2K context with 4 parallel requests will result in an 8K context and additional memory allocation.

Estimates for these models

Q4_K_M and FP16 estimates in GB. The 4096 column matches Ollama's documented default context. FP16 is the full estimate at this site's default context. FP16 weights alone are the parameter count times 2 bytes.
Model Q4_K_M at 4096 Q4_K_M at default FP16 at default Max context
Gemma 4FP16 at default 60.34 · Max context 262144 18.32 18.48 60.34 262144
gpt-oss-120bFP16 at default 218.9 · Max context 131072 62.35 62.49 218.9 131072
gpt-oss-20bFP16 at default 40.15 · Max context 131072 12.05 12.15 40.15 131072
Kimi K2FP16 at default 1913.4 · Max context 131072 539.0 539.2 1913.4 131072
DeepSeek R1FP16 at default 1276.5 · Max context 163840 359.9 360.1 1276.5 163840
Llama 3.3 70BFP16 at default 134.9 · Max context 131072 39.21 40.46 134.9 131072
Llama 3.1 8BFP16 at default 16.96 · Max context 131072 5.71 6.21 16.96 131072
Llama 3.1 70BFP16 at default 134.9 · Max context 131072 39.21 40.46 134.9 131072
DeepSeek V3FP16 at default 1276.6 · Max context 163840 359.9 360.1 1276.6 163840
Llama 3 8BFP16 at default 16.96 · Max context 8192 5.71 6.21 16.96 8192