Short answer
Ollama does not publish a per-model RAM table that this site can copy, so this page does not invent one. The numbers below are this site's weight-plus-cache estimates. The sentences in quotation marks are from Ollama's FAQ, fetched 2026-10-10.
What the Ollama FAQ states
The 4096-token default is why the table has a column at 4096 as well as at this site's 8192-token default. A parallel setting can multiply the context Ollama allocates. That multiplier is not included in the estimate.
System RAM is the pool Ollama uses for a CPU load. The FAQ points at system memory for CPU inference and VRAM for GPU inference. It does not give a RAM formula, and neither does this page.
By default, Ollama uses a context window size of 4096 tokens.
By default models are kept in memory for 5 minutes before being unloaded.
When using GPU inference new models must be able to completely fit in VRAM to allow concurrent model loads.
Parallel request processing for a given model results in increasing the context size by the number of parallel requests. For example, a 2K context with 4 parallel requests will result in an 8K context and additional memory allocation.
Estimates for these models
| Model | Q4_K_M at 4096 | Q4_K_M at default | FP16 at default | Max context |
|---|---|---|---|---|
| Gemma 4FP16 at default 60.34 · Max context 262144 | 18.32 | 18.48 | 60.34 | 262144 |
| gpt-oss-120bFP16 at default 218.9 · Max context 131072 | 62.35 | 62.49 | 218.9 | 131072 |
| gpt-oss-20bFP16 at default 40.15 · Max context 131072 | 12.05 | 12.15 | 40.15 | 131072 |
| Kimi K2FP16 at default 1913.4 · Max context 131072 | 539.0 | 539.2 | 1913.4 | 131072 |
| DeepSeek R1FP16 at default 1276.5 · Max context 163840 | 359.9 | 360.1 | 1276.5 | 163840 |
| Llama 3.3 70BFP16 at default 134.9 · Max context 131072 | 39.21 | 40.46 | 134.9 | 131072 |
| Llama 3.1 8BFP16 at default 16.96 · Max context 131072 | 5.71 | 6.21 | 16.96 | 131072 |
| Llama 3.1 70BFP16 at default 134.9 · Max context 131072 | 39.21 | 40.46 | 134.9 | 131072 |
| DeepSeek V3FP16 at default 1276.6 · Max context 163840 | 359.9 | 360.1 | 1276.6 | 163840 |
| Llama 3 8BFP16 at default 16.96 · Max context 8192 | 5.71 | 6.21 | 16.96 | 8192 |