Methodology
Sources in this file were checked 2026-10-10. A model page's Last updated date is the Hub lastModified date of its data, not the build date.
Every total is an estimate unless the row is labeled as published files. Published-file rows still add an estimated cache and the fixed overhead. Nothing on this site is a measurement of allocator-reserved VRAM.
Units
1 GB in these tables means 1 GiB, 1024 cubed bytes. GPU listings that say 24 GB are treated as 24 of those units. Manufacturer pages are not always explicit about GiB versus decimal GB; the cited figure is used as printed.
Weights
The parameter count is safetensors.total from the official Hub API, including vision or audio tensors when those configs exist. Weights for a formula row are that count times the nominal bits of one llama.cpp block, divided by 8, rounded half up with integer arithmetic.
The nominal bits come from ggml-common.h at commit f2918cabbffe8abf8e2a90c1c089085fa116c2cf. Q4_K_M, Q5_K_M, and Q3_K_M recipes keep some tensors at a higher type, so a real GGUF file is larger than the base block applied to every parameter. This site does not print a GGUF file size unless the bytes were measured.
| Quantization | Bits | Byte ratio |
|---|---|---|
| FP16/BF16 | 16 | 2/1 |
| Q8_0 | 8.5 | 17/16 |
| Q6_K | 6.5625 | 105/128 |
| Q5_K_M | 5.5 | 11/16 |
| Q4_K_M | 4.5 | 9/16 |
| Q3_K_M | 3.4375 | 55/128 |
| Q2_K | 2.625 | 21/64 |
Cache
Ordinary attention stores 2 vectors per layer: 2 × layers × key heads × head dimension × context × 2 bytes. Head dimension is the config field when it exists, otherwise hidden size divided by query heads.
MLA models (a kv_lora_rank and a qk_rope_head_dim) store the compressed latent: (kv_lora_rank + qk_rope_head_dim) × 2 bytes × layers × context. A reference implementation may expand that latent into full key and value tensors. This estimate does not bill the expansion.
Hybrid configs with a layer_types list charge sliding layers for min(context, sliding window) and full layers for the full context. Gemma 3 has no layer_types list; a sliding_window_pattern of N marks a layer sliding unless its index plus 1 is divisible by N.
Gemma 4 full layers use global_head_dim. When attention_k_eq_v is set, those layers store 1 vector, and they use num_global_key_value_heads when that field is set. num_kv_shared_layers drops the cache on later layers. Those Gemma 4 rules are not applied to other families.
Llama 4 no_rope_layers uses chunked attention (the flag value 1) or full attention (the flag value 0). The chunk length is attention_chunk_size. Phi-3 with a sliding window below its context max charges every layer as sliding. A Qwen config with use_sliding_window set to false ignores a sliding_window number that is still present.
Context tables show 2048, 8192, 32768, and 131072 when those values are within the config max, and they add the config max when it is at most 262144 and not already listed. Longer maxima are stated in prose. The quant table uses min(8192, config max).
Overhead, fit, and the GPU pick
Overhead is 1 GiB on every row, a fixed allowance for allocator waste and runtime buffers. It is not fitted from a benchmark.
Fits means the total is at or under usable memory and at least 10 percent of that memory remains. Tight fit means the total fits with less than 10 percent free. CPU offload means usable memory is under the total but at least 6 GB and at least a quarter of the weight bytes. Won't fit means neither. The words are the signal. Color only repeats the word.
The recommended card is the smallest usable memory in the 20-GPU list whose Q4_K_M total at the table context is at most 90 percent of that usable memory. A discrete card's usable memory is the listed figure. A unified pool's usable memory is 75 percent of the listed figure. Ties follow list order. There is no price in the rule. If the configured maximum is under 8,192 tokens, the rule uses that maximum.
Published files
A published-file row appears only when the Hub tree API's safetensors byte sum is under 75 percent of parameters times 2. That test keeps a BF16 dump on the FP16 formula row. The row's weight bytes are the file sum. Bits per parameter are that sum times 8, divided by the parameter count. Cache and overhead are still estimates.
gpt-oss quant_method mxfp4 is labeled from the config and the model card. DeepSeek fp8 is labeled from quant_method fp8. Kimi K2.6 compressed-tensors packing is labeled as published checkpoint files, not as a GGUF type.
Gated configs
Some official config.json files return 401 without a login. Those pages use a public mirror for architecture fields and the official API for the parameter total, license tag, gated flag, and dates. The mirror's quantization block is ignored when the official tensors are BF16.
Llama 3.1 405B is the case that needs a numeric override. The AWQ mirror lists 16 key heads. An untied Llama reconstruction matches the official safetensors.total only at 8 key heads. The page uses 8.
- meta-llama/Llama-3.3-70B-Instruct config fields from unsloth/Llama-3.3-70B-Instruct
- meta-llama/Llama-3.1-8B-Instruct config fields from unsloth/Llama-3.1-8B-Instruct
- meta-llama/Llama-3.1-70B-Instruct config fields from NousResearch/Meta-Llama-3.1-70B-Instruct
- meta-llama/Meta-Llama-3-8B-Instruct config fields from NousResearch/Meta-Llama-3-8B-Instruct
- meta-llama/Meta-Llama-3-70B-Instruct config fields from NousResearch/Meta-Llama-3-70B-Instruct
- meta-llama/Llama-3.1-405B-Instruct config fields from hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4; key heads set to 8 to match the official parameter total
- meta-llama/Llama-4-Scout-17B-16E-Instruct config fields from unsloth/Llama-4-Scout-17B-16E-Instruct
- meta-llama/Llama-4-Maverick-17B-128E-Instruct config fields from unsloth/Llama-4-Maverick-17B-128E-Instruct
- google/gemma-3-27b-it config fields from unsloth/gemma-3-27b-it
- google/gemma-3-12b-it config fields from unsloth/gemma-3-12b-it
GPU memory
Each row is one configuration named in the manufacturer excerpt stored under data/sources/gpu/. RTX 3080 uses the 10 GB figure in the 3080 column. RTX 4060 Ti uses the 16 GB configuration. RTX 3060 uses the 12 GB configuration. RTX 4070 uses the 12 GB column whose type is listed as GDDR6 / GDDR6X. H100 SXM uses the 80 GB column, which the page does not label with a memory type.
Apple's excerpt lists 48GB unified memory for M5 Pro and M5 Max with a 40-core GPU. The table's Apple row is that 48 GB pool. The memory column shows 48 GB. Fit, headroom, and the smallest-GPU pick use 75 percent of a unified pool, so this row is treated as 36 GB usable. The spec page does not print an operating-system reserve. 75 percent is this site's rule for every unified row.
| GPU | Memory | Type | Source |
|---|---|---|---|
| GeForce RTX 3060 12GBType GDDR6 | 12 GB | GDDR6 | Manufacturer page |
| GeForce RTX 3070 8GBType GDDR6 | 8 GB | GDDR6 | Manufacturer page |
| GeForce RTX 3080 10GBType GDDR6X | 10 GB | GDDR6X | Manufacturer page |
| GeForce RTX 3090 24GBType GDDR6X | 24 GB | GDDR6X | Manufacturer page |
| GeForce RTX 4060 8GBType GDDR6 | 8 GB | GDDR6 | Manufacturer page |
| GeForce RTX 4060 Ti 16GBType GDDR6 | 16 GB | GDDR6 | Manufacturer page |
| GeForce RTX 4070 12GBType GDDR6 / GDDR6X | 12 GB | GDDR6 / GDDR6X | Manufacturer page |
| GeForce RTX 4070 Ti Super 16GBType GDDR6X | 16 GB | GDDR6X | Manufacturer page |
| GeForce RTX 4080 16GBType GDDR6X | 16 GB | GDDR6X | Manufacturer page |
| GeForce RTX 4090 24GBType GDDR6X | 24 GB | GDDR6X | Manufacturer page |
| GeForce RTX 5060 Ti 16GBType GDDR7 | 16 GB | GDDR7 | Manufacturer page |
| GeForce RTX 5070 12GBType GDDR7 | 12 GB | GDDR7 | Manufacturer page |
| GeForce RTX 5070 Ti 16GBType GDDR7 | 16 GB | GDDR7 | Manufacturer page |
| GeForce RTX 5080 16GBType GDDR7 | 16 GB | GDDR7 | Manufacturer page |
| GeForce RTX 5090 32GBType GDDR7 | 32 GB | GDDR7 | Manufacturer page |
| Radeon RX 7900 XTX 24GBType GDDR6 | 24 GB | GDDR6 | Manufacturer page |
| Radeon RX 9070 XT 16GBType GDDR6 | 16 GB | GDDR6 | Manufacturer page |
| NVIDIA A100 80GBType HBM2e | 80 GB | HBM2e | Manufacturer page |
| NVIDIA H100 SXM 80GBType Not stated | 80 GB | Not stated | Manufacturer page |
| Apple M5 Max 48GB unifiedType unified | 48 GB | unified | Manufacturer page |