16 GB machineFits
14.4 GB usable of 16 GB
Uses 7.4 GB · 8.6 GB free
A quick memory check before you download a local AI model.
e.g. “Llama 3.1 8B” → 8.03
1,000 tokens ≈ 750 words of text.
Depends on the model’s layers and attention heads. Presets fill this in.
Llama 3.1 8B · Q4_K_M · 8.2k tokens
Yes — it runs on all three. Even a 16 GB card has 8.6 GB to spare.
14.4 GB usable of 16 GB
Uses 7.4 GB · 8.6 GB free
28.8 GB usable of 32 GB
Uses 7.4 GB · 24.6 GB free
115.2 GB usable of 128 GB
Uses 7.4 GB · 120.6 GB free
A device “fits” when the model leaves 10 % of its memory free — that headroom keeps everything responsive.
| Bits per weight | 4.85 bpw |
| Model weights | 4.87 GB |
| Context memory | 1.06 GB |
| Overhead | 1.50 GB |
| Total | 7.43 GB |
| Fits on 16 GB (usable 14.4 GB) | Yes · up to 61.8k tokens |
| Fits on 32 GB (usable 28.8 GB) | Yes · up to 172.5k tokens |
| Fits on 128 GB (usable 115.2 GB) | Yes · up to 837.2k tokens |
| Total with Q3_K_M (3.91 bpw) | 6.5 GB |
| Total with Q4_K_M (4.85 bpw) | 7.4 GB |
| Total with Q5_K_M (5.69 bpw) | 8.3 GB |
| Total with Q6_K (6.56 bpw) | 9.1 GB |
| Total with Q8_0 (8.50 bpw) | 11.1 GB |
| Total with FP16 (16.00 bpw) | 18.6 GB |
| Total with NVFP4 (4.50 bpw) | 7.1 GB |
weights = params × bpw ÷ 8 · kv = context ÷ 1000 × kv-per-1k · overhead = 1.5 GB · fits when total ≤ 0.9 × device memory