Will it fit on my graphics card?

A quick memory check before you download a local AI model.

7.4 GB needed ✓ 16✓ 32✓ 128

Your setup

Popular models
B

e.g. “Llama 3.1 8B” → 8.03

1,000 tokens ≈ 750 words of text.

GB/1k

Depends on the model’s layers and attention heads. Presets fill this in.

Memory needed

7.4 GB

Llama 3.1 8B · Q4_K_M · 8.2k tokens

Yes — it runs on all three. Even a 16 GB card has 8.6 GB to spare.

Model weights · 8.03 B × 4.85 ÷ 84.87 GB
Context (KV cache) · 8.2k tokens1.06 GB
Runtime overhead · fixed allowance1.50 GB

Does it fit?

16 GB machineFits

14.4 GB usable of 16 GB

Uses 7.4 GB · 8.6 GB free

32 GB machineFits

28.8 GB usable of 32 GB

Uses 7.4 GB · 24.6 GB free

128 GB machineFits

115.2 GB usable of 128 GB

Uses 7.4 GB · 120.6 GB free

A device “fits” when the model leaves 10 % of its memory free — that headroom keeps everything responsive.

Expert details
Memory breakdown
Bits per weight4.85 bpw
Model weights4.87 GB
Context memory1.06 GB
Overhead1.50 GB
Total7.43 GB
Fits on 16 GB (usable 14.4 GB)Yes · up to 61.8k tokens
Fits on 32 GB (usable 28.8 GB)Yes · up to 172.5k tokens
Fits on 128 GB (usable 115.2 GB)Yes · up to 837.2k tokens
Total with Q3_K_M (3.91 bpw)6.5 GB
Total with Q4_K_M (4.85 bpw)7.4 GB
Total with Q5_K_M (5.69 bpw)8.3 GB
Total with Q6_K (6.56 bpw)9.1 GB
Total with Q8_0 (8.50 bpw)11.1 GB
Total with FP16 (16.00 bpw)18.6 GB
Total with NVFP4 (4.50 bpw)7.1 GB

weights = params × bpw ÷ 8 · kv = context ÷ 1000 × kv-per-1k · overhead = 1.5 GB · fits when total ≤ 0.9 × device memory