Will It Fit?

See how much VRAM your model needs.

Your setup

Total parameters, e.g. 8 for an 8B model.

Fewer bits per weight = smaller model, slightly less quality.

How much text the model can see at once.

Memory the context needs while generating. Typical: 0.05–0.5.

Fits on all three cards
7.4 GB of VRAM needed

About 7.4 GB in total: 4.9 GB for the weights, 1.1 GB for the context and 1.5 GB of overhead. On a 128 GB card (115.2 GB usable), that leaves 107.8 GB free.

Does it fit?

16 GB card Fits
14.4 GB usable · 6.9 GB to spare
32 GB card Fits
28.8 GB usable · 21.4 GB to spare
128 GB card Fits
115.2 GB usable · 107.8 GB to spare

Details

Weights4.85 GB8 B × 4.85 bits/weight
Context (KV cache)1.06 GB8,192 tokens
Overhead1.50 GBruntime, buffers, CUDA
Total7.41 GBfits when ≤ 90 % of VRAM

All quantizations, same context

QuantBitsTotal16 GB32 GB128 GB
Q3_K_M3.916.47 GB✓✓✓
Q4_K_M4.857.41 GB✓✓✓
Q5_K_M5.698.25 GB✓✓✓
Q6_K6.569.12 GB✓✓✓
Q8_08.5011.06 GB✓✓✓
FP1616.0018.56 GB✗✓✓
NVFP44.507.06 GB✓✓✓