Your setup
Total parameters, e.g. 8 for an 8B model.
Fewer bits per weight = smaller model, slightly less quality.
How much text the model can see at once.
Memory the context needs while generating. Typical: 0.05–0.5.
Fits on all three cards
7.4
GB of VRAM needed
About 7.4 GB in total: 4.9 GB for the weights, 1.1 GB for the context and 1.5 GB of overhead. On a 128 GB card (115.2 GB usable), that leaves 107.8 GB free.
Does it fit?
16 GB card
Fits
14.4 GB usable · 6.9 GB to spare
32 GB card
Fits
28.8 GB usable · 21.4 GB to spare
128 GB card
Fits
115.2 GB usable · 107.8 GB to spare
Details
Weights4.85 GB8 B × 4.85 bits/weight
Context (KV cache)1.06 GB8,192 tokens
Overhead1.50 GBruntime, buffers, CUDA
Total7.41 GBfits when ≤ 90 % of VRAM
All quantizations, same context
| Quant | Bits | Total | 16 GB | 32 GB | 128 GB |
|---|---|---|---|---|---|
| Q3_K_M | 3.91 | 6.47 GB | ✓ | ✓ | ✓ |
| Q4_K_M | 4.85 | 7.41 GB | ✓ | ✓ | ✓ |
| Q5_K_M | 5.69 | 8.25 GB | ✓ | ✓ | ✓ |
| Q6_K | 6.56 | 9.12 GB | ✓ | ✓ | ✓ |
| Q8_0 | 8.50 | 11.06 GB | ✓ | ✓ | ✓ |
| FP16 | 16.00 | 18.56 GB | ✗ | ✓ | ✓ |
| NVFP4 | 4.50 | 7.06 GB | ✓ | ✓ | ✓ |