Your setup
Try a well-known model
the model file will be roughly 4.85 GB
Quantization (bits per weight)
Not sure what to put for context memory? Pick a preset above — each one fills in a realistic value. Longer conversations need more of it.
The verdict
7.4GB
needed in total — model + context + overhead
Fits easily.
16 GB card gaming PC
FITS
32 GB card RTX 5090
FITS
128 GB machine Mac Studio / server
FITS
Expert details
Weights: 8 B × 4.85 bpw ÷ 84.85 GB
Context (KV cache): 8,192 tokens × 0.128 GB/1k1.05 GB
Overhead (runtime, fixed)1.50 GB
Total7.40 GB
A device with D GB counts as fitting while the total stays at or below 0.9 × D — 10 % of the memory is kept free for the system and apps.