Your setup

Try a well-known model

the model file will be roughly 4.85 GB

Quantization (bits per weight)

Not sure what to put for context memory? Pick a preset above — each one fills in a realistic value. Longer conversations need more of it.

The verdict

7.4GB
needed in total — model + context + overhead

Fits easily.

16 GB card gaming PC FITS
32 GB card RTX 5090 FITS
128 GB machine Mac Studio / server FITS
Expert details
Weights: 8 B × 4.85 bpw ÷ 84.85 GB
Context (KV cache): 8,192 tokens × 0.128 GB/1k1.05 GB
Overhead (runtime, fixed)1.50 GB
Total7.40 GB

A device with D GB counts as fitting while the total stays at or below 0.9 × D — 10 % of the memory is kept free for the system and apps.