Will this model fit on my GPU?

Enter the model size, quantization and context you plan to run. You get the memory it needs and whether it fits on a 16 GB, 32 GB or 128 GB machine — with 10 % of the card kept free, like real runtimes do.

Presets

Result

0.0GB of memory needed

16 GB · laptop / small desktop GPU
32 GB · high-end desktop GPU
128 GB · Mac Studio / server board
Expert details

Formula: weights = params × bits-per-weight ÷ 8 · context = tokens ÷ 1 000 × GB per 1k · overhead = 1.5 GB · a device must keep 10 % of its memory free.