Will It Fit?

Does this AI model fit on your graphics card?

billion params
4.85 bits per weight
tokens
GB per 1,000 tokens

This is the model's KV cache — the memory used to remember the conversation. Typical values: 0.1–0.2 GB per 1k tokens for 7–30B models, around 0.3–0.9 for 70B.

0.0GB
total memory needed

Under the hood

total = weights + context + 1.5 GB overhead  ·  fits when total ≤ 90% of the card