billion params
4.85 bits per weight
tokens
GB per 1,000 tokens
This is the model's KV cache — the memory used to remember the conversation. Typical values: 0.1–0.2 GB per 1k tokens for 7–30B models, around 0.3–0.9 for 70B.
0.0GB
total memory needed
Under the hood
total = weights + context + 1.5 GB overhead · fits when total ≤ 90% of the card