The GGUF Size Calculator estimates how large a quantized GGUF model file will be for a given parameter count and quantization level, plus how much RAM it needs to run. Enter parameters in billions and choose from Q2_K through Q8_0; the page returns estimated file size, RAM with overhead and download time at 100 Mbps. It is aimed at people planning local inference on a laptop or workstation. Sizes are estimates, so check the actual file before downloading.
The page estimates RAM at roughly 1.2 times the file size, which covers runtime overhead. Long context windows and larger batches add more, so leave headroom on a shared machine.
Q4_K is a common balance of size and quality. Q5_K and Q6_K keep more quality at a larger size, while Q2_K and Q3_K shrink the file with a bigger quality trade-off.
No. They are estimates based on parameter count and bit width. Real GGUF files vary by architecture, tokenizer and how each tensor is quantized, so verify the published file size.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs