Calculate VRAM requirements for any LLM including model weights, KV cache, and overhead.
VRAM = weights + KV cache + overhead
Weights = P ร Q / 8
KV cache = B ร L ร P ร 2 ร Q / 8
Where: P=parameters, Q=bits, B=batch, L=context
The LLM VRAM Calculator estimates the GPU memory an LLM needs for inference. It sums model weights, KV cache and overhead, and reports the total plus how many 80 GB GPUs are required. Inputs are parameter count, quantization, context length and batch size. It is aimed at engineers choosing hardware or planning local deployment. The page states its formula, and results are estimates, so leave headroom.
Weights alone need roughly parameter count times bits per parameter divided by eight. KV cache grows with context and batch, and runtimes add overhead. A 7B model at 4-bit can fit in a few gigabytes, while long-context or high-batch serving needs multiples of that. The calculator shows each part.
4-bit quantization cuts weight memory by roughly 75% against 16-bit and is usually fine for inference, with modest quality loss on some tasks. 8-bit is closer to full precision. Test on your workload, because coding, maths and long-context tasks are often most sensitive.
Yes. A hosted API removes hardware planning, and Plugsky exposes 30+ models through an OpenAI-compatible interface with a free plan including 2 free models. Compare that against GPU purchase, power and maintenance before deciding to self-host.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs