Check if any LLM fits on your GPU. Enter model size, quantization, context length, and VRAM.
The calculator estimates total VRAM needed including model weights, KV cache, and overhead. If the total fits within your GPU's VRAM, the model can run. If not, try a smaller quantization, shorter context, or smaller model.
| GPU | VRAM | Max 7B Q4 | Max 70B Q4 |
|---|---|---|---|
| RTX 4060 | 8 GB | ✅ Yes | ❌ No |
| RTX 4090 | 24 GB | ✅ Yes | ⚠️ 4-bit |
| A100 | 80 GB | ✅ Yes | ✅ Yes |
The GPU Fit Checker estimates whether a model will run on a given graphics card. Enter parameter count, quantization level, available VRAM, and context length, and it calculates the memory needed for weights plus the KV cache, then tells you whether the model fits. It suits anyone planning local inference who wants a quick sanity check before downloading large weights. Results are planning estimates, not benchmarks.
The tool estimates weights as parameter count times bytes per value for the chosen quantization, then adds a KV cache allowance that grows with context length. Real usage also includes runtime overhead, so leave headroom beyond the number shown.
The KV cache stores attention keys and values for every token in context, so memory grows roughly linearly with context length and model size. Long conversations or large documents can push a model past the card's limit even when the weights alone fit.
Lower the quantization from 16-bit to 8-bit or 4-bit, reduce the context length, or choose a smaller model. Quantization trades some quality for roughly halved or quartered weight memory and is the usual first lever.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs