Pick a quantization level from your VRAM and model size.
—
Plugsky is OpenAI-compatible with flat-rate plans — see pricing or start on the free plan (2 free models, no card).
The Model Quantization Selector compares quantization formats such as Q4, Q8, GPTQ, AWQ and FP8 to help match a model to your GPU. It is for people running local inference who need to balance file size, memory use and output quality, and who want a plain explanation of which format suits which hardware. Use it to narrow the choice, then measure throughput and quality on your own setup. Quantization trade-offs vary by model, runtime and task.
There is no universal winner. Fit the model into memory first, then test quality. Q4_K is a common balance, while higher bit widths preserve more quality at a larger size.
GGUF K-quants suit llama.cpp, Ollama and similar runtimes that mix CPU and GPU. GPTQ and AWQ target GPU inference servers such as vLLM, where they can improve throughput.
Usually yes, and lower bit widths lose more. The drop varies by model and task, so measure it on your own evaluation set rather than assuming a fixed penalty.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs