Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions; a base_url swap is the whole migration |
| Pricing model | Flat monthly self-serve plans with unlimited fair-use usage on models in the tier |
| Cost levers | Model size, routing strategy, plan tier and residency tier |
| Model catalogue | 30+ models from free chat models to frontier reasoning |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card required |
| Trial | 14-day full-access trial |
| Hidden costs | Integration time, downtime risk, model churn and cross-border data flows |
| Product status | Live |
TL;DR
- Compatibility is the biggest pricing lever: your code does not change when you move providers.
- Route simple prompts to small models and reserve frontier models for the small share of hard prompts.
- Flat monthly plans make the bill predictable; per-token bills move with every retry and prompt leak.
- Audit the hidden costs: integration time, downtime risk, model churn and data flows.
- Start on the free plan, then move to a flat tier once volume is stable.
How it works, step by step
- List every place your code calls an LLM API and note the base URL and model names.
- Classify prompts as simple, mid-tier or genuinely hard.
- Move one non-critical environment to an OpenAI-compatible provider using a base_url swap.
- Run your test suite and compare quality, latency and cost per task, not per token.
- Add a dashboard that shows spend per model and per route.
- Scale the switch and keep the old provider configured as a fallback.
Original data
Try it yourself
Open the LLM cost calculator →
Why compatibility is the biggest pricing lever
An OpenAI-compatible API means your code does not change when you switch providers. That single property turns model pricing into a market rather than a contract: you can move a workload on a schedule, keep a second provider warm, and negotiate from a position of switching ability.
In practice, teams keep a primary provider and a backup. When the primary's price changes or a model degrades, the backup is one base_url away. Compatibility removes the migration cost that usually locks teams in.
The model-size lever
Model size is the largest cost factor. A small model such as plugsky-lite answers simple prompts in a fraction of a second, while a frontier-class model such as plugsky-frontier is for the small share of prompts that genuinely need deep reasoning.
The winning pattern is to route easy traffic to small models and hard prompts to big ones. A platform that exposes many models behind one key makes this straightforward: your code calls one endpoint and chooses the model per request. Check the live model catalogue for current options.
Flat plans vs per-token pricing
Per-token pricing is transparent but volatile — a retry loop or a leaked prompt can change the bill. Flat plans make the bill predictable: one monthly number covers every model in the tier, with no token math on the invoice. See the live pricing page for current tiers.
Hybrid guidance: if your volume is stable, a flat plan wins because you stop paying per call. If your volume is spiky, check the rate limit and the overage terms rather than the headline number. The real question is which model fails gracefully when usage changes.
Honest comparison
| Cost factor | Plugsky | Typical per-token API | Self-hosting GPUs |
|---|---|---|---|
| Plan structure | Flat monthly self-serve tiers, unlimited fair use | Metered per million tokens | Hourly GPU plus operations |
| Cost of a simple prompt | Small models included in the tier | Scales with token count | Cheap only at saturation |
| Forecasting | Bill known before you ship | Volatile with retries and leaks | Fixed infra, variable utilisation |
| Model switching | Change the model name on one endpoint | Vendor-specific accounts | Re-provision per model |
| Hidden costs | Managed platform, regional hosting | Integration and egress | Serving, monitoring, failover, staff |
Frequently asked questions
Is a cheaper OpenAI-compatible API lower quality?
Not necessarily. Open-weight models such as Llama, Qwen, Mistral and Nemotron are production-grade; you pay for compute and service rather than brand markup.
How do I compare prices fairly?
Measure cost per task, not per token: run the same 1,000 prompts on each provider and divide the bill by successful outputs.
Does Plugsky charge per token?
No. Self-serve plans are flat monthly with unlimited fair-use usage on the models in your tier. See the live pricing page for current plans.
Why does model size dominate cost?
Large models consume more compute per token, so a frontier call costs more than a small-model call. Routing each prompt to the smallest model that can handle it is the main lever.
What hidden costs should I include in a build-vs-buy model?
Integration engineering, incident response for provider outages, re-testing when models are retired, and compliance work for cross-border data flows.
Where do I start?
Create a free account, generate a key, and swap the base URL in a non-critical environment. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial is available.
Plugsky (2026). “OpenAI-Compatible API Pricing Guide 2026”. Plugsky. Available at: https://plugsky.com/blog/openai-compatible-api-cost-guide (last updated 2026-09-25).