Key facts
| Routing inputs | Task type, input length, user tier and validation result |
| Latency metrics | Time to first token for streaming plus total completion time |
| Quality gates | Schemas, tests, citations and judge rubrics |
| Cost metric | Cost per successful task, not cost per request |
| Catalogue | 30+ models behind one OpenAI-compatible endpoint |
| Policy | Keep routing rules in configuration and version them |
| Free plan | plugsky-micro and plugsky-lite, no card required |
| Product status | Live |
TL;DR
- Score the request first, then choose the cheapest model that passes.
- Rules cover most routing; add a classifier only when rules cannot decide.
- Latency means time to first token for streaming, not one average.
- Escalation needs an objective trigger and a hard cap.
- Review routing outcomes monthly — untested routers rot.
How it works, step by step
- Define quality gates and latency budgets per workload before choosing models.
- Collect routing signals you already have: task type, input length, tenant tier, validator output.
- Start with explicit rules and a small model ladder for each workload.
- Add a classifier only for the requests your rules cannot separate.
- Shadow-test the router against the current single-model path before switching traffic.
- Track escalation rate, latency percentiles and cost per successful task.
- Re-tune thresholds after model-card changes and review them on a schedule.
Try it yourself
Open the AI workload router simulator →
Signals you can score
Three dimensions decide the route: cost, quality and latency. Each is estimated from signals available before the request runs, then corrected by what happens after.
- Difficulty: task type, input length, number of retrieved passages, structured-output requirement.
- Quality bar: whether a wrong answer is embarrassing, recoverable or costly.
- Latency budget: interactive chat, autocomplete and batch jobs have very different limits.
- Outcome signals: validation result, retry count, user feedback — the data that tunes thresholds.
Implementing the router
Begin with a ladder, not a matrix: define two or three tiers per workload and move up only on failure. Rules such as input length or task type handle most traffic; a small classifier earns its place when easy and hard requests look identical on the surface.
Keep the router stateless and the policy in configuration. Every path needs a fallback — if the chosen model errors or times out, retry on the next tier before surfacing a failure. Fast first hops such as plugsky-lite keep interactive paths responsive, while stronger tiers handle escalations.
Testing and tuning
Run the router in shadow mode against your current path and compare quality, latency and cost on the same traffic. Only then switch real users. Once live, watch the escalation rate: a sudden rise usually means a prompt or validator change, not a model problem.
Report latency as percentiles — p50 hides the p95 that users complain about — and keep time to first token separate from total completion time. Routing is a living policy; schedule a monthly review, and re-run your evaluation set whenever a model card changes. See model routing for the wider implementation patterns.
Honest comparison
| Routing signal | How to detect | Cheap action | Strong action |
|---|---|---|---|
| Short, simple request | Input length and task type | Free or small tier | Escalate only on validation failure |
| Structured output | Schema requirement | Small tier with JSON mode | Stronger tier on schema failure |
| Long input | Token count above threshold | Retrieval instead of full context | Long-context tier |
| Low-confidence answer | Validator or judge score | Retry once on the same tier | Escalate a tier |
| High-stakes request | Workload or tenant flag | Draft on the cheap tier | Top tier plus human review |
Frequently asked questions
How do I route between AI models?
Define quality gates and latency budgets per workload, score each incoming request on difficulty and stakes, then send it to the cheapest model that meets the bar, with a capped escalation path for failures.
What are the best routing signals?
Task type, input length, tenant tier, structured-output requirements and validator results. Use signals available before the call to choose, and outcomes after the call to tune thresholds.
Should I use rules or a classifier?
Start with rules — they are explainable and easy to audit. Add a classifier only for requests where the rules genuinely cannot separate easy from hard.
How should I measure latency?
Report percentiles for time to first token and total completion time separately, at realistic concurrency. A single average hides the slow tail that users notice.
How do I stop escalation loops?
Give escalation an objective trigger, cap attempts per request, and return a safe fallback or human handover after the cap rather than retrying indefinitely.
Can I test a router before using it in production?
Yes — run it in shadow mode against existing traffic, comparing quality, latency and cost, and only move real traffic after the evaluation passes.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage rather than per-token billing, which makes routing policy easier to reason about. See the live pricing page for current plans.
Can I build a router on the free plan?
Yes — plugsky-micro and plugsky-lite are free with no card required and work as fast first hops. Use the 14-day full-access trial for the escalation tiers.
Plugsky (2026). “Route AI Models by Cost, Quality and Latency”. Plugsky. Available at: https://plugsky.com/articles/how-to-route-different-ai-models-by-cost-quality-and-latency (last updated 2026-09-25).