Key facts
| Tier | Top capability tier in the Plugsky model ladder |
| Best for | Hard reasoning, high-stakes drafting and multi-document synthesis |
| Use as | Escalation target and quality baseline, not a default |
| Capabilities | Streaming, function calling, JSON mode and long-context |
| Context class | 128K-class window; live limits are published per model |
| Pricing tier | Paid-plan model; free plan covers plugsky-micro and plugsky-lite |
| Resilience | Backup upstream plus same-profile fallback peers |
| Product status | Live |
TL;DR
- Frontier is the quality ceiling — use it where a wrong answer is expensive.
- Keep it off the default path; route to it on validation failure or complexity.
- Use it to build the evaluation baseline other models are measured against.
- Send complete, well-structured context so the call is not wasted.
- Score escalations regularly so the tier stays where it earns its keep.
How it works, step by step
- Define what "hard" means for your product: failure cost, ambiguity, input length.
- Build an escalation eval set from your hardest real requests.
- Run plugsky-frontier against your current production model on that set.
- Add a validator so escalation fires on evidence, not guesswork.
- Cap escalation attempts per request and log the escalation rate.
- Keep a faster model as the default path and cache repeated premium outputs.
- Review escalation triggers monthly and after model-card changes.
Try it yourself
Open the LLM cost calculator →
What the frontier tier is for
plugsky-frontier is the top capability tier in the catalogue, intended for requests where answer quality matters more than response time. Typical jobs are complex analysis, high-stakes drafting, difficult reasoning chains and synthesis across several documents. It speaks the same OpenAI-compatible chat surface as every other model, including streaming, function calling and JSON mode.
The frontier model guide covers the model itself; this article is about deciding when a request deserves it. Read the live model card at /models before hard-coding window or capability assumptions.
Escalation triggers that work
- Validation failure: a schema, test, citation or rubric check fails on the cheaper tier.
- Complexity flags: long inputs, multiple documents or reasoning-heavy task types.
- Premium workflows: contract review, financial analysis or anything customer-facing that must be right.
- Offline jobs: batch analysis where throughput does not matter but quality does.
Triggers should be observable and testable. "The first answer looked weak" is not a trigger; "the verifier returned false" is. Keep the escalation evidence compact and send the original request unchanged so context is preserved through the chain.
Cost and quality discipline
Frontier calls are the ones worth measuring most carefully. Build a baseline: what does your current model score on your hardest set, and what does frontier add? If the delta is small, the escalation rule — not the model — is the problem.
Cache outputs for repeated premium tasks, store them with the source inputs, and review the escalation rate monthly. Self-serve plans are flat monthly with fair-use usage rather than per-token billing, so use the live pricing page to map escalation volume to plan limits, and protect quality by keeping the frontier path well fed rather than overused.
Honest comparison
| Dimension | Frontier tier | Workhorse tier | Small tier |
|---|---|---|---|
| Quality ceiling | Highest in the catalogue | Strong general purpose | Good for narrow, simple tasks |
| Latency profile | Slowest | Balanced | Fastest |
| Best use | Escalation and baseline | Default production traffic | Cheap first hop |
| Cost profile | Highest per call | Moderate | Lowest |
| Failure impact | Lowest risk | Manageable with validators | Needs escalation path |
| Failover | Same-profile top-tier peers | Peer fallback | Peer fallback |
Frequently asked questions
When should I use plugsky-frontier?
Use it for hard reasoning, high-stakes drafting, evaluation baselines and multi-document synthesis — and as an escalation target when a cheaper model fails validation. Keep it off the default path.
How is plugsky-frontier different from plugsky-max?
Both sit at the top of the catalogue. plugsky-frontier is the frontier-profile flagship, while plugsky-max is the largest general tier. Compare them on your hardest prompts before committing a workload to either.
Should I use frontier for everything?
No. It is the slowest and most expensive tier. Routine chat, classification, extraction and agent steps belong on faster models with a defined escalation path.
How do I measure whether frontier is worth it?
Run it against your production model on your hardest real requests, then compare quality, latency and escalation rate. If the quality gain does not change outcomes, tighten the trigger.
Does it support tools and structured output?
Yes — streaming, function calling and JSON mode are part of the OpenAI-compatible surface, so existing tool-using code works without changes.
Can I use it on the free plan?
No. The free plan covers plugsky-micro and plugsky-lite with no card required. A 14-day full-access trial lets you evaluate frontier and other paid tiers.
What happens if the upstream has an incident?
Requests retry through a backup upstream and same-profile top-tier fallback peers. Live component health is published on the status page.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage rather than per-token billing. See the live pricing page for current plans.
Plugsky (2026). “When to Use a Frontier-Tier Model”. Plugsky. Available at: https://plugsky.com/articles/plugsky-frontier-when-to-use-a-frontier-tier-model (last updated 2026-09-25).