Key facts
| Agent behavior | Hundreds of model calls per multi-step task |
| Forecast | Goldman Sachs expects token use to multiply 24x by 2030 |
| Cost control | Flat, plan-based pricing on self-serve |
| Routing | Route routine steps to smaller in-house models |
| Runtime | Sovereign AI agent in the Playground Beta |
| Free plan | 2 free AI models (plugsky-micro, plugsky-lite), no card |
| Trial | 14-day full-access trial |
| Product status | Playground Beta (agent runtime) |
TL;DR
- Agents consume far more tokens than chat because every task is a loop.
- Metered third-party models convert that volume into vendor pricing power.
- Owning the model converts the same volume into predictable infrastructure cost.
- Route routine steps to smaller in-house models to bend the cost curve.
- Cost per completed task, not cost per token, is the metric that matters.
How it works, step by step
- Instrument agents to log tokens and calls per completed task.
- Calculate cost per completed task, not cost per thousand tokens.
- Identify steps that can run on smaller, cheaper in-house models.
- Move orchestration and routine reasoning to your own sovereign endpoint.
- Design bounded loops with iteration limits and budget ceilings.
- Re-measure cost per task after routing changes and track it over time.
- Review agent scope quarterly as horizon tasks grow.
Original data
Try it yourself
Open the agent run cost calculator →
Why agents consume more tokens than chat
A chat turn is usually one request and one response. An agent planning a task, checking intermediate results, and retrying failed steps may issue many calls before it finishes. Each call carries context, so token consumption compounds with the number of iterations.
That is the point of agents — they do more work — but it also means chat-era cost intuition underestimates them badly.
The margin math nobody runs before deploying
If an agent runs on a metered third-party model, every additional step adds a billable call whose price you do not set. Goldman Sachs expects token consumption to multiply 24x by 2030; an agent-heavy product built on metered inference scales its largest cost line with exactly that curve.
The result is a product whose margin erodes as usage grows, with pricing power on the supplier's side.
Metered third-party models vs your own
Running agents on models you control changes the shape of the cost. Capacity is planned rather than metered, routine steps can use smaller models, and usage growth does not automatically trigger a price change from someone else. Flat, plan-based pricing on self-serve Plugsky plans keeps the bill predictable while you scale.
Owning the economics of autonomy
Start by measuring cost per completed task, then route routine steps to smaller in-house models and keep the frontier model for genuinely hard steps. Bound loops with iteration limits and budget ceilings. The Playground Beta gives you a sovereign agent runtime to test this; the free plan includes two free AI models, a 14-day full-access trial covers the catalog, and current plans are on the live pricing page. Measure first, then route.
Honest comparison
| Capability | Plugsky | Metered third-party agent | Building in-house |
|---|---|---|---|
| Cost shape | Flat, plan-based on self-serve | Per-token, vendor-set | Infrastructure plus operations |
| Cost per task at scale | Predictable with routing | Rises with every iteration | Depends on utilization |
| Model routing | Route steps to smaller models | Limited by provider catalog | You build it |
| Data path | Stays in your region | Leaves to the provider | You control |
| Time to first agent | Start in the Playground Beta | Fast, with data leaving | Months |
Frequently asked questions
Why do agents cost more than chat?
Every agent task involves multiple model calls for planning, execution, and verification, so token consumption is a multiple of a single chat exchange.
What does the 24x forecast mean?
Goldman Sachs projects token consumption will multiply 24x by 2030; treat it as a planning signal, not a guarantee, and model your own usage.
How do I control agent cost?
Measure cost per completed task, bound loops, and route routine steps to smaller in-house models instead of sending every step to a frontier model.
Does flat pricing exclude per-token charges?
Self-serve Plugsky plans are flat monthly with fair-use terms rather than per-token billing; see the live pricing page for the current details.
Can agents run on my own infrastructure?
Yes. Plugsky supports Plugsky cloud, VPC, on-prem, and air-gapped deployment, and the Playground Beta agent runs on Plugsky's own models.
How much does Plugsky cost?
The free plan includes two free AI models and a 14-day full-access trial; current plans are on the live pricing page at /#sec-pricing.
What should I track first?
Tokens and calls per completed task, then cost per task by workflow. That baseline shows where routing will have the largest effect.
Plugsky (2026). “Self-Hosted AI Agent Economics”. Plugsky. Available at: https://plugsky.com/news/self-hosted-ai-agent-economics (last updated 2026-09-25).