News

Why do self-hosted AI agents change the economics of autonomy?

Agents do not send one prompt; they send hundreds of calls in a loop for a single task. Goldman Sachs expects token consumption to multiply 24x by 2030, so a metered third-party model turns that curve into a cost problem — and pricing power for the vendor. Running agents on models you control, with routine steps routed to smaller in-house models, keeps margins predictable.

Key facts

Agent behaviorHundreds of model calls per multi-step task
ForecastGoldman Sachs expects token use to multiply 24x by 2030
Cost controlFlat, plan-based pricing on self-serve
RoutingRoute routine steps to smaller in-house models
RuntimeSovereign AI agent in the Playground Beta
Free plan2 free AI models (plugsky-micro, plugsky-lite), no card
Trial14-day full-access trial
Product statusPlayground Beta (agent runtime)

TL;DR

  • Agents consume far more tokens than chat because every task is a loop.
  • Metered third-party models convert that volume into vendor pricing power.
  • Owning the model converts the same volume into predictable infrastructure cost.
  • Route routine steps to smaller in-house models to bend the cost curve.
  • Cost per completed task, not cost per token, is the metric that matters.

How it works, step by step

  1. Instrument agents to log tokens and calls per completed task.
  2. Calculate cost per completed task, not cost per thousand tokens.
  3. Identify steps that can run on smaller, cheaper in-house models.
  4. Move orchestration and routine reasoning to your own sovereign endpoint.
  5. Design bounded loops with iteration limits and budget ceilings.
  6. Re-measure cost per task after routing changes and track it over time.
  7. Review agent scope quarterly as horizon tasks grow.
1Instrument agentsto log tokens andcalls per completed2Calculate cost percompleted task, notcost per thousand3Identify steps thatcan run on smaller,cheaper in-house4Move orchestrationand routinereasoning to your5Design boundedloops withiteration limits6Re-measure cost pertask after routingchanges and track

Original data

Goldman Sachs Forecast2 free AI modeFree plan14-day full-acTrialSource: Plugsky facts table · updated 2026-09-25

Try it yourself

Open the agent run cost calculator →

Why agents consume more tokens than chat

A chat turn is usually one request and one response. An agent planning a task, checking intermediate results, and retrying failed steps may issue many calls before it finishes. Each call carries context, so token consumption compounds with the number of iterations.

That is the point of agents — they do more work — but it also means chat-era cost intuition underestimates them badly.

The margin math nobody runs before deploying

If an agent runs on a metered third-party model, every additional step adds a billable call whose price you do not set. Goldman Sachs expects token consumption to multiply 24x by 2030; an agent-heavy product built on metered inference scales its largest cost line with exactly that curve.

The result is a product whose margin erodes as usage grows, with pricing power on the supplier's side.

Metered third-party models vs your own

Running agents on models you control changes the shape of the cost. Capacity is planned rather than metered, routine steps can use smaller models, and usage growth does not automatically trigger a price change from someone else. Flat, plan-based pricing on self-serve Plugsky plans keeps the bill predictable while you scale.

Owning the economics of autonomy

Start by measuring cost per completed task, then route routine steps to smaller in-house models and keep the frontier model for genuinely hard steps. Bound loops with iteration limits and budget ceilings. The Playground Beta gives you a sovereign agent runtime to test this; the free plan includes two free AI models, a 14-day full-access trial covers the catalog, and current plans are on the live pricing page. Measure first, then route.

Honest comparison

CapabilityPlugskyMetered third-party agentBuilding in-house
Cost shapeFlat, plan-based on self-servePer-token, vendor-setInfrastructure plus operations
Cost per task at scalePredictable with routingRises with every iterationDepends on utilization
Model routingRoute steps to smaller modelsLimited by provider catalogYou build it
Data pathStays in your regionLeaves to the providerYou control
Time to first agentStart in the Playground BetaFast, with data leavingMonths

Frequently asked questions

Why do agents cost more than chat?

Every agent task involves multiple model calls for planning, execution, and verification, so token consumption is a multiple of a single chat exchange.

What does the 24x forecast mean?

Goldman Sachs projects token consumption will multiply 24x by 2030; treat it as a planning signal, not a guarantee, and model your own usage.

How do I control agent cost?

Measure cost per completed task, bound loops, and route routine steps to smaller in-house models instead of sending every step to a frontier model.

Does flat pricing exclude per-token charges?

Self-serve Plugsky plans are flat monthly with fair-use terms rather than per-token billing; see the live pricing page for the current details.

Can agents run on my own infrastructure?

Yes. Plugsky supports Plugsky cloud, VPC, on-prem, and air-gapped deployment, and the Playground Beta agent runs on Plugsky's own models.

How much does Plugsky cost?

The free plan includes two free AI models and a 14-day full-access trial; current plans are on the live pricing page at /#sec-pricing.

What should I track first?

Tokens and calls per completed task, then cost per task by workflow. That baseline shows where routing will have the largest effect.

Cite this page

Plugsky (2026). “Self-Hosted AI Agent Economics”. Plugsky. Available at: https://plugsky.com/news/self-hosted-ai-agent-economics (last updated 2026-09-25).