Key facts
| Breach share | Agents involved in more than 1 in 8 reported AI breaches, as reported by HiddenLayer |
| Why it grows | Agents hold credentials and call tools, so one bad instruction can act at machine speed |
| Containment | Ephemeral sandbox per agent run, destroyed afterwards |
| Permissions | Scoped, least-privilege identity per agent and per task |
| Auditability | Replayable audit logs of prompts, tool calls and results |
| Deployment | Plugsky cloud, your VPC, on-prem or air-gapped |
| Product status | Agents live; Playground available in beta |
TL;DR
- Agents now cause more than 1 in 8 reported AI breaches, per HiddenLayer.
- Autonomy expands the blast radius: credentials plus tool access equals exposure.
- Containment beats detection: ephemeral sandboxes and least-privilege scopes.
- Replayable audit logs make every agent action attributable.
- Sovereign deployment keeps agent data and logs in your jurisdiction.
How it works, step by step
- Inventory every agent in production, including ones individual teams adopted without review.
- Move agent runs into ephemeral, isolated sandboxes with no standing credentials.
- Scope each agent's permissions to the minimum tools and data the task needs.
- Require approval gates for irreversible actions such as payments, deletes and sends.
- Turn on replayable audit logs and alert on anomalous tool calls.
- Rehearse containment: assume one agent is compromised and time the response.
- Review agent identities quarterly, revoking stale keys and unused scopes.
Try it yourself
Open the API key security checklist →
Why autonomy expands the blast radius
An agent is not a chatbot with a longer prompt. It holds credentials, calls tools, writes files and chains actions at machine speed. That is what makes it useful — and what lets a compromised prompt turn a helpful worker into an insider. HiddenLayer's 2026 AI Threat Landscape Report, which tied agents to more than one in eight reported AI breaches, describes exactly this shift: the agent itself is the new attack surface.
Contain first: sandboxes and scoped identity
Treat every agent run as untrusted. Start it in an ephemeral sandbox that is destroyed after the task, with no standing access to production secrets. Give the agent its own identity, and scope permissions to the minimum tools and data needed for that run. If a prompt injection or poisoned tool result steers the agent, the blast radius stays inside a disposable box instead of your network.
Audit logs that make accountability real
Accountability is not a policy document; it is a replayable record. Log the goal, the plan, every tool call, the result, and who approved each irreversible action. Store those logs in your own jurisdiction with configurable retention. When an incident review asks what the agent did and who authorized it, you should answer from evidence, not memory. Plugsky agents emit this trail for every run.
A containment checklist for this quarter
Inventory every agent in production, including tools individual teams adopted without review. Move risky agents into sandboxes, cut standing credentials, add approval gates for payments and deletes, and alert on anomalous tool calls. Then rehearse: assume one agent is compromised and time how quickly you contain it. Teams that practice containment recover faster than teams that only detect.
Honest comparison
| Control | Plugsky agent stack | Typical agent framework | Unmanaged agent sprawl |
|---|---|---|---|
| Run isolation | Ephemeral sandbox per run | Depends on your hosting | None |
| Identity | Scoped, least-privilege agent identity | Often shared API keys | Personal accounts |
| Approval gates | Confirmation before irreversible actions | Custom code required | None |
| Audit trail | Replayable logs of prompts, tool calls and results | Partial logging | Unknown |
| Data residency | In-region deployment options | Provider-dependent | Unknown |
| Product status | Agents live; Playground in beta | Not applicable | Not applicable |
Frequently asked questions
What does '1 in 8 AI breaches' actually mean?
It comes from HiddenLayer's 2026 AI Threat Landscape Report, which found autonomous agents involved in more than one in eight reported AI breaches. Treat it as a directional industry signal, not a Plugsky measurement.
Do sandboxes make agents slower?
Isolation adds a little startup overhead per run, but the containment benefit usually outweighs it. Long-running agents can use a persistent sandbox scoped to one task and destroyed afterwards.
Can I keep using my existing agent framework?
Yes. Plugsky exposes an OpenAI-compatible API, so frameworks that speak chat completions or tool calling keep working; you change the base URL and model names.
Are audit logs configurable?
Retention and storage location are configurable, and logs can stay in your chosen jurisdiction. Check the docs for current log fields and retention controls.
Is the agent playground production-ready?
No — the playground is in beta. Use it to evaluate planning, tool use and approval gates before wiring agents into production systems.
Does Plugsky train on my agent data?
No. Prompts and completions stay in your chosen jurisdiction, and training-data isolation is both a contract and architecture guarantee. See the docs and legal terms for details.
How do I start?
Create a free account with no card and two models, plugsky-micro and plugsky-lite, or start a 14-day full-access trial. See the live pricing page for plans.
Plugsky (2026). “Agents Cause 1 in 8 AI Breaches — Contain It”. Plugsky. Available at: https://plugsky.com/news/agentic-ai-breach-risk (last updated 2026-09-25).