Key facts
| Event | NVIDIA NIM retired google/gemma-3n-e2b-it on 27 July 2026 |
| Response | Watchdog swapped 4 affected models to healthy peers |
| Detection | Every chat model receives an 80-token probe every 5 minutes |
| Threshold | Three consecutive failures mark a model broken |
| Peer selection | Healthy model with the same context window and tier |
| Runtime cascade | Mid-request failures route to a healthy upstream |
| Uptime target | 99.9% uptime SLA |
| Product status | Live |
TL;DR
- An upstream retirement broke several models without any user-visible outage.
- The watchdog probes every chat model every 5 minutes with a real request.
- Three consecutive failures trigger a swap to a same-profile healthy peer.
- Runtime cascades also catch failures that happen mid-request.
- Users received responses tagged with the fallback model instead of errors.
How it works, step by step
- Probe every model with a real request on a fixed interval, not a health ping alone.
- Mark a model broken only after repeated consecutive failures.
- Select a replacement with the same context window and tier.
- Swap the upstream automatically and record why the swap happened.
- Cascade mid-request failures to a healthy upstream so callers get an answer.
- Force failures in controlled tests to prove the failover path works.
- Publish the incident and the test data so customers can verify it.
Original data
Try it yourself
What happened
On 27 July 2026, several models on the Plugsky API lost their primary upstream when NVIDIA NIM retired the google/gemma-3n-e2b-it model. Users saw no interruption. The watchdog detected the failures, selected healthy same-profile peers with the same context window and tier, and swapped the upstreams automatically.
The affected models — including plugsky-vision-fast, plugsky-qwen-vl, and plugsky-gemma-4 — now run on meta/llama-3.2-11b-vision-instruct.
How the watchdog works
Every five minutes, each chat model receives a real 80-token probe. Three consecutive failures mark a model broken. The watchdog then finds a healthy model with the same context window and tier and swaps the broken model's upstream to that peer — no manual intervention required.
Requests also cascade at runtime: if a model's upstream fails mid-request, the call routes to a healthy upstream and still returns.
Proof it works
In a controlled test, a model was deliberately pointed at a nonexistent upstream. The API still returned HTTP 200 with content in under a second, tagged with the fallback model. A ten-message conversation survived the failover intact, with history preserved across the swap. The latency report holds the full data.
Why this matters
AI outages cost businesses heavily, and single-provider lock-in turns an upstream change into an outage. Automatic peer-healing is what makes an independent AI cloud operationally boring in the good sense: models are replaced by configuration and code paths keep returning answers.
The platform targets a 99.9% uptime SLA under the published terms; see /legal/sla and the status page for current commitments and component health. Boring infrastructure is a feature, and making upstream changes invisible to callers is the goal.
Honest comparison
| Capability | Plugsky | Single-provider API | Self-hosted without a watchdog |
|---|---|---|---|
| Failure detection | Real probe every 5 minutes per model | Provider status page only | Whatever you script |
| Failover trigger | Three consecutive failures | Manual intervention | Manual intervention |
| Replacement model | Same context window and tier | Not applicable | You choose under pressure |
| Mid-request failure | Cascades to a healthy upstream | Error returned | Error returned |
| User impact during swap | None observed in the July event | Outage until fixed | Outage until fixed |
Frequently asked questions
What happened on 27 July 2026?
NVIDIA NIM retired the google/gemma-3n-e2b-it upstream, breaking the primary path for several Plugsky models; the watchdog auto-healed them.
Did users see errors?
No. The affected models were swapped to healthy same-profile peers automatically, and users continued to receive responses.
How does the watchdog detect broken models?
Each chat model receives a real 80-token probe every five minutes; three consecutive failures mark it broken.
How is the replacement chosen?
The watchdog prefers a healthy model with the same context window and tier, so speed and capability stay comparable.
What happens if a model fails mid-request?
The request cascades to a healthy upstream and still returns, tagged with the fallback model.
Can we verify the failover?
Yes. A controlled test forced a failure and still returned HTTP 200 with content in under a second, with conversation history preserved.
What uptime commitment applies?
Plugsky publishes a 99.9% uptime SLA under its terms; check /legal/sla and the status page for current details.
Plugsky (2026). “Watchdog Auto-Healed 4 Models”. Plugsky. Available at: https://plugsky.com/news/watchdog-peer-heal-july-2026 (last updated 2026-09-25).