News

How did Plugsky's watchdog auto-heal 4 models after an upstream retirement?

On 27 July 2026, NVIDIA NIM retired the upstream that several Plugsky models depended on. The watchdog detected the failures, found healthy same-profile peers with the same context window and tier, and swapped the upstreams automatically. Affected models, including plugsky-vision-fast, plugsky-qwen-vl, and plugsky-gemma-4, moved to meta/llama-3.2-11b-vision-instruct with no user interruption.

Key facts

EventNVIDIA NIM retired google/gemma-3n-e2b-it on 27 July 2026
ResponseWatchdog swapped 4 affected models to healthy peers
DetectionEvery chat model receives an 80-token probe every 5 minutes
ThresholdThree consecutive failures mark a model broken
Peer selectionHealthy model with the same context window and tier
Runtime cascadeMid-request failures route to a healthy upstream
Uptime target99.9% uptime SLA
Product statusLive

TL;DR

  • An upstream retirement broke several models without any user-visible outage.
  • The watchdog probes every chat model every 5 minutes with a real request.
  • Three consecutive failures trigger a swap to a same-profile healthy peer.
  • Runtime cascades also catch failures that happen mid-request.
  • Users received responses tagged with the fallback model instead of errors.

How it works, step by step

  1. Probe every model with a real request on a fixed interval, not a health ping alone.
  2. Mark a model broken only after repeated consecutive failures.
  3. Select a replacement with the same context window and tier.
  4. Swap the upstream automatically and record why the swap happened.
  5. Cascade mid-request failures to a healthy upstream so callers get an answer.
  6. Force failures in controlled tests to prove the failover path works.
  7. Publish the incident and the test data so customers can verify it.
1Probe every modelwith a real requeston a fixed2Mark a model brokenonly after repeatedconsecutive3Select areplacement withthe same context4Swap the upstreamautomatically andrecord why the swap5Cascade mid-requestfailures to ahealthy upstream so6Force failures incontrolled tests toprove the failover

Original data

NVIDIA NIM retEventWatchdog swappResponseEvery chat modDetection99.9% uptime SUptime targetSource: Plugsky facts table · updated 2026-09-25

Try it yourself

Open the API latency tester →

What happened

On 27 July 2026, several models on the Plugsky API lost their primary upstream when NVIDIA NIM retired the google/gemma-3n-e2b-it model. Users saw no interruption. The watchdog detected the failures, selected healthy same-profile peers with the same context window and tier, and swapped the upstreams automatically.

The affected models — including plugsky-vision-fast, plugsky-qwen-vl, and plugsky-gemma-4 — now run on meta/llama-3.2-11b-vision-instruct.

How the watchdog works

Every five minutes, each chat model receives a real 80-token probe. Three consecutive failures mark a model broken. The watchdog then finds a healthy model with the same context window and tier and swaps the broken model's upstream to that peer — no manual intervention required.

Requests also cascade at runtime: if a model's upstream fails mid-request, the call routes to a healthy upstream and still returns.

Proof it works

In a controlled test, a model was deliberately pointed at a nonexistent upstream. The API still returned HTTP 200 with content in under a second, tagged with the fallback model. A ten-message conversation survived the failover intact, with history preserved across the swap. The latency report holds the full data.

Why this matters

AI outages cost businesses heavily, and single-provider lock-in turns an upstream change into an outage. Automatic peer-healing is what makes an independent AI cloud operationally boring in the good sense: models are replaced by configuration and code paths keep returning answers.

The platform targets a 99.9% uptime SLA under the published terms; see /legal/sla and the status page for current commitments and component health. Boring infrastructure is a feature, and making upstream changes invisible to callers is the goal.

Honest comparison

CapabilityPlugskySingle-provider APISelf-hosted without a watchdog
Failure detectionReal probe every 5 minutes per modelProvider status page onlyWhatever you script
Failover triggerThree consecutive failuresManual interventionManual intervention
Replacement modelSame context window and tierNot applicableYou choose under pressure
Mid-request failureCascades to a healthy upstreamError returnedError returned
User impact during swapNone observed in the July eventOutage until fixedOutage until fixed

Frequently asked questions

What happened on 27 July 2026?

NVIDIA NIM retired the google/gemma-3n-e2b-it upstream, breaking the primary path for several Plugsky models; the watchdog auto-healed them.

Did users see errors?

No. The affected models were swapped to healthy same-profile peers automatically, and users continued to receive responses.

How does the watchdog detect broken models?

Each chat model receives a real 80-token probe every five minutes; three consecutive failures mark it broken.

How is the replacement chosen?

The watchdog prefers a healthy model with the same context window and tier, so speed and capability stay comparable.

What happens if a model fails mid-request?

The request cascades to a healthy upstream and still returns, tagged with the fallback model.

Can we verify the failover?

Yes. A controlled test forced a failure and still returned HTTP 200 with content in under a second, with conversation history preserved.

What uptime commitment applies?

Plugsky publishes a 99.9% uptime SLA under its terms; check /legal/sla and the status page for current details.

Cite this page

Plugsky (2026). “Watchdog Auto-Healed 4 Models”. Plugsky. Available at: https://plugsky.com/news/watchdog-peer-heal-july-2026 (last updated 2026-09-25).