Blog

How fast is the Plugsky LLM API in 2026?

On 7 August 2026 Plugsky measured end-to-end latency for every chat model on api.plugsky.com/v1/chat/completions. All measured chat models responded, with fastest-tier medians of 0.23-0.30s, fast-tier medians of 0.46-0.78s and mid-tier medians of 0.67-0.93s. A forced upstream failure still returned content in 0.93s via automatic failover.

Key facts

MethodEnd-to-end, non-streaming /v1/chat/completions calls with one fixed prompt and max_tokens=20
CoverageEvery chat model on the public API returned a valid completion
Published medians0.23-0.30s fastest tier, 0.46-0.78s fast tier, 0.67-0.93s mid tier
Health watchdogAll probed chat models reported healthy at publication, with none failed or rate-limited
Failover behaviourAutomatic fallback kept a request alive through a forced upstream failure
EmbeddingsValid vectors returned for English and Arabic input on the same endpoint
API surfaceOpenAI-compatible /v1/chat/completions and /v1/embeddings
Product statusLive

TL;DR

  • All measured chat models responded; fastest-tier medians landed at 0.23-0.30s.
  • Mid-tier models measured 0.67-0.93s, and the slowest reasoning model measured 3.4s.
  • A forced upstream failure still returned content in 0.93s via automatic fallback.
  • Multilingual embeddings handled Arabic input on the same endpoint as English.
  • Latency varies with load and prompt length; treat these as single-run medians.

How it works, step by step

  1. Run the same fixed prompt against each model you are considering.
  2. Measure end to end, from request to final response, not time to first byte.
  3. Record medians across several runs at different times of day.
  4. Test failover by pointing a client at a bad upstream or simulating a 503.
  5. Check embeddings separately with English and Arabic input if you serve both.
  6. Re-run the benchmark after provider or model changes.
1Run the same fixedprompt against eachmodel you are2Measure end to end,from request tofinal response, not3Record mediansacross several runsat different times4Test failover bypointing a clientat a bad upstream5Check embeddingsseparately withEnglish and Arabic6Re-run thebenchmark afterprovider or model

Original data

End-to-end, noMethod0.23-0.30s fasPublished mediansOpenAI-compatiAPI surfaceSource: Plugsky facts table · updated 2026-09-25

Try it yourself

Open the API latency tester →

Method

On 7 August 2026 Plugsky ran real inference requests against every chat model on the public API at api.plugsky.com/v1/chat/completions using a production key. Each model received the same prompt with max_tokens=20, non-streaming. Latency was measured end to end, from request to final response.

A watchdog probed the same models every five minutes with an 80-token prompt, and embedding models were tested with both English and Arabic input so the results cover more than English chat.

Results by tier

Every chat model returned a valid completion. Fastest-tier models measured 0.23-0.30s median, fast-tier models 0.46-0.78s, and mid-tier models 0.67-0.93s. The slowest model in the run, a long-form reasoning model, measured 12.1s and was still marked healthy.

The practical reading: sub-300ms responses are achievable on hosted open-weight models, and most production chat workloads land between 0.5s and 0.9s. Latency varies with load and prompt length, so treat these as single-run medians rather than guarantees.

Failover and what it means for buyers

Plugsky forced one model's primary upstream to a nonexistent endpoint. The request still returned content in 0.93s, routed automatically to a healthy fallback, and a ten-message conversation recalled earlier context through the failover. Embeddings similarly returned valid vectors, including Arabic input.

For buyers, the lesson is to test failover, not just happy-path latency. Independent AI clouds can be fast, but the differentiator is what happens when an upstream disappears. Re-run your own benchmark with the latency tester and check the status page before you commit.

Re-run the same test whenever a provider changes an upstream, because a single swap can move latency by hundreds of milliseconds.

Honest comparison

TierExample modelsMedian latencyBest for
Fastestplugsky-phi, plugsky-lite, plugsky-gemma3-nano-2b0.23-0.30sHigh-volume chat, classification
Fastplugsky-micro, plugsky-kimi, plugsky-minimax0.46-0.78sGeneral chat with some reasoning
Midplugsky-pro, plugsky-frontier, plugsky-reasoning0.67-0.93sProduction chat, code and agents
Slowerplugsky-deepseek-pro, plugsky-mistral-medium1.4-12.1sHard reasoning, long-form analysis

Frequently asked questions

How was latency measured?

End to end, from request to final response, with one fixed prompt and max_tokens=20, non-streaming, against the public API.

Do these numbers include network time?

The measurement is end-to-end from the client request to the final response, so transport time is included. Your own network path may differ.

Why is one model much slower?

The report records plugsky-mistral-medium at 12.1s and still marks it OK. Long-form and reasoning-heavy models trade latency for depth, so route them away from interactive paths.

Does failover preserve conversation history?

Yes. In the forced-failure test a ten-message conversation correctly recalled earlier context through the automatic failover.

Are embeddings measured too?

Yes. The embedding models returned valid vectors for both English and Arabic input on the same endpoint.

How often are these measurements repeated?

A watchdog probes the models every five minutes, and the report notes that latency varies with load and prompt length, so current figures should be re-checked.

Can I test my own workload?

Yes. Use the API latency tester to run your own prompts and compare model tiers with your prompt shapes and network path.

Cite this page

Plugsky (2026). “LLM API Latency Report 2026: Real Data”. Plugsky. Available at: https://plugsky.com/blog/llm-api-latency-report-2026 (last updated 2026-09-25).