Key facts
| Method | End-to-end, non-streaming /v1/chat/completions calls with one fixed prompt and max_tokens=20 |
| Coverage | Every chat model on the public API returned a valid completion |
| Published medians | 0.23-0.30s fastest tier, 0.46-0.78s fast tier, 0.67-0.93s mid tier |
| Health watchdog | All probed chat models reported healthy at publication, with none failed or rate-limited |
| Failover behaviour | Automatic fallback kept a request alive through a forced upstream failure |
| Embeddings | Valid vectors returned for English and Arabic input on the same endpoint |
| API surface | OpenAI-compatible /v1/chat/completions and /v1/embeddings |
| Product status | Live |
TL;DR
- All measured chat models responded; fastest-tier medians landed at 0.23-0.30s.
- Mid-tier models measured 0.67-0.93s, and the slowest reasoning model measured 3.4s.
- A forced upstream failure still returned content in 0.93s via automatic fallback.
- Multilingual embeddings handled Arabic input on the same endpoint as English.
- Latency varies with load and prompt length; treat these as single-run medians.
How it works, step by step
- Run the same fixed prompt against each model you are considering.
- Measure end to end, from request to final response, not time to first byte.
- Record medians across several runs at different times of day.
- Test failover by pointing a client at a bad upstream or simulating a 503.
- Check embeddings separately with English and Arabic input if you serve both.
- Re-run the benchmark after provider or model changes.
Original data
Try it yourself
Method
On 7 August 2026 Plugsky ran real inference requests against every chat model on the public API at api.plugsky.com/v1/chat/completions using a production key. Each model received the same prompt with max_tokens=20, non-streaming. Latency was measured end to end, from request to final response.
A watchdog probed the same models every five minutes with an 80-token prompt, and embedding models were tested with both English and Arabic input so the results cover more than English chat.
Results by tier
Every chat model returned a valid completion. Fastest-tier models measured 0.23-0.30s median, fast-tier models 0.46-0.78s, and mid-tier models 0.67-0.93s. The slowest model in the run, a long-form reasoning model, measured 12.1s and was still marked healthy.
The practical reading: sub-300ms responses are achievable on hosted open-weight models, and most production chat workloads land between 0.5s and 0.9s. Latency varies with load and prompt length, so treat these as single-run medians rather than guarantees.
Failover and what it means for buyers
Plugsky forced one model's primary upstream to a nonexistent endpoint. The request still returned content in 0.93s, routed automatically to a healthy fallback, and a ten-message conversation recalled earlier context through the failover. Embeddings similarly returned valid vectors, including Arabic input.
For buyers, the lesson is to test failover, not just happy-path latency. Independent AI clouds can be fast, but the differentiator is what happens when an upstream disappears. Re-run your own benchmark with the latency tester and check the status page before you commit.
Re-run the same test whenever a provider changes an upstream, because a single swap can move latency by hundreds of milliseconds.Honest comparison
| Tier | Example models | Median latency | Best for |
|---|---|---|---|
| Fastest | plugsky-phi, plugsky-lite, plugsky-gemma3-nano-2b | 0.23-0.30s | High-volume chat, classification |
| Fast | plugsky-micro, plugsky-kimi, plugsky-minimax | 0.46-0.78s | General chat with some reasoning |
| Mid | plugsky-pro, plugsky-frontier, plugsky-reasoning | 0.67-0.93s | Production chat, code and agents |
| Slower | plugsky-deepseek-pro, plugsky-mistral-medium | 1.4-12.1s | Hard reasoning, long-form analysis |
Frequently asked questions
How was latency measured?
End to end, from request to final response, with one fixed prompt and max_tokens=20, non-streaming, against the public API.
Do these numbers include network time?
The measurement is end-to-end from the client request to the final response, so transport time is included. Your own network path may differ.
Why is one model much slower?
The report records plugsky-mistral-medium at 12.1s and still marks it OK. Long-form and reasoning-heavy models trade latency for depth, so route them away from interactive paths.
Does failover preserve conversation history?
Yes. In the forced-failure test a ten-message conversation correctly recalled earlier context through the automatic failover.
Are embeddings measured too?
Yes. The embedding models returned valid vectors for both English and Arabic input on the same endpoint.
How often are these measurements repeated?
A watchdog probes the models every five minutes, and the report notes that latency varies with load and prompt length, so current figures should be re-checked.
Can I test my own workload?
Yes. Use the API latency tester to run your own prompts and compare model tiers with your prompt shapes and network path.
Plugsky (2026). “LLM API Latency Report 2026: Real Data”. Plugsky. Available at: https://plugsky.com/blog/llm-api-latency-report-2026 (last updated 2026-09-25).