Generate a latency test script that measures p50/p95 for your endpoint.
—
Plugsky is OpenAI-compatible — see docs and start on the free plan (2 free models, no card).
The API Latency Tester is a browser page for measuring LLM API performance: time-to-first-token, total response time, throughput and regional variation. It is aimed at developers and platform teams who need evidence before choosing a provider or region for latency-sensitive work such as chat and voice. The page frames what to measure and why averages hide tail latency. Results depend on your network path, region and payload size, so test from an environment that resembles production.
Time-to-first-token, total latency, throughput in tokens per second and error rate. Track percentiles such as p50 and p95, because averages hide the slow requests that users actually notice.
In streaming chat, users see the first words before the answer is complete, so time-to-first-token drives perceived speed more than total completion time. A slow first token feels broken even if the full answer is fast.
Yes. Physical distance, peering and provider capacity all add latency, which is why the same model can feel fast in one region and slow in another. Test from where your users are.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs