Calculate maximum throughput within API rate limits.
The Rate Limit Calculator converts API rate limits into practical throughput numbers. Enter your requests-per-minute cap, tokens-per-minute cap, and average tokens per request, and it shows which limit binds first, the real maximum requests per minute, the interval between requests, and the daily ceiling. It is for developers sizing clients, queues, and retry logic before hitting 429 errors. Use it with your provider's published limits.
Providers typically return HTTP 429 and may throttle or reject requests until the window resets. Well-behaved clients retry with exponential backoff, honor any Retry-After header, and cap concurrency so they stay under both RPM and TPM.
It depends on request size. With short requests, RPM binds first; with long prompts or large outputs, TPM binds first. The calculator shows both, so you can see the effective maximum rather than assuming the lower headline number.
Reduce tokens per request by trimming context or summarizing, spread load evenly instead of bursting, request a higher limit from the provider, or route simple requests to a smaller model on a separate limit.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs