The KV Cache Calculator estimates how much memory key-value cache consumes for a given LLM configuration. Enter the number of layers, hidden dimension, context length and batch size, and it returns KV cache memory. It is for engineers sizing GPUs or tuning serving parameters, because KV cache grows with context and batch and can exceed model weights at long contexts. Treat the output as a planning estimate.
During generation, a transformer stores key and value tensors for every token already processed so it does not recompute the whole sequence. That cache grows linearly with context length and batch size. At long contexts it can dominate memory, limiting how many requests a GPU can serve.
Shorter contexts, smaller batches and grouped-query attention reduce cache size. Quantising the cache and using paged attention also help. If you cannot change the model, cap context with retrieval instead of sending whole documents, and lower concurrency on memory-constrained hardware.
It is an estimate based on the numbers you enter for layers, hidden dimension, context and batch. Real runtimes add allocator overhead and differ with attention layout. Use the figure to compare configurations and set a safety margin rather than to predict exact GPU memory.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs