Key facts
| 128K window | Roughly a 200-page book; covers chat, support, summarisation and most agent loops |
| 256K window | Long documents plus images plus reasoning in one call, as with plugsky-max |
| 1M window | Whole codebase, contract or audit log, as with plugsky-micro and plugsky-longctx |
| Main trade-off | Longer context raises compute cost and latency |
| Best pattern | Short model for chat, long-context model for analysis, RAG for retrieval |
| Agent impact | Less aggressive summarisation of conversation history |
| API surface | OpenAI-compatible /v1/chat/completions |
| Product status | Live |
TL;DR
- 128K is enough for chat, support, summarisation and most agent loops.
- 1M windows remove chunking for whole-codebase and whole-contract analysis.
- Longer context costs more compute and adds latency, so budget for it.
- The winning pattern is sending only what matters, not the biggest window available.
- Keep RAG for retrieval and long-context models for analysis passes.
How it works, step by step
- Classify tasks by how much source material they genuinely need.
- Estimate tokens for your worst-case prompt, not the average.
- Keep chat and support on a 128K-class model.
- Route whole-document analysis to a 256K or 1M window model.
- Use RAG when the corpus is larger than the prompt budget.
- Measure latency and cost after switching to a longer window.
Try it yourself
Open the context window calculator →
The basics
A context window is how many tokens a model can see in one call. 128K tokens is roughly a 200-page book; 1M tokens is roughly a full codebase or a long financial quarter. The window matters less than you think, and more than you would guess, depending on the task.
If the task fits in 128K, a bigger window buys nothing and costs more compute. If the task does not fit, the window decides the architecture: chunking and retrieval, or a single pass.
When size changes the product
A 1M-token window lets you load an entire codebase, contract or audit log without chunking, retrieval or the model forgetting earlier material. A 256K window handles long documents plus images plus reasoning in one call. 128K covers chat, support, summarisation and most agent loops.
Longer context is not free: models with very large windows are slower and more expensive per token than compact models. The winning pattern is not to use the biggest window, but to send only what matters and reserve long-context passes for work that needs them.
Agents and context
Agents fail most often from context loss: the model forgets the original goal after a long tool-calling loop. A large window removes the need for aggressive summarisation of conversation history, which is why long-context models are the quiet backbone of reliable agent systems.
Keep the short model for chat, the long-context model for analysis passes, and RAG for retrieval over large corpora. One OpenAI-compatible endpoint makes that mix practical: you change the model name per request instead of running separate vendors.
The same endpoint also exposes 30+ models, so the long-context pass and the chat model can sit behind one key and one bill.Honest comparison
| Use case | 128K model | 256K-1M model | RAG pipeline |
|---|---|---|---|
| Chat and support | Fits comfortably | Overkill | Unnecessary |
| Whole-codebase review | Chunking required | Single pass | Chunk and retrieve |
| Long document Q&A | Multiple calls | Single call | Best when the corpus is large |
| Cost per call | Lowest | Higher compute | Embedding plus retrieval plus small model |
| Latency | Fastest | Slower as context grows | Retrieval adds a step |
| Agent memory | Summarise history | Keep full history | Vector recall |
Frequently asked questions
Is bigger context always better?
No. Cost and latency grow with context length, and irrelevant tokens can distract the model. Match the window to the task and send only the material that matters.
Which Plugsky models have the biggest windows?
plugsky-micro and plugsky-longctx are listed with 1M context, and plugsky-max is positioned for long documents plus images plus reasoning. Check the live model catalogue for the current list.
What does 128K mean in pages?
Roughly a 200-page book. 1M tokens is closer to a full codebase or a long financial quarter of material.
Does long context replace RAG?
Not for large corpora. RAG retrieves only the relevant chunks and keeps the prompt small; long context is better when the whole document genuinely matters.
How does context length affect latency?
Latency generally grows with the number of tokens processed. Long-context passes should be asynchronous or reserved for analysis rather than interactive chat.
How do I stop agents forgetting the goal?
Keep the original goal pinned in the prompt, use a longer window for the loop, and store durable state in agent memory rather than relying on summarisation.
Plugsky (2026). “LLM Context Windows 2026: 128K vs 1M”. Plugsky. Available at: https://plugsky.com/blog/llm-context-window-guide (last updated 2026-09-25).