Blog

How do LLM context windows change what you can build in 2026?

A context window is how many tokens a model can see in one call. 128K covers chat, support and most agent loops; 256K to 1M windows let you load whole codebases, contracts or audit logs without chunking. Bigger is not automatically better: cost and latency grow with context, so send only what matters and route long analysis passes to long-context models.

Key facts

128K windowRoughly a 200-page book; covers chat, support, summarisation and most agent loops
256K windowLong documents plus images plus reasoning in one call, as with plugsky-max
1M windowWhole codebase, contract or audit log, as with plugsky-micro and plugsky-longctx
Main trade-offLonger context raises compute cost and latency
Best patternShort model for chat, long-context model for analysis, RAG for retrieval
Agent impactLess aggressive summarisation of conversation history
API surfaceOpenAI-compatible /v1/chat/completions
Product statusLive

TL;DR

  • 128K is enough for chat, support, summarisation and most agent loops.
  • 1M windows remove chunking for whole-codebase and whole-contract analysis.
  • Longer context costs more compute and adds latency, so budget for it.
  • The winning pattern is sending only what matters, not the biggest window available.
  • Keep RAG for retrieval and long-context models for analysis passes.

How it works, step by step

  1. Classify tasks by how much source material they genuinely need.
  2. Estimate tokens for your worst-case prompt, not the average.
  3. Keep chat and support on a 128K-class model.
  4. Route whole-document analysis to a 256K or 1M window model.
  5. Use RAG when the corpus is larger than the prompt budget.
  6. Measure latency and cost after switching to a longer window.
1Classify tasks byhow much sourcematerial they2Estimate tokens foryour worst-caseprompt, not the3Keep chat andsupport on a128K-class model.4Routewhole-documentanalysis to a 256K5Use RAG when thecorpus is largerthan the prompt6Measure latency andcost afterswitching to a

Try it yourself

Open the context window calculator →

The basics

A context window is how many tokens a model can see in one call. 128K tokens is roughly a 200-page book; 1M tokens is roughly a full codebase or a long financial quarter. The window matters less than you think, and more than you would guess, depending on the task.

If the task fits in 128K, a bigger window buys nothing and costs more compute. If the task does not fit, the window decides the architecture: chunking and retrieval, or a single pass.

When size changes the product

A 1M-token window lets you load an entire codebase, contract or audit log without chunking, retrieval or the model forgetting earlier material. A 256K window handles long documents plus images plus reasoning in one call. 128K covers chat, support, summarisation and most agent loops.

Longer context is not free: models with very large windows are slower and more expensive per token than compact models. The winning pattern is not to use the biggest window, but to send only what matters and reserve long-context passes for work that needs them.

Agents and context

Agents fail most often from context loss: the model forgets the original goal after a long tool-calling loop. A large window removes the need for aggressive summarisation of conversation history, which is why long-context models are the quiet backbone of reliable agent systems.

Keep the short model for chat, the long-context model for analysis passes, and RAG for retrieval over large corpora. One OpenAI-compatible endpoint makes that mix practical: you change the model name per request instead of running separate vendors.

The same endpoint also exposes 30+ models, so the long-context pass and the chat model can sit behind one key and one bill.

Honest comparison

Use case128K model256K-1M modelRAG pipeline
Chat and supportFits comfortablyOverkillUnnecessary
Whole-codebase reviewChunking requiredSingle passChunk and retrieve
Long document Q&AMultiple callsSingle callBest when the corpus is large
Cost per callLowestHigher computeEmbedding plus retrieval plus small model
LatencyFastestSlower as context growsRetrieval adds a step
Agent memorySummarise historyKeep full historyVector recall

Frequently asked questions

Is bigger context always better?

No. Cost and latency grow with context length, and irrelevant tokens can distract the model. Match the window to the task and send only the material that matters.

Which Plugsky models have the biggest windows?

plugsky-micro and plugsky-longctx are listed with 1M context, and plugsky-max is positioned for long documents plus images plus reasoning. Check the live model catalogue for the current list.

What does 128K mean in pages?

Roughly a 200-page book. 1M tokens is closer to a full codebase or a long financial quarter of material.

Does long context replace RAG?

Not for large corpora. RAG retrieves only the relevant chunks and keeps the prompt small; long context is better when the whole document genuinely matters.

How does context length affect latency?

Latency generally grows with the number of tokens processed. Long-context passes should be asynchronous or reserved for analysis rather than interactive chat.

How do I stop agents forgetting the goal?

Keep the original goal pinned in the prompt, use a longer window for the loop, and store durable state in agent memory rather than relying on summarisation.

Cite this page

Plugsky (2026). “LLM Context Windows 2026: 128K vs 1M”. Plugsky. Available at: https://plugsky.com/blog/llm-context-window-guide (last updated 2026-09-25).