Skip to main content

What Are AI Tokens and Context Windows? How Context Length, Token Pricing and Limits Work

What does a 1-million-token AI context window actually mean? Learn how tokens are counted, how AI companies charge for them, why context length matters and why bigger is not always better.

By Manuj Gupta
Published: Aug 21, 2026
1 mins read
👁️ 151 Unique Views
What Are AI Tokens and Context Windows? How Context Length, Token Pricing and Limits Work
The scale of inference: Optimized for multimodal workloads.

AI companies increasingly boast about models with context windows of 128,000, 200,000 or even a million tokens. The numbers sound impressive, partly because anything involving a million of something tends to do well in a specification sheet.

But what exactly have you been given a million of?

Understanding tokens and context windows explains not only how AI models “remember” a conversation, but also why API bills rise, why dumping an entire company archive into a prompt may be a terrible idea, and why a larger context window does not automatically mean a smarter AI.

First, what on earth is a token?

Large language models do not read sentences in quite the way humans do. Before text reaches the model, a tokenizer breaks it into smaller pieces called tokens.

A token might be a whole short word, part of a longer word, punctuation or another frequently occurring character sequence. OpenAI gives a useful English-language rule of thumb: 100 tokens is roughly 75 words, although the ratio varies considerably with language, numbers, code and formatting.

So this sentence:

“The robot walked into the kitchen.”

is not necessarily processed as six neat words. The tokenizer converts it into a sequence of token IDs that the neural network actually processes.

This matters because AI companies generally meter API usage in tokens, not words.

Then what is a context window?

Think of the context window as the model's working desk.

Everything it needs for the current task has to fit on that desk: your instructions, previous conversation, documents you supplied, retrieved search results, code, tool information and, depending on the model/API design, room for its response.

A model advertised with a one-million-token context window therefore has an extraordinarily large desk. Current OpenAI GPT-5.6 models, for example, list a 1.05-million-token context window and a maximum output of 128,000 tokens.

But context is not permanent memory. Remove something from the model's supplied context and, unless another memory or retrieval system provides it again, the model cannot simply rummage around in a mysterious mental filing cabinet and retrieve it.

Context window, training knowledge and persistent memory are three different things.

Why do AI companies charge by tokens?

Because processing more tokens requires computation.

API providers therefore commonly quote separate prices for input tokens and output tokens. Input is what you send; output is what the model generates. Output is often considerably more expensive because generating text requires the model to repeatedly compute the next token.

Take OpenAI's GPT-5.6 Sol pricing in August 2026: standard pricing below the long-context threshold is $5 per million input tokens and $30 per million output tokens. Cached input costs less.

Suppose an application sends 100,000 uncached input tokens and receives a 2,000-token answer.

The approximate bill is:

Input: 100,000 ÷ 1,000,000 × $5 = $0.50
Output: 2,000 ÷ 1,000,000 × $30 = $0.06
Total: $0.56

Crucially, having access to a million-token context window does not mean you pay for a million tokens every time. You normally pay for what you actually use.

And very long prompts can cost more. GPT-5.6 Sol currently applies higher rates once input exceeds 272,000 tokens. Anthropic similarly publishes premium pricing for Claude requests crossing specified long-context thresholds.

The giant desk, apparently, comes with giant-desk pricing.

So is a huge context window important?

Sometimes, enormously.

Long context allows an AI system to analyse lengthy codebases, compare dozens of documents, work through legal contracts, retain much longer conversations, examine research papers together or let an agent operate with far more information available at once.

Previously, developers often had to chop documents into pieces, summarise earlier conversations or build retrieval systems that selected only the most relevant material. Larger windows make some of those workflows much easier.

But there is a trap.

Being able to fit information into context is not the same as being able to use it perfectly.

Research famously described a “lost in the middle” effect: models could perform worse when important information was buried in the middle of a long context than when it appeared near the beginning or end. More information could even add distraction instead of improving the answer.

Modern models have improved substantially, but the underlying lesson remains useful: maximum context length is a capacity specification, not an intelligence benchmark.

A well-designed AI system that retrieves 8,000 highly relevant tokens may outperform one indiscriminately swallowing 500,000.

That is why the more interesting question is no longer simply, “How big is the context window?”

It is: How much useful information can the model reliably find, understand and reason over once we fill it?

A bigger desk is excellent. It still helps enormously if somebody knows where the paperwork is.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Tags & Topics

Discussion

Leave a Comment

No comments yet. Be the first to start the conversation!

Link copied to clipboard!