Language models read and write tokens, and you are billed for both. Token optimization means spending them where they improve the answer and nowhere else. Done well, it cuts cost and latency and often improves quality, because a focused prompt is easier to follow.
What a token is
A token is a chunk of text from the model’s vocabulary: a word, part of a word, a punctuation mark, or a space plus a word. For English prose, a rough rule is four characters, or three-quarters of a word, per token. The rule fails often:
- Code, JSON, and unusual identifiers use more tokens than prose.
- Languages other than English often use more tokens per word.
- Each provider has its own tokenizer, so the same text has a different count on each model.
Count with the real tokenizer
Estimates are fine for a first guess and wrong for a budget. OpenAI publishes its tokenizers, so counts can be exact offline. Anthropic and Google provide token-counting endpoints, and without them a count is an approximation.
PromptFlowEngine follows this distinction: counts for OpenAI models are computed with their tokenizer and labeled exact, while other providers are labeled approximate. You can check a prompt in the workspace or over the HTTP API:
curl -s https://promptflowengine.com/api/v1/tokens \
-H 'content-type: application/json' \
-d '{"prompt":"Summarize the notes below in 5 bullets.","model":"gpt-4.1"}'
Where the cost comes from
Cost per call is input tokens times the input price, plus output tokens times the output price. Three facts shape every optimization:
- Output tokens usually cost several times more than input tokens. Limiting answer length often saves more than trimming the prompt.
- The fixed part of a prompt is paid on every call. Fifty wasted tokens in a system prompt that runs a million times is fifty million tokens.
- In a conversation or agent loop, history is re-sent each turn. Cost grows faster than the number of turns.
Prompt compression: what to remove
Work from the safest cut to the riskiest.
Filler and politeness. “Hi”, “please”, “I would like you to”, “it is important to note that”. These carry no instruction.
Before: Hello! I would like you to please summarize the notes below.
Please make sure to keep it brief.
After: Summarize the notes below in 3 sentences.
Repetition. The same instruction stated twice, often in different sections after months of edits.
Wordy constructions. “In order to” becomes “to”. “Due to the fact that” becomes “because”.
Unused context. Background that no test case depends on. Remove it and rerun your tests.
Excess examples. Try three examples instead of eight and measure.
Verbose formats. Indented JSON with long keys costs more than compact JSON with short keys. For data inserted into a prompt, a compact format saves tokens on every call.
The first three are mechanical. The PromptFlowEngine prompt optimizer applies them and verifies that every sentence, requirement, code block, and variable in the original is still present, which is the part that is tedious to check by hand.
What not to remove
- Requirements and constraints. A shorter prompt that drops a rule is a different prompt.
- The output format. Removing it saves ten tokens and costs you a parser.
- Delimiters around input. They are cheap and they prevent injection.
- Structure. Section headings cost a few tokens and can improve adherence. Over-compressed text, with articles and connectives stripped, becomes ambiguous to the model and unreadable to the next person who edits it.
Context window optimization
The context window is the maximum number of tokens a model can consider at once, input and output together. Windows have grown very large, which has changed the problem from “will it fit” to “should it be there”.
More context is not free. Beyond cost and latency, quality can fall as the input grows: a relevant passage surrounded by thousands of irrelevant tokens is harder for the model to use. Send what the task needs.
Retrieve instead of including everything. Search for the passages relevant to the question and insert only those. See RAG prompting.
Order matters in long prompts. Many models attend best to the beginning and end of a long input. Put long documents first and the question or instructions after them, or repeat the key instruction at the end.
Manage conversation history. Summarize older turns, drop tool results that are no longer needed, and start a new conversation when the topic changes.
Reserve room for the answer. If input fills the window, the output is cut off. Leave space for the longest response you expect.
Split the work. A prompt chain gives each step only the context it needs.
Use caching
Major providers offer prompt caching: when the beginning of a prompt matches a recent request, the repeated portion is billed at a reduced rate and processed faster. To benefit:
- Put stable content first: system prompt, tool definitions, examples, reference documents.
- Put variable content last: the user’s question and per-request data.
- Avoid changing anything in the stable part, such as a timestamp, between calls.
Caching makes a long, fixed prompt much cheaper. It does not remove the quality cost of irrelevant context.
Control output length
- Give a number: “under 100 words”, “exactly 5 bullets”.
- Ask for only what you need: a label, not a label with an explanation.
- Set the API’s maximum output tokens as a hard stop.
- For reasoning models, lower the reasoning effort on simple tasks.
Set a budget and enforce it
Decide how many tokens each production prompt may use, and check it automatically. The prompt testing gate fails a build when a prompt file exceeds its token limit:
{ "include": ["prompts/**/*.md"], "model": "gpt-4.1", "maxTokens": 1500 }
That turns token growth from something you discover on an invoice into something you review in a pull request. For the wider process of improving a prompt by measurement, see prompt optimization.