Prompt optimization is the process of improving a prompt against a defined goal: better answers, fewer tokens, lower latency, or more consistent format. What separates it from tinkering is measurement. You decide what “better” means, change one thing, and check.
Decide what you are optimizing for
A prompt can be optimized along several axes, and they pull against each other.
| Goal | Measured by | Often costs |
|---|---|---|
| Quality | Pass rate on a test set | Tokens |
| Consistency | Same input, same output shape | Flexibility |
| Cost | Input and output tokens per call | Context, examples |
| Latency | Time to first and last token | Reasoning depth |
| Safety | Resistance to injection and leaks | Some helpfulness |
Pick a primary goal and a limit for the others, for example “raise the pass rate without exceeding 800 input tokens”.
The optimization loop
1. Fix the static problems first
Some defects need no model to find: no objective, vague words, sizes without numbers, conflicting instructions, no output format, repeated sentences, undelimited variables. Clear these before spending on model calls. The PromptFlowEngine prompt analyzer lists them, each with a suggested fix.
2. Build a test set
Collect 20 to 50 inputs that represent real use. Include:
- Typical cases.
- Edge cases: empty, very long, mixed language, malformed.
- Adversarial cases: text that tries to change the instructions.
- Past failures. Every bug becomes a test.
For each input, write down what a correct output must contain or satisfy.
3. Establish a baseline
Run the current prompt on every input and record the results. This number is what every change is compared against. Without it you cannot tell an improvement from noise.
4. Read the failures
Group them. Ten failures usually have two or three causes. Look for the pattern: is it always long inputs, always one category, always the same missing field?
5. Form one hypothesis and make one change
“The model adds commentary because the prompt does not say to return only JSON.” Make that change and nothing else.
6. Rerun the whole set
Check that the failures are fixed and, just as important, that nothing that passed before now fails. A fix for one case that breaks two others is a regression.
7. Keep or revert, then repeat
Record what you changed and why. See prompt management for versioning.
Refinement moves that usually work
| Symptom | Refinement |
|---|---|
| Output varies in length | Replace size words with numbers |
| Output varies in shape | Add an explicit format or a schema |
| Ignores a rule | Move it out of a paragraph into a list; remove conflicts |
| Too generic | Add audience and domain context |
| Wrong on edge cases | Add a rule or an example for that case |
| Invents facts | Supply sources and require “not stated” when absent |
| Verbose | Give a word limit and say what to leave out |
Refinement moves that usually do not
- Adding emphasis (capitals, “very important”) to every line.
- Repeating the same instruction in several places.
- Appending a new special-case rule for every failure without checking for conflicts.
- Changing the wording of something that was not the cause.
Optimizing for tokens
Once quality is acceptable, reduce cost without changing behavior:
- Remove greetings, politeness, and filler.
- Delete repeated and near-repeated sentences.
- Cut context that does not change any answer in the test set.
- Reduce the number of examples and retest.
- Put the stable part of the prompt first so provider caching applies.
Steps 1 and 2 are mechanical and safe to automate. The PromptFlowEngine prompt optimizer does them and then verifies that every original sentence, requirement, and protected span is still present, so the shorter prompt asks for exactly what the longer one did. Details in token optimization.
Automated optimization
Tools exist that search for better prompts automatically, by having a model propose variants and scoring them on your test set. They can find improvements people miss. They need three things to work: a good test set, a reliable scoring function, and a review of the winning prompt, since an optimizer will exploit any weakness in the metric.
A related caution applies to asking a chat model to “improve this prompt”. It will produce something that reads better. It may also drop a constraint or add an assumption. Diff the result against the original, and confirm every requirement survived.
When to stop
Stop when the prompt meets the goal on the test set and further changes move results within the noise. Past that point, more gains come from better context, a different technique such as prompt chaining, or a different model.
What a rule-based score can and cannot tell you
A static quality score, such as the one from prompt evaluation, measures the prompt as written: clarity, completeness, consistency, and safety. It is fast and reproducible, so it suits every edit and every commit. It does not measure what a model does with the prompt. Use it to keep the prompt sound, and a test set to confirm the output. Both are covered in prompt testing and evaluation.