Techniques

Prompt optimization: an iterative method

Prompt optimization is improving a prompt by measurement. A repeatable loop for LLM prompt optimization, including refinement, test sets, and when to stop.

Prompt optimization is the process of improving a prompt against a defined goal: better answers, fewer tokens, lower latency, or more consistent format. What separates it from tinkering is measurement. You decide what “better” means, change one thing, and check.

Decide what you are optimizing for

A prompt can be optimized along several axes, and they pull against each other.

Goal Measured by Often costs
Quality Pass rate on a test set Tokens
Consistency Same input, same output shape Flexibility
Cost Input and output tokens per call Context, examples
Latency Time to first and last token Reasoning depth
Safety Resistance to injection and leaks Some helpfulness

Pick a primary goal and a limit for the others, for example “raise the pass rate without exceeding 800 input tokens”.

The optimization loop

1. Fix the static problems first

Some defects need no model to find: no objective, vague words, sizes without numbers, conflicting instructions, no output format, repeated sentences, undelimited variables. Clear these before spending on model calls. The PromptFlowEngine prompt analyzer lists them, each with a suggested fix.

2. Build a test set

Collect 20 to 50 inputs that represent real use. Include:

  • Typical cases.
  • Edge cases: empty, very long, mixed language, malformed.
  • Adversarial cases: text that tries to change the instructions.
  • Past failures. Every bug becomes a test.

For each input, write down what a correct output must contain or satisfy.

3. Establish a baseline

Run the current prompt on every input and record the results. This number is what every change is compared against. Without it you cannot tell an improvement from noise.

4. Read the failures

Group them. Ten failures usually have two or three causes. Look for the pattern: is it always long inputs, always one category, always the same missing field?

5. Form one hypothesis and make one change

“The model adds commentary because the prompt does not say to return only JSON.” Make that change and nothing else.

6. Rerun the whole set

Check that the failures are fixed and, just as important, that nothing that passed before now fails. A fix for one case that breaks two others is a regression.

7. Keep or revert, then repeat

Record what you changed and why. See prompt management for versioning.

Refinement moves that usually work

Symptom Refinement
Output varies in length Replace size words with numbers
Output varies in shape Add an explicit format or a schema
Ignores a rule Move it out of a paragraph into a list; remove conflicts
Too generic Add audience and domain context
Wrong on edge cases Add a rule or an example for that case
Invents facts Supply sources and require “not stated” when absent
Verbose Give a word limit and say what to leave out

Refinement moves that usually do not

  • Adding emphasis (capitals, “very important”) to every line.
  • Repeating the same instruction in several places.
  • Appending a new special-case rule for every failure without checking for conflicts.
  • Changing the wording of something that was not the cause.

Optimizing for tokens

Once quality is acceptable, reduce cost without changing behavior:

  1. Remove greetings, politeness, and filler.
  2. Delete repeated and near-repeated sentences.
  3. Cut context that does not change any answer in the test set.
  4. Reduce the number of examples and retest.
  5. Put the stable part of the prompt first so provider caching applies.

Steps 1 and 2 are mechanical and safe to automate. The PromptFlowEngine prompt optimizer does them and then verifies that every original sentence, requirement, and protected span is still present, so the shorter prompt asks for exactly what the longer one did. Details in token optimization.

Automated optimization

Tools exist that search for better prompts automatically, by having a model propose variants and scoring them on your test set. They can find improvements people miss. They need three things to work: a good test set, a reliable scoring function, and a review of the winning prompt, since an optimizer will exploit any weakness in the metric.

A related caution applies to asking a chat model to “improve this prompt”. It will produce something that reads better. It may also drop a constraint or add an assumption. Diff the result against the original, and confirm every requirement survived.

When to stop

Stop when the prompt meets the goal on the test set and further changes move results within the noise. Past that point, more gains come from better context, a different technique such as prompt chaining, or a different model.

What a rule-based score can and cannot tell you

A static quality score, such as the one from prompt evaluation, measures the prompt as written: clarity, completeness, consistency, and safety. It is fast and reproducible, so it suits every edit and every commit. It does not measure what a model does with the prompt. Use it to keep the prompt sound, and a test set to confirm the output. Both are covered in prompt testing and evaluation.

Put it into practice

Try it with the prompt optimizer

Cut tokens and add structure, with proof that every requirement survived. It is free, runs deterministically, and needs no account.

Related PromptFlowEngine tools

More in techniques

Browse every prompt engineering guide