Optimization, testing, and security

Prompt testing and evaluation: how to measure a prompt

How to test and evaluate LLM prompts. Build a test set, choose metrics, combine static checks, code assertions, and model graders, and run them in CI.

“It looked good when I tried it” is how most prompt regressions reach production. Prompt testing and evaluation replace that with evidence: a fixed set of inputs, a definition of a correct output, and a check that runs every time the prompt changes.

The two terms overlap. Evaluation is measuring how good a prompt or its output is. Testing is running those measurements automatically and failing when a threshold is missed.

Two things you can evaluate

The prompt itself. Is it clear, complete, consistent, and safe as written? This is static analysis. It needs no model call, so it is instant, free, and gives the same result every time.

The model’s output. Given this prompt and this input, is the answer correct? This needs model calls and a way to judge the result.

They catch different problems, and a good process uses both. Static checks run on every keystroke and every commit. Output evaluation runs when a change could affect behavior.

Layer 1: static checks on the prompt

A static check reads the prompt text and looks for known defects:

  • No stated objective, or no output format.
  • Vague words and sizes without numbers.
  • Instructions that contradict or repeat each other.
  • Template variables that are not delimited.
  • Injection phrasing, credentials, or personal data.
  • Token count over a budget.

This is what PromptFlowEngine prompt evaluation does. The score is 100 minus a fixed penalty per finding (30 for critical, 10 for warning, 3 for info), capped at 50 when anything is critical, with the same arithmetic for each of eight dimensions. Because the formula is published, a score change always traces to a specific finding.

A static score does not tell you whether the model’s answer is right. It tells you whether the prompt gives the model a fair chance.

Layer 2: a test set

A test set is a list of inputs with expectations. Start small.

Kind of case Why
Typical inputs Confirms the main path works
Edge cases Empty, very long, wrong language, malformed
Adversarial inputs Attempts to override instructions
Past failures Every bug found becomes a permanent test

Twenty well-chosen cases beat two hundred near-duplicates. Draw them from real traffic where you can, with personal data removed.

Layer 3: deciding whether an output passes

Use the cheapest method that can judge the property.

Code assertions. Deterministic and free. Use them wherever possible.

  • Parses as JSON and matches the schema.
  • Contains, or does not contain, a given string.
  • Length within bounds.
  • Label equals the expected label.
  • Numeric answer within tolerance.

Reference comparison. Compare against a known good answer with exact match or a similarity measure. Suited to extraction and classification, weak for open-ended writing.

Model-graded evaluation. A second model judges the output against a rubric. Use it for qualities code cannot check: faithfulness to a source, tone, completeness.

You are grading an answer against a source document.

Does the answer contain any claim not supported by the source?
Reply with a JSON object: {"unsupported_claims": [string], "pass": boolean}.

<source>{{source}}</source>
<answer>{{answer}}</answer>

Model graders are themselves prompts, with their own failure modes. Keep each rubric to one question, ask for the evidence before the verdict, and check the grader against a sample of human judgments before trusting it.

Human review. The reference standard, and the most expensive. Use it to calibrate the automatic methods and to review a sample of production output.

Metrics worth tracking

Metric What it tells you
Pass rate Share of test cases that meet expectations
Format validity Share of outputs that parse
Faithfulness Share of answers supported by the supplied context
Refusal accuracy Declines what it should, answers what it should
Consistency Agreement across repeated runs of one input
Tokens and latency Cost of the quality you are getting

Report pass rate with the number of cases. “90%” on 10 cases is one failure.

Handling non-determinism

The same prompt and input can produce different outputs. Account for it:

  • Set temperature low for tasks that need consistency.
  • Run each case several times and report the pass rate, not a single result.
  • Assert on properties (“mentions the refund window”) instead of exact text.
  • Pin the model version, and rerun the suite when it changes.

Regression testing in CI

Tests only protect you if they run automatically.

On every commit: static checks. They are fast and need no API key. The prompt testing gate fails a pull request when a prompt file drops below a minimum score, exceeds a token limit, or contains a critical finding such as a leaked key.

- run: promptflow check "prompts/**/*.md" --min-score 70 --max-tokens 1500

On changes to a prompt or model: the output test set, with a pass-rate threshold.

On a schedule: the full set against the live model, to catch provider-side changes.

PromptFlowEngine covers the first of these. It does not run prompts against a model; output evaluation needs a test runner and model access of your own.

Comparing two versions

When choosing between drafts, compare them on the same inputs and look at what changed, not only the totals: which cases flipped from pass to fail, how tokens moved, and whether any new problems were introduced. The compare view in the workspace reports the quality, token, and cost differences between two prompts and lists the findings each version resolves or introduces.

A minimal process to start with

  1. Run the static check and fix warnings.
  2. Write 20 test cases with expectations.
  3. Add code assertions for format and required content.
  4. Record the baseline pass rate.
  5. Gate pull requests on the static check.
  6. Rerun the test set whenever the prompt or the model changes.

For how to use these results to improve a prompt, see prompt optimization. For keeping versions and results organized, see prompt management.