Prompt testing

Prompt testing

Treat prompts as source files. One command checks every prompt in a repository against a quality bar, a token budget, and a severity threshold, and fails the build when one of them is broken.

  • CLI
  • Offline
  • Build from source

What it does

A prompt that worked last month can be broken by a well-meant edit: a duplicated paragraph, a removed output format, a pasted key. Prompt testing catches that at review time instead of in production.

The promptflow check command reads the prompt files you point it at and fails when a file scores below your minimum, uses more tokens than your maximum, or has a finding at or above a chosen severity. The default threshold is critical, so a leaked credential always fails. Settings live in a .promptflowrc.json file at the repository root.

These are static tests. They examine the prompt text and never call a model, which is why they are fast, free, and identical on every run. They do not replace output testing against a model; they sit in front of it and remove the failures that do not need a model to find. Model-based testing is on the roadmap and is not available today.

Practical example

A regression caught by compare

Two versions of the same prompt, compared by the engine when this page was built. The edit looked harmless and removed the output format.

Version 1

100/100 · 46 tokens
Summarize the incident report below for the on-call engineering team in 5 bullet points. Include the root cause, the customer impact, and the fix. Return a markdown list.

<report>
{{report}}
</report>

Version 2

84/100 · 36 tokens
Summarize the incident report below for the on-call engineering team. Include the root cause, the customer impact, the fix, and some other details if relevant.

{{report}}
Quality
-16
Tokens
-10
New issues
3
  • warningTemplate variable not delimited: {{report}}.
  • infoNo output format is specified.
  • infoNo length limit for generated content.

Computed by the engine when this page was built. With the configuration below, version 2 passes the gate at a minimum score of 70.

.promptflowrc.json
{
  "include": ["prompts/**/*.md"],
  "model": "gpt-4.1",
  "minScore": 70,
  "maxTokens": 2000,
  "failOn": "critical"
}
shell
promptflow check
promptflow compare prompts/summary.v1.md prompts/summary.v2.md

Key benefits

Features

Who it is for

Use cases

Frequently asked questions

Does prompt testing run my prompts against a model?

No. The checks are static: they analyze the prompt text with deterministic rules. Running prompts against a model and scoring the output is on the roadmap and is not available today.

Is the CLI on npm?

Not yet. Until the package is published, build the CLI from a checkout of the repository. The CLI documentation has the commands.

What file types can it check?

Any text file. Point the include globs at the files that hold your prompts, such as Markdown or .prompt.txt files.

Can I ignore a rule?

Yes. List rule codes under ignore in .promptflowrc.json to skip them for the whole repository.

Every product runs the same engine, so they combine without surprises.

Learn the technique behind it

Guides from the prompt engineering knowledge center that explain the ideas this product applies.

Try it on your own prompt.

Free, deterministic, and private. No sign-up.