Version 1
100/100 · 46 tokensSummarize the incident report below for the on-call engineering team in 5 bullet points. Include the root cause, the customer impact, and the fix. Return a markdown list.
<report>
{{report}}
</report>Prompt testing
Treat prompts as source files. One command checks every prompt in a repository against a quality bar, a token budget, and a severity threshold, and fails the build when one of them is broken.
A prompt that worked last month can be broken by a well-meant edit: a duplicated paragraph, a removed output format, a pasted key. Prompt testing catches that at review time instead of in production.
The promptflow check command reads the prompt files you point it at and fails when a file scores below your minimum, uses more tokens than your maximum, or has a finding at or above a chosen severity. The default threshold is critical, so a leaked credential always fails. Settings live in a .promptflowrc.json file at the repository root.
These are static tests. They examine the prompt text and never call a model, which is why they are fast, free, and identical on every run. They do not replace output testing against a model; they sit in front of it and remove the failures that do not need a model to find. Model-based testing is on the roadmap and is not available today.
Practical example
Two versions of the same prompt, compared by the engine when this page was built. The edit looked harmless and removed the output format.
Summarize the incident report below for the on-call engineering team in 5 bullet points. Include the root cause, the customer impact, and the fix. Return a markdown list.
<report>
{{report}}
</report>Summarize the incident report below for the on-call engineering team. Include the root cause, the customer impact, the fix, and some other details if relevant.
{{report}}Computed by the engine when this page was built. With the configuration below, version 2 passes the gate at a minimum score of 70.
{
"include": ["prompts/**/*.md"],
"model": "gpt-4.1",
"minScore": 70,
"maxTokens": 2000,
"failOn": "critical"
}promptflow check
promptflow compare prompts/summary.v1.md prompts/summary.v2.mdThe same file always produces the same result, so a red build means the prompt changed.
Nothing is sent to a provider, so there are no secrets to manage and nothing to pay per run.
A repository of prompt files is checked in seconds.
Each failure names the file, the rule, and the fix, and links to the rule reference.
The gate: minimum score, maximum tokens, and fail-on severity across a set of globs.
Exit code 1 on any critical finding, for a simple pass or fail.
Token, quality, and findings difference between two versions of a prompt.
Security-only run that fails on a critical security finding.
Every command accepts --json for use in scripts and dashboards.
Include globs, model, thresholds, and ignored rules in .promptflowrc.json.
Run check in a GitHub Actions job so a prompt change cannot merge below the bar.
Set maxTokens to keep system prompts within the allowance you planned for.
Leave fail-on at critical so a committed key or password fails immediately.
Run validate on staged prompt files for feedback before a push.
No. The checks are static: they analyze the prompt text with deterministic rules. Running prompts against a model and scoring the output is on the roadmap and is not available today.
Not yet. Until the package is published, build the CLI from a checkout of the repository. The CLI documentation has the commands.
Any text file. Point the include globs at the files that hold your prompts, such as Markdown or .prompt.txt files.
Yes. List rule codes under ignore in .promptflowrc.json to skip them for the whole repository.
Every product runs the same engine, so they combine without surprises.
Guides from the prompt engineering knowledge center that explain the ideas this product applies.
Free, deterministic, and private. No sign-up.