Prompt evaluation

Prompt evaluation

A quality score you can check by hand. The score is 100 minus a fixed penalty per issue, every dimension lists the checks behind it, and two versions of a prompt can be compared line by line.

  • Free
  • Published formula
  • No model calls

What it does

Evaluation answers one question: is this version of the prompt better than that one? A number from a model that will not explain itself does not answer it. PromptFlowEngine scores prompts with arithmetic that fits on one line.

Each finding subtracts a fixed penalty by severity. The overall score is 100 minus the sum, and it is capped when any finding is critical. Each of the eight dimensions is scored the same way from its own findings, and the report lists every check that ran in that dimension and whether it passed.

This is an evaluation of the prompt as written, not of model output. It tells you whether the instructions are clear, complete, consistent, and safe. Measuring what a model does with them takes a test set and model calls, which this engine does not make.

Practical example

A scored prompt

The engine evaluated this prompt when the page was built. The table shows each dimension and the checks that failed in it.

The prompt

71/100 · good
Please write a short welcome email for new customers. It should be friendly and professional. Mention a few of our features and stuff. Keep it brief but include plenty of detail.

Computed by the engine when this page was built. 100 minus 29 points of penalties across 5 findings.

Key benefits

Features

Who it is for

Use cases

Frequently asked questions

How is the prompt quality score calculated?

The score starts at 100. Each finding subtracts a fixed penalty by severity: 30 for critical, 10 for warning, and 3 for info. If any finding is critical, the overall score is capped at 50. Each dimension is scored the same way from the findings in that dimension.

Does a high score mean the model will answer well?

It means the prompt has none of the problems the rules look for. Answer quality also depends on the model, the task, and the input, which only output testing can measure.

Is an LLM used as a judge?

No. The built-in evaluator is deterministic and makes no model calls.

What is a good score?

Scores of 85 and above are graded strong. Short prompts for simple tasks often score lower on completeness and still work, so read the findings and not only the number.

Every product runs the same engine, so they combine without surprises.

Learn the technique behind it

Guides from the prompt engineering knowledge center that explain the ideas this product applies.

Try it on your own prompt.

Free, deterministic, and private. No sign-up.