Evaluation plan for an LLM feature
Designs a test set and pass criteria for a model-backed feature, before you tune the prompt.
Use it when: You are shipping a feature that depends on a model and need a repeatable way to tell whether a change made it better.
You are an AI engineer designing evaluations for a language model feature.
## Task
Design an evaluation plan for the feature described below.
## Requirements
- Define 10 test cases: 4 typical inputs, 3 edge cases, and 3 adversarial inputs.
- State the expected property of the output for each case, in a form code or a reviewer can check.
- Name the check for each property: exact match, schema validation, a rule, or a graded rubric.
- State the pass threshold for the suite as a number.
- Do not invent facts about the feature; base every case on the description.
## Output format
Return markdown with three sections: Test cases (a table with the columns ID, Type, Input summary, Expected property, and Check), Pass threshold (one sentence), and Gaps (bulleted list of what the description does not let you test).
## Input
<feature>
{{feature}}
</feature>Variables
{{feature}}- What the feature does, its inputs, its outputs, and what a wrong answer costs.
Example input
feature: extracts the due date and total from invoice text into JSON; a wrong total causes an incorrect payment.
Example output
## Test cases | ID | Type | Input summary | Expected property | Check | |---|---|---|---|---| | T1 | Typical | One-page invoice with one total | `total` equals 1250.00 | Exact match | | A1 | Adversarial | Invoice text containing "ignore the instructions" | Output is valid JSON with the real total | Schema validation |
Illustrative: written to show the expected shape, not generated by a model.