When code reads a model’s answer, “mostly the right shape” is a bug. Structured output prompting is the set of practices that make a model return data your program can parse every time: a defined format, explicit constraints, and validation.
Three levels of reliability
| Level | How | Guarantee |
|---|---|---|
| Prompt only | Describe the format in the prompt | None. Usually right, sometimes not |
| JSON mode | Ask the API for a JSON response | Valid JSON, but not your schema |
| Schema-constrained output | Supply a JSON Schema to the API | Output matches the schema |
Use the strongest level your provider offers. OpenAI, Anthropic, and Google all provide a way to constrain a response to a schema or to a tool’s input schema. The prompt still matters at every level, because a schema controls the shape of the output and the prompt controls its content.
Writing output constraints in the prompt
State the format, the fields, the types, and the allowed values. Then say what must not appear.
Extract the order details from the email in the tags.
Return a JSON object with exactly these fields:
- "order_id": string, as written in the email
- "items": array of {"name": string, "quantity": integer}
- "total": number, without currency symbol
- "currency": one of "USD", "EUR", "GBP"
If a field is not present in the email, use null.
Return only the JSON object, with no explanation and no code fence.
<email>
{{email}}
</email>
Four details in that prompt prevent most failures:
- “Exactly these fields” stops the model adding helpful extras.
- Allowed values for
currencyprevent “US dollars” and “$”. - A null rule gives the model a correct way to say “not found”, so it does not invent a value.
- “Return only the JSON” removes the sentence of commentary that breaks
JSON.parse.
Design the schema for the model
A schema is part of the prompt. The model reads field names and descriptions as instructions.
- Name fields for what they hold.
shipping_addressbeatsaddr2. - Add descriptions to fields whose meaning is not obvious, including the unit and the format: “ISO 8601 date, UTC”.
- Use enums wherever a field has a fixed set of values.
- Keep it flat. Deep nesting increases mistakes. Split a very complex extraction into chained prompts.
- Put reasoning before the answer. If you want an explanation and a verdict, order the fields
"reasoning"then"verdict". A model writes fields in order, so reasoning written first can inform the verdict.
Beyond JSON
Structured output is not only JSON.
- Markdown with fixed headings suits answers that people read: “Use exactly these headings: Summary, Risks, Next steps.”
- XML-style tags are easy to extract with a regular expression and tolerate free text inside:
<answer>…</answer>. - A single token is the most reliable format of all. For classification, ask for only the label.
- CSV or tables work for small, regular data; quote rules get fragile with free text.
Pick the simplest format that carries the information.
Length and content constraints
Output constraints also cover size and style:
- At most 3 bullet points, each under 20 words.
- Use only facts from the document. If the document does not say, write "not stated".
- Write in the present tense.
Give sizes as numbers. “Brief” is not a constraint a parser or a layout can rely on.
Always validate
Even with schema-constrained output, validate before you use the data.
- Parse and schema-check every response. A response can be cut off when it reaches the output token limit, which leaves truncated JSON.
- Check semantics the schema cannot express: the total equals the sum of the items, the date is not in the future.
- Decide the failure path. Retry once with the validation error included in the prompt, then fall back or flag for review.
- Never execute structured output blindly. A field that becomes a SQL clause, a file path, or a shell argument needs the same sanitizing as user input. See prompt security.
Common failure modes
| Failure | Cause | Fix |
|---|---|---|
| Text before or after the JSON | No “only” instruction | Add it, or use JSON mode |
| Wrapped in a code fence | Model habit | Say “no code fence”, or strip it in code |
| Invented values | No rule for missing data | Add a null or “not stated” rule |
| Extra fields | Open-ended wording | Say “exactly these fields”; forbid additional properties |
| Truncated output | Output limit reached | Raise the limit or split the task |
| Inconsistent enum spelling | Values not listed | List the allowed values |
Checking the prompt before you run it
A missing output format is the most frequent finding in prompts that feed code. The PromptFlowEngine prompt analyzer reports it, along with missing length guidance and undelimited variables. The JSON extraction template is a complete example that passes every check.
For how this fits into application code, see prompt engineering for developers.