AI agents and MCP

Prompts for AI agents: instructions, tools, and workflows

How to write prompts for AI agents. System instructions, tool descriptions, context management, stopping rules, and when a fixed AI workflow beats an agent.

An AI agent is a language model running in a loop with tools: it decides what to do, calls a tool, reads the result, and continues until the task is done. Prompting an agent differs from prompting for a single answer. You are no longer writing a question. You are writing the operating instructions for something that will make many decisions without you.

Workflow or agent

First decide whether you need an agent at all.

AI workflow AI agent
Who decides the steps Your code The model
Predictability High Lower
Cost and latency Lower, known in advance Higher, variable
Good for Tasks with a known procedure Tasks where the path depends on what is found
Debugging Step by step Read the whole trace

If you can draw the flowchart, build a workflow with prompt chaining. Use an agent when the number and order of steps cannot be known ahead of time, such as investigating a bug or researching a question.

The parts of an agent prompt

An agent’s behavior comes from four things it reads: the system prompt, the tool definitions, the task, and the results that come back. Each needs deliberate writing.

1. The system prompt

Cover six things.

Role and goal. What the agent is for, in two sentences.

Operating rules. How to work: investigate before changing anything, prefer small steps, verify results.

Boundaries. What it must not do, and what needs approval.

How to handle problems. Tool errors, missing information, ambiguity.

When to stop. The condition that means the task is done, and a limit on effort.

How to report. What the final answer contains.

You are a coding agent working in the user's repository.

## How to work
- Read the relevant files before proposing a change.
- Make the smallest change that satisfies the task.
- After editing, run the tests that cover the change.

## Boundaries
- Do not modify files outside the repository.
- Do not delete files or run destructive commands without asking.
- Do not change public interfaces unless the task says so.

## When things go wrong
- If a command fails, read the error and try a different approach once.
  If it fails again, stop and report what you tried.
- If the task is ambiguous, ask one question before starting.

## When to stop
- Stop when the acceptance criteria are met and the tests pass.

## Final report
- What changed and why, files touched, how it was verified, and
  anything left undone.

Explain the reason behind a rule when it is not obvious. An agent meets situations you did not list, and a rule with a reason generalizes better than a bare prohibition.

2. Tool definitions

The model chooses tools by reading their names and descriptions. Vague definitions are the most common cause of agents using the wrong tool.

  • Name tools for what they do: search_orders, not query2.
  • Describe when to use each one, and when not to.
  • Say what comes back, including the format and the limits.
  • Make parameters unambiguous: types, formats, allowed values, an example.
  • Avoid overlapping tools. If two tools could do the job, the model will pick inconsistently.
  • Return actionable errors: “date must be YYYY-MM-DD”, not “invalid input”.

Fewer, well-described tools outperform a large set of similar ones. The Model Context Protocol standardizes how tools are described and offered to agents.

3. The task prompt

Give the agent a specification, not a sentence. The more autonomy it has, the more the starting brief matters.

## Objective
Add rate limiting to the public API.

## Requirements
- 100 requests per minute per API key.
- Return 429 with a Retry-After header when exceeded.

## Constraints
- Existing clients must keep working: no changes to response schemas.

## Acceptance criteria
- New tests cover the limit and the 429 response.
- The existing test suite passes.

Acceptance criteria give the agent a way to know it is finished and give you a way to check. The PromptFlowEngine VS Code extension turns a one-line task into this structure and attaches the relevant files.

4. Context management

Everything the agent reads stays in its context and is re-read at every step. Left alone, the context fills with stale tool output, cost climbs, and quality drops.

  • Return compact tool results. The fields needed, not the whole payload.
  • Summarize or drop old results once they have been used.
  • Keep durable notes outside the context, in a file the agent can read and update, for long tasks.
  • Delegate. Hand a self-contained sub-task to a sub-agent with a fresh context and take back only its conclusion.

See token optimization.

Writing prompts for sub-agents

When one agent briefs another, the brief is all the recipient knows. A good delegation prompt states:

  1. The goal and why it matters.
  2. What is already known, so work is not repeated.
  3. The exact output wanted, and its format.
  4. The boundaries: what to touch and what to leave alone.

An orchestrating agent that writes vague briefs gets vague work back. This is a place where an automatic check helps: an agent can pass its draft brief through the PromptFlowEngine MCP tools to find a missing objective or output format before delegating.

Safety for agents

Agents act, so mistakes and attacks have consequences.

  • Least privilege. Each tool has the narrowest permission that works. Read-only by default.
  • Approval for irreversible actions. Deleting, sending, paying, deploying.
  • Treat tool results as untrusted. A web page or file can contain instructions aimed at the agent. See prompt injection.
  • Limits. Maximum steps, maximum spend, and timeouts.
  • Keep secrets out of the context. Give credentials to tools directly. See PII and secret detection.
  • Log every tool call so a run can be reviewed.

Common failure modes

Symptom Likely cause Fix
Loops on the same action No stopping rule or step limit Add both
Declares success too early No acceptance criteria Define them; require verification
Uses the wrong tool Overlapping or vague descriptions Rewrite tool descriptions
Does far more than asked Scope not stated State boundaries explicitly
Forgets earlier instructions Context overloaded Trim results, summarize, use notes
Follows text found in a file Untrusted content not marked Mark it as data; restrict tools

Evaluating agents

Score the outcome, not the path: did the tests pass, was the right record updated, is the answer correct. Then review traces for wasted steps and risky actions. Keep a set of tasks with known good outcomes and rerun it when the prompt, the tools, or the model changes. See prompt testing and evaluation.

For the reasoning loop underneath all of this, see ReAct prompting.