Gemini, Google’s family of models, is used through the Gemini app, Google AI Studio, and the Gemini API on Google Cloud. The fundamentals of prompting apply unchanged. This guide covers what is particular to Gemini, following Google’s published prompt design guidance.
Model names and limits change, so the advice here concerns stable behavior. Check Google’s documentation for current models.
System instructions
The Gemini API accepts system instructions separately from the conversation. Use them for the role, the standing rules, and the output conventions, and keep the task and its data in the user content. See system prompts and role prompting.
In the Gemini app, saved instructions and custom Gems serve the same purpose: a Gem is a reusable assistant with its own standing instructions.
Include few-shot examples
Google’s guidance is direct on this point: it recommends including few-shot examples in prompts wherever possible, and notes that prompts without examples are likely to be less effective. In practice:
- Two to five examples are usually enough. Too many can cause the model to overfit to them.
- Keep the formatting identical across examples, including delimiters and spacing.
- Show the pattern to follow instead of patterns to avoid.
- Vary the content so only the intended pattern is shared.
Classify the request as question, complaint, or praise.
Text: "How do I export my data?"
Category: question
Text: "The app crashed twice during checkout."
Category: complaint
Text: "{{text}}"
Category:
Ending the prompt with the start of the expected answer, as above, is a completion-style cue that Gemini follows well. More in few-shot prompting.
Use consistent structure
Gemini works with XML-style tags or Markdown headings for separating instructions, context, and input. Choose one style and use it consistently throughout a prompt. Mixing styles makes the boundaries less clear.
Prefixes are another lightweight structure: labels such as Text:, Category:, or JSON: mark what each part is and what should come next.
Long context: question last
Gemini models accept very large inputs, including long documents, codebases, and hours of audio or video. When you use that capacity:
- Put the question or instruction after the material. Google’s guidance for long context is to place the query at the end.
- Do not include material only because it fits. Irrelevant context costs money and can dilute attention. See token optimization.
- Use context caching for large content reused across requests.
Multimodal prompting
Gemini accepts text, images, audio, video, and documents such as PDFs in one prompt. A few practices improve results:
- Say what to look at. “In the second chart” or “between 01:20 and 02:00 in the video”.
- For a single image, put the image before the text that asks about it.
- Ask for a description first when a task needs careful reading of an image, then ask the question. This is a two-step chain.
- Give a rule for the unreadable. “If a value is not legible, use null.”
- State the output format, as you would for text.
The image is a photo of a utility bill. Extract the account number,
billing period, and amount due as JSON. Use null for any field you
cannot read with confidence.
Structured output
For output your code parses, set the response type to JSON and supply a response schema through the API. The response is then constrained to that schema.
- Use enums for fields with fixed values; an enum response type is also available for pure classification.
- Describe fields in the schema, since the model reads it.
- Keep schemas reasonably simple; very large or deeply nested schemas can be rejected or degrade quality.
- Validate the result in your code.
See structured output prompting.
Thinking
Current Gemini models can reason before responding, with a setting that controls how much. For these models, describe the goal and constraints and let the model plan. Lower the setting for simple, high-volume tasks to reduce latency and cost, and raise it for multi-step analysis. Explicit “think step by step” instructions add little when thinking is on. See chain-of-thought prompting.
Grounding with Google Search
The Gemini API can ground a response in Google Search results and return citations. Use it for questions about recent events or facts that change. For answers that must come from your data, supply the documents yourself and require the model to answer from them: see RAG prompting.
Whatever the source, tell the model what to do when the supplied context does not contain the answer.
Function calling
Tools are declared with a name, a description, and a parameter schema. As with every model, the description is what the model uses to decide when to call. Be specific about purpose, parameters, and return values. See prompts for AI agents.
Sampling settings
For tasks that need consistency, such as extraction and classification, the usual advice is a low temperature. Google’s guidance for some recent Gemini models is to leave temperature at its default, because lowering it can cause looping or weaker reasoning. Check the documentation for the model you use before changing it.
Token counting
Gemini provides a token-counting API method, and multimodal inputs are counted in tokens too, with images, audio, and video each having their own rates. There is no offline tokenizer, so PromptFlowEngine labels its Gemini counts approximate. Use the API’s count for budgets.
Prompting in the Gemini app
- Create a Gem for a task you repeat, with its instructions written once.
- Attach files or connect the relevant Google Workspace sources instead of pasting.
- Ask for the format: a table, a list, or a word limit.
- Start a new chat when the subject changes.
- Keep credentials and customer data out of the chat.
The PromptFlowEngine Chrome extension works inside Gemini: it analyzes and restructures the draft in the prompt box and warns when it contains a secret.
Common mistakes with Gemini
| Mistake | What happens | Fix |
|---|---|---|
| No examples | Format and style drift | Add two to five consistent examples |
| Question before a long document | Weaker use of the material | Put the question last |
| Mixed tag and heading styles | Blurred boundaries | Pick one and keep to it |
| Vague multimodal request | Generic description | Say what to look at and what to return |
| JSON requested in prose only | Occasional invalid output | Use a response schema |
| Filling the window because you can | Higher cost, diluted focus | Send what the task needs |
Optimizing a Gemini prompt
- Move standing rules into system instructions.
- State the task, then add consistent examples.
- Place long material first and the question last.
- Define the output with a schema where code reads it.
- Remove filler and repetition, and retest.
The PromptFlowEngine prompt optimizer handles the last step with verification, and the prompt analyzer flags a missing output format before you run anything.
Compare with the ChatGPT and OpenAI prompting guide and the Claude prompting guide.