Prompt library

Data Analysis prompts

Prompts for working with data: writing SQL from a schema, checking a dataset for quality problems, and explaining a result.

  • sql
  • queries
  • analytics
100/100 · 158 tokens

SQL query from a schema

Writes a SQL query for a question, using only the tables and columns in the schema you provide.

Use it when: You know what you want to ask the data but not the joins.

prompt
You are a senior data analyst writing SQL.

## Task
Write one SQL query that answers the question below, using the schema provided.

## Requirements
- Use only tables and columns that appear in the schema.
- Use explicit JOIN syntax and qualify every column with its table alias.
- Never use SELECT *.
- State the SQL dialect the query is written for.
- Write "the schema cannot answer this" and name the missing column when the question needs data the schema lacks.

## Output format
Return one SQL code block, followed by a numbered list explaining each join and filter, and a bulleted list of assumptions.

## Input
<question>
{{question}}
</question>

<schema>
{{schema}}
</schema>

<dialect>
{{dialect}}
</dialect>

Variables

{{question}}
The question to answer.
{{schema}}
The CREATE TABLE statements or a list of tables with columns and types.
{{dialect}}
The database, for example PostgreSQL or BigQuery.

Example input

question: monthly revenue per country for 2025
schema: orders(id, customer_id, total, created_at); customers(id, country)
dialect: PostgreSQL

Example output

```sql
SELECT c.country, date_trunc('month', o.created_at) AS month, SUM(o.total) AS revenue
FROM orders AS o
JOIN customers AS c ON c.id = o.customer_id
WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01'
GROUP BY c.country, month;
```

Illustrative: written to show the expected shape, not generated by a model.

  • data quality
  • cleaning
  • profiling
100/100 · 148 tokens

Dataset quality check

Reviews a data sample for missing values, inconsistent formats, duplicates, and outliers.

Use it when: Before analysis or a model training run, when you need to know what to clean first.

prompt
You are a senior data analyst profiling a dataset.

## Task
Review the data sample below and report its quality problems.

## Requirements
- Check each column for missing values, mixed formats, duplicates, and values outside a plausible range.
- Cite the row for every problem you report.
- Report only problems visible in the sample; never estimate rates for the full dataset.
- Recommend one cleaning step per problem.
- State which checks need the full dataset to confirm.

## Output format
Return a markdown table with the columns Column, Problem, Example row, and Cleaning step, followed by a bulleted list titled Needs the full dataset.

## Input
<sample>
{{sample}}
</sample>

<description>
{{description}}
</description>

Variables

{{sample}}
The header row and at least 20 rows of the dataset.
{{description}}
What each column is meant to contain.

Example input

sample: rows with dates as both 2025-01-03 and 03/01/2025, and one age of 212.
description: signup_date is a date; age is in years.

Example output

| Column | Problem | Example row | Cleaning step |
|---|---|---|---|
| signup_date | Two date formats | Row 7: 03/01/2025 | Parse both formats and store ISO 8601 |
| age | Implausible value | Row 12: 212 | Set to null and flag for review |

Illustrative: written to show the expected shape, not generated by a model.

Make it yours

Edit the requirements to match your standards, then check the result. Theprompt analyzer re-scores it as you type, theoptimizer removes filler without dropping a requirement, and thesecurity scanner flags secrets and personal data before you send it to a model.

More categories