Every prompt sent to a hosted model leaves your environment. If it contains an API key, a password, or a customer’s personal data, that information is now in someone else’s systems and possibly in logs on both sides. Sensitive information detection is the practice of finding that content before the prompt is sent.
How sensitive data gets into prompts
Rarely on purpose. The usual routes:
- Pasted logs and stack traces that include tokens, connection strings, or session cookies.
- Configuration files shared for debugging:
.env, YAML, Terraform. - Code with a hard-coded credential.
- Customer records in support tickets, emails, and chat transcripts.
- Retrieved documents that contain more than the question needs.
- Prompt files in a repository where someone tested with a real key.
- Tool results in an agent loop, forwarded to the model unfiltered.
Two categories, two kinds of risk
Secrets grant access: API keys, tokens, passwords, private keys, connection strings. A leaked secret can be used immediately. The response is to revoke and rotate it.
Personal data (PII) identifies people: names, email addresses, phone numbers, postal addresses, payment card numbers, national ID numbers. Sending it may breach your privacy commitments or the law, and it cannot be rotated.
How detection works
Detectors combine several methods, because each has blind spots.
Known formats. Many providers issue keys with a recognizable prefix and length. These can be matched with high confidence.
Context patterns. An assignment such as password = "..." or api_key: ... marks the value as a credential, whatever its format.
Entropy. Random strings have high Shannon entropy. A long, high-entropy value next to a word such as “token” or “secret” is probably a credential, even in an unknown format.
Structured identifiers. Email addresses, phone numbers, and card numbers follow patterns. Card numbers can be confirmed with a checksum.
Named entity recognition. Statistical models find names, organizations, and locations that have no fixed pattern. More coverage, more false positives, and it needs a model.
The PromptFlowEngine prompt security scanner uses the first four. It is deterministic and runs locally, so the text being checked for secrets is not sent anywhere in order to check it.
Report locations, never values
A detector that prints the secret it found has created a second copy of the secret, in a log or a terminal. A well-behaved scanner returns the category and the position only:
{
"hasSensitiveData": true,
"countsByCategory": { "api_key": 1, "email": 1 },
"matches": [
{ "category": "api_key", "start": 118, "end": 164 },
{ "category": "email", "start": 176, "end": 196 }
]
}
This is the shape PromptFlowEngine returns in every interface. See the sensitive data rule for a worked example.
Redaction
Redaction replaces each detected value with a placeholder that keeps the meaning of the surrounding text.
Before: Connect with postgres://admin:hunter2@db.internal:5432/app
After: Connect with [REDACTED_CONNECTION_STRING]
Typed placeholders such as [REDACTED_API_KEY] are better than a row of asterisks, because the model still knows what kind of thing was there and can reason about it (“the connection string looks malformed”) without seeing it.
For personal data that the task needs to refer to, use consistent pseudonyms within a request: every occurrence of one email becomes [EMAIL_1], the next address [EMAIL_2]. The model can then tell people apart, and your application can map the placeholders back after the response.
Where to put the check
Detect as early as possible, and in more than one place.
| Where | What it catches |
|---|---|
| In the editor or chat box | A paste, before it is sent |
| In application code, before the API call | Data from users, records, and retrieval |
| Between agent steps | Tool results about to enter the context |
| In CI, on prompt files | Credentials committed to a repository |
| Before logging | Content about to be stored |
The Chrome extension warns while a draft in ChatGPT, Claude, or Gemini contains a secret, and can redact it on your device. The VS Code extension runs a privacy guard before a prompt is sent to chat, with strict, redact, and warn modes. In CI, the check gate fails on any critical finding by default, so a committed key stops the build.
Policy choices: block, redact, or warn
- Block when a secret should never be sent. Safest, most disruptive.
- Redact as the default. The prompt goes through without the sensitive values.
- Warn when a person should decide, for example a possible false positive in code.
Whichever you choose, apply secret detection locally. A check that sends the text to a remote service to find out whether it is safe to send defeats the purpose.
Limits you should know
- False negatives. A secret in an unusual format, or personal data with no pattern such as a description of a person, can pass. Detection reduces risk; it does not certify a prompt as clean.
- False positives. Hashes, UUIDs, and test fixtures can look like secrets. Allow-list known safe values instead of disabling the check.
- Images and files. Text scanners do not read screenshots or PDFs unless they are converted first.
- Context. A name is public in one setting and confidential in another. Only your policy can decide.
Beyond detection: send less
The most reliable control is not including the data at all.
- Pass an ID and let a tool look up the record, so the model never sees fields it does not need.
- Strip fields before building the prompt.
- Keep credentials in your application and give them to tools directly. A model never needs your API key to ask a tool to call an API.
- Use test data in examples and prompt files.
If a secret has already been sent
- Revoke and rotate it immediately. Treat it as compromised.
- Check access logs for use.
- Remove it from the repository history and prompt files.
- Add a check so the same route is caught next time.
For the surrounding threat model, see prompt security. For attacks that try to extract data once it is in the context, see prompt injection.