You're three minutes from a deadline. There's a messy customer complaint in your inbox with a name, a phone number and an order history, and a chatbot would turn it into a polite reply in ten seconds. Do you paste it?
Short answer: not as-is. Paste what the model needs to do the job, and remove what it doesn't. Names, ID numbers, credentials, client documents and anything under an NDA stay out of a personal chatbot account. Everything else is a judgment call that depends on which account you're using and what its settings are. This guide gives you a three-tier sorting rule, a redaction script I ran (with the places it failed), the settings that matter in ChatGPT, Claude and Gemini as of 2026-10-08, and a template for a workplace policy.
For the attack side of this problem, where someone else's text tricks your tool into leaking data, see prompt injection explained. This post is about what you choose to hand over.
What should I never paste into a chatbot?
Sort by what happens if the text ends up somewhere you didn't intend: a breach, a human reviewer, a training set, a subpoena, a screenshot.
| Tier | What it is | Rule |
|---|---|---|
| Red | Passwords, API keys, private keys, one-time codes, card and bank numbers, government ID numbers (Aadhaar, PAN, passport), medical records tied to a named person, unpublished financials, source code you don't own, anything under NDA | Never in a consumer chatbot. Not even "just this once". |
| Amber | Customer or employee names with context, internal strategy docs, contracts, support tickets, your own code at work, meeting transcripts | Only with identifiers removed, or in an approved company tool. |
| Green | Public information, your own writing, generic questions, text you'd be fine seeing quoted on a call | Fine. |
The test I'd use: would I be comfortable if a trained reviewer at the provider read this? Google's own Gemini page says it plainly: "Please don't enter confidential information that you wouldn't want a reviewer to see." That sentence is a decent policy for every tool, because reviewer access, breach and legal process all end up at the same place.
Two cases people underestimate:
- Credentials inside code. You paste a config file to debug it and the database password goes along. Search for keys before you paste, and rotate anything that went.
- Other people's data. Your own details are your risk. A client's, patient's or colleague's details are not yours to give away, whatever the tool's settings say.
How do I anonymise text before pasting it?
The reliable pattern is replace, ask, restore. Swap identifiers for placeholders, ask your question, then map the real values back on your own machine.
I wrote a small redaction script to see how far regexes get you. First attempt, run on an invented sample (every value below is fake):
Hi Priya, please summarise the dispute with Rohan Mehta.
He wrote from [EMAIL_1] and called +91 98765 43210.
PAN [PAN_1], Aadhaar [AADHAAR_1], card [AADHAAR_2] 1111.
Our key is [API_KEY_1]. Invoice total was 4,82,500 rupees.
Three problems in four lines. The phone number slipped through because my pattern didn't allow the space after the first five digits. The card number was chewed up by the Aadhaar rule, which matched its first twelve digits and left the last four. And the script labelled it wrong, so a reader skimming the output would think no card was present. I reordered the patterns (most specific first) and fixed the phone rule. Second run:
Hi Priya, please summarise the dispute with Rohan Mehta.
He wrote from [EMAIL_1] and called [PHONE_1].
PAN [PAN_1], Aadhaar [AADHAAR_1], card [CARD_1].
Our key is [API_KEY_1]. Invoice total was 4,82,500 rupees.
Better, but look at what's still there: "Priya", "Rohan Mehta", and the invoice amount. Regex can catch things with a shape. It can't catch names, company names, or "the Pune warehouse deal that's about to fall through". Those are the identifiers that matter in a dispute. Here's the core of the second version, which I ran with Python 3:
import re
# Order matters: longest/most specific patterns first.
PATTERNS = [
("EMAIL", r"[\w.+-]+@[\w-]+\.[\w.-]+"),
("API_KEY", r"\b(?:sk|pk|ghp|AKIA)[-_A-Za-z0-9]{16,}\b"),
("CARD", r"\b\d{4}[ -]\d{4}[ -]\d{4}[ -]\d{4}\b"),
("AADHAAR", r"\b\d{4}\s\d{4}\s\d{4}\b"),
("PAN", r"\b[A-Z]{5}[0-9]{4}[A-Z]\b"),
("PHONE", r"(?:\+91[\s-]?)?\b[6-9]\d{4}[\s-]?\d{5}\b"),
]
def redact(text):
counts = {}
for label, pat in PATTERNS:
def sub(m, label=label):
counts[label] = counts.get(label, 0) + 1
return f"[{label}_{counts[label]}]"
text = re.sub(pat, sub, text)
return text
Use it as a safety net after you've done the human part, which is read the text and replace names and specifics yourself. Don't treat "the script found nothing" as "nothing sensitive is here". It only knows the formats I taught it, tuned for Indian ID formats and a few key prefixes.
A practical habit: write the placeholders into your prompt. "Customer is [CLIENT_1], order [ORDER_1], amount [AMOUNT_1]" works fine for drafting a reply, and you fill them in locally.
Which settings should I check in ChatGPT, Claude and Gemini?
I checked vendor pages on 2026-10-08. OpenAI's help pages returned HTTP 403 to my fetch tool, so the ChatGPT consumer details below come from search-result summaries of OpenAI's help articles and secondary guides, and I've marked what I couldn't confirm. Settings move, so open the page in your own account before relying on any of this.
Claude
Anthropic's consumer privacy page (it showed "last updated March 16, 2026" when I read it) says:
- The control is called Model Improvement, in Privacy Settings. If it's on, chats and coding sessions from Free, Pro and Max (including Claude Code on those plans) may be used to train models. The page doesn't state the default, so look at your own toggle.
- Incognito chats aren't used for improvement even if Model Improvement is on.
- Deleted conversations leave your history immediately and are deleted from back-end storage within 30 days. With the setting on, de-identified chats may be kept up to 5 years. Thumbs up/down feedback keeps the whole conversation for up to 5 years, so don't click it on a chat that holds anything sensitive.
- Flagged policy violations can be retained for up to 2 years, with safety classification scores up to 7 years.
- Turning the setting off stops future training use. Data already in a training run isn't pulled back out.
- Commercial products (Claude for Work, the API, Claude Gov) are different: Anthropic says it won't use inputs or outputs from them for training by default, except where feedback is submitted or you otherwise allow it. Team and Enterprise owners can turn off member feedback under Organization settings.
Gemini
Google's Gemini Apps Activity page says:
- With Keep Activity on, chats are saved to your Activity and may be used to improve Google services including training generative AI models, and human reviewers may see some of it. Auto-delete defaults to 18 months (3 or 36 months, or none, are options).
- Reviewed chats are disconnected from your account but can be kept up to three years, and deleting your activity doesn't remove chats already reviewed.
- Temporary chats and chats with Keep Activity off are kept 72 hours, and temporary chats aren't used for training.
- Work or school Google accounts may be subject to different data-handling terms, which Google points to its Workspace privacy hub for.
ChatGPT
What I could and couldn't confirm:
- Training toggle. Search summaries of OpenAI's Data Controls help page describe a setting named "Improve the model for everyone" under Settings, Data Controls, which applies account-wide. I could not load the page itself. Turning it off applies going forward only.
- Temporary Chat. Described as not saved to history and not used for training, with a copy possibly retained for up to 30 days for safety. That 30-day figure came from secondary sources and a search summary, not a page I could read. Verify it.
- API. OpenAI's developer docs, which I could read, say data sent to the API isn't used to train models unless you opt in, abuse-monitoring logs are kept up to 30 days by default, and Zero Data Retention needs prior approval and isn't available on every endpoint.
- Business and Enterprise plans. I could not load OpenAI's enterprise privacy page. Read it directly if your company is deciding on a plan.
What the three have in common
Opt-outs control training. They don't promise instant deletion, and they don't stop a provider's staff reviewing flagged content for safety. Every vendor above keeps some copy for some period. If the data would hurt you in a breach, the setting doesn't fix that. The sorting rule does.
Does paying or using a work account change the answer?
Often yes, and that's the main reason companies buy business plans. The general pattern across the vendors I read is that commercial products carry different default terms from consumer apps: Anthropic's page says it doesn't train on commercial inputs by default, and OpenAI's developer docs say the same for the API. Google says work and school accounts may be governed by different terms.
Two catches. "Not used for training" is not "private from everyone": admins can have logs and exports, and your contract, not a blog post, defines what's allowed. And a personal subscription doesn't become a business plan because you used it for work. If your employer hasn't approved the tool, an enterprise-grade privacy policy elsewhere doesn't help you.
What should a workplace AI policy actually say?
If you run a team, a policy that fits on one page gets read. Here's an untested template to adapt (get legal and security to sign off, it isn't legal advice):
AI tool use policy (v1)
1. Approved tools: [list]. Anything else needs approval from [owner].
2. Never enter: credentials, customer personal data, employee HR data,
unreleased financials, regulated data ([list your categories]),
third-party confidential material.
3. Allowed with removal of identifiers: drafts, summaries, analysis of
internal documents using placeholders.
4. Outputs: a human reviews anything sent to a customer or published.
5. Accounts: use your work account on approved tools. Do not use personal
accounts for work content.
6. If you pasted something you shouldn't have: tell [contact] the same day.
Reporting is not punished; hiding it is.
Point 6 matters most. People paste things by mistake. If reporting is safe, you find out about it while it's cheap to fix: rotate a key, ask the vendor about deletion, tell the client.
What if I need AI help with data I can't share?
Three real options, in order of effort:
- Redact and restore. Placeholders, as above. Works for drafting, summarising and reformatting.
- Use the approved enterprise tool. If your company pays for one, that's the point of it.
- Run a model locally. Nothing leaves your machine, at the cost of weaker models and setup time. The OpenClaw with Ollama post shows one way to run a local model.
If none works, do that task by hand. Some documents aren't worth the risk.
A 30-second check before you hit enter
- Is there a password, key, ID number or card number in it? Remove it.
- Is there a real person's name with context? Replace it.
- Would I mind a stranger reading it? If yes, stop.
- Is this a work account approved for this data? If not, stop.
- If I use memory or Projects, does this stay stored somewhere? The ChatGPT Projects and memory guide explains how stored context carries over, and it's a reason to be pickier about what you save there.
For what to do once you have an answer and need to know whether to trust it, read the hallucinations lesson and the companion post on fact-checking AI answers.



