Can I paste client data into ChatGPT?
The short answer: if the information belongs to a client and you are signed in to a personal ChatGPT account, no. If your organisation is on a business tier with a signed data processing agreement, sometimes, for some categories, and you need to be able to say which. The rest of this is how to tell which case you are in, and what to do when the answer is no but the work still has to get done.
"Does it train on my data" is the wrong first question
It is the question everyone asks, because it is the one with a switch attached.
As of August 2026, OpenAI states that it does not use business data from ChatGPT Enterprise, Business, Edu or the API to train its models, and that consumer accounts can be opted out through data controls. Those statements are real and they matter. They also answer a narrower question than the one you are actually asking.
Turning training off changes what a future model learns. It does not change whether your text was stored, for how long, who inside the vendor can read it during that window, or what happens if a third party later obtains lawful access to the store.
That last item stopped being hypothetical. In the copyright litigation against OpenAI in the Southern District of New York, a preservation order in May 2025 required the company to retain consumer ChatGPT output logs, including conversations users had deleted. The forward-looking part of that obligation was lifted on 26 September 2025, but data already preserved stayed preserved. Then a magistrate judge ordered production of a sample of user logs, and on 5 January 2026 the district judge affirmed it: twenty million de-identified conversations, prompts and model outputs, handed to opposing parties under a protective order.
De-identified, sampled, and covered by court-ordered safeguards. Also: twenty million private conversations that their authors believed were between themselves and a chatbot.
No vendor setting could have prevented that, because it was not the vendor's decision. This is the gap between "we will not train on it" and "nobody will ever read it". A company can promise the first. Nobody can promise the second about text that still exists somewhere.
The four questions that actually decide it
1. Is the data yours to send at all?
If the information was entrusted to you by a client, a patient, or a candidate, you are usually holding it under a duty of confidentiality that exists independently of data protection law. Legal privilege, medical confidentiality, the secrecy obligations that apply to accountants, auditors and tax advisers in most jurisdictions: these are owed to the person, not to a regulator.
The important property of that duty is that disclosure itself is the breach. There is no incident to wait for and no harm threshold to cross. If you sent a client's contract to a third party without a basis for doing so, the exposure exists whether or not anything bad ever happens to the file.
Data protection obligations sit on top of that, not instead of it. We covered where a prompt crosses the GDPR line in a separate post.
2. What does the account promise, in writing?
Not the marketing page. The agreement your organisation actually signed.
A personal account operating under consumer terms gives you no data processing agreement, no defined controller and processor relationship, no contractual retention commitment and no transfer safeguards you can point to. Business tiers change that, and the change is real: training off by default, administrator control over retention, contractual terms a compliance function can review.
The failure mode is not choosing the wrong tier. It is a firm that bought a business tier and has staff still pasting into personal logins on the same laptops, which is the ordinary state of most organisations that have not checked.
3. What happens to the text after you get your answer?
The default for API traffic is retention for up to thirty days for abuse monitoring, after which it is deleted unless a legal obligation requires otherwise. Zero data retention exists, applies only to eligible endpoints, and is granted on approval for qualifying enterprise agreements. It is not a checkbox on a normal account.
So for most people, most of the time, the honest description is: the text sits with a third party for a period you did not choose, in a jurisdiction that may not be yours, and "delete" removes it from your view before it removes it from theirs.
4. Could you show, a year from now, what was sent?
Accountability under Article 5(2) is a demonstration requirement. If your position is that staff exercised judgement, the follow-up question is which data reached which tool on which date. Chat interfaces produce no record you control. "We were careful" is a claim, not evidence.
The ten second version
| What you want to paste | Personal account | Business tier with DPA | After masking |
|---|---|---|---|
| Client contract naming real parties | No | Only with a lawful basis and a record of it | Yes |
| CV or candidate notes | No | Only with a lawful basis and a record of it | Yes |
| Patient notes, health details | No | Rarely, and not without a DPIA | Yes |
| Customer support email thread | No | Usually, with a record | Yes |
| Source file carrying credentials or customer rows | No | No | Yes, once the identifiers and secrets are out |
| Internal draft with no people in it | Yes | Yes | Not needed |
The right-hand column is the one worth noticing. It does not depend on the tier, the vendor, the jurisdiction or the year, because there is nothing in the text for any of those to apply to.
Four beliefs that do not survive contact
"Training is off, so it is private." Training is one use of your data among several. Storage, abuse monitoring, human review of flagged content and legal process are the others, and the training switch touches none of them.
"I deleted the chat." Deletion starts a removal process on the vendor's side, subject to their stated window. A preservation obligation overrides it, as the 2025 order demonstrated for conversations users had already deleted.
"I took the name out, so it is anonymous." A job title plus an employer plus a city identifies one person. So does a birth date plus a postcode. Removing the direct identifier while leaving the context is the most common way people convince themselves a prompt is safe. We wrote about what actually counts as personal data in a prompt.
"IT approved ChatGPT." Approving a tool is not the same as approving a data class for it. Most approvals were granted for general productivity use and were never scoped to client files, and staff reasonably read the approval as covering whatever they type.
"My client consented." Consent to your engagement is not consent to transfer their file to a third-party processor abroad. If consent is your basis, it has to be consent to that, and you have to be able to produce it.
The version that is always allowed
Every question above exists because identifiable personal data leaves your control. Remove that premise and the questions do not need better answers, they stop applying.
That is what masking before send does. Names, contact details, identifiers, account numbers and the sensitive details that slipped in for context get replaced before the text reaches the model. The model works on the structure of the problem, which is what made the answer good in the first place. The identities stay on your machine, and the reply comes back with the real values restored so the output is usable rather than full of gaps.
The word doing the work is reversible. One-way redaction produces a draft somebody has to reassemble by hand, and people who have to do that twice go back to pasting the raw version. We compared the two masking styles, opaque tokens and format-valid stand-ins, here.
Velum does this locally. The browser extension makes no network calls at all. The desktop app makes five enumerated outbound requests, listed inside the app itself, and none of them carry your document text. If you want the version that becomes a written rule for staff, the AI usage policy template is free and has a data classification table you can adopt as it stands.
Common questions
Can lawyers use ChatGPT with client information?
Not on a personal account. On a business tier with a data processing agreement, the professional conduct rules in your jurisdiction still govern disclosure of client confidences to a third party, and those rules sit outside data protection law entirely. The American Bar Association's Formal Opinion 512, issued 29 July 2024, reads Model Rule 1.6 as requiring a client's informed consent before information relating to the representation goes into a self-learning generative AI tool, and states that boilerplate consent buried in an engagement letter does not do the job. Bar and law society guidance elsewhere has landed in similar territory. Masking the client's identifying details before the prompt leaves is the route that does not depend on any of it.
Is ChatGPT Enterprise GDPR compliant?
A product cannot be compliant on your behalf. The tier gives you the contractual pieces compliance requires: a processing agreement, training off, retention controls, transfer terms. Your lawful basis, your minimisation, your records and your DPIA where one is needed remain yours to establish.
Does turning off chat history stop the data being stored?
No. It changes what appears in your sidebar and what feeds model improvement. Retention for safety and abuse purposes runs on its own schedule, and legal preservation obligations override both.
Is it enough to remove names before pasting?
Rarely. Direct identifiers are the easiest part to catch and the least likely to be the only thing identifying someone. Combinations of ordinary details do the identifying, which is why manual redaction under time pressure has a poor record.
Can I paste client data if the client agreed?
Only if what they agreed to was this. Consent has to be specific, informed and demonstrable, which means consent to the transfer of their information to a named third party for this purpose, not a general clause about using technology in your practice.
What if I only need help with the shape of the problem?
That is the usual case, and it is the whole argument. If the model does not need to know who the person is in order to do the task, the person has no reason to be in the prompt.
If you want to see where this lands across legal, recruitment, healthcare, finance and support work, see the use cases.