Confidentiality8 min read

How to anonymise a document before using AI (step by step for lawyers)

Swapping out names isn't enough. Here's the process that actually holds up, including the step almost every guide skips: metadata and fact-pattern identification.

Minimal illustration of a document with redacted lines and a red redline mark, representing anonymisation before AI use.
By The Redline Editors

Deleting a client's name before you paste text into ChatGPT is not anonymisation. It's the first step of one, and most of the guidance circulating right now stops exactly there. The real process has four parts: strip structural identifiers, rewrite the fact pattern so it can't be traced back to one matter, check the document itself for hidden metadata, and only then decide whether you still need the client's consent anyway.

That last part surprises people. Oregon's bar ethics committee has said directly that anonymisation does not automatically end the inquiry, because some fact patterns are distinctive enough to identify a client with or without a name attached.

Why removing the name isn't enough

In Formal Opinion No. 2025-205, the Oregon State Bar addressed AI use directly and pointed back to its own earlier guidance on hypotheticals, Formal Ethics Opinion No. 2011-184. That opinion warned that "framing a question as a hypothetical is not a perfect solution," because lawyers face a real risk of violating the duty of confidentiality "when the facts are so unique... that other circumstances might reveal the identity of the consulting lawyer's client even without the client being named."

That's the exact failure mode of most AI prompts lawyers write today. A solo practitioner in a mid-size town describes "a $2.3 million earn-out dispute involving a regional dental practice acquisition" without naming anyone, and anyone who knows the local market can work out who that is in about ten seconds. The dollar figure, the deal type, the sector, and the geography together do the identifying work a name would have done. Generalise all four, not just the name, or don't share the fact pattern at all.

The step almost every guide skips: document metadata

Redacting the visible text in a Word document or PDF does nothing to the metadata riding underneath it. Author names, tracked changes, comment threads, and prior edit history all persist inside the file even after you've blacked out or deleted every name on the page. If you upload that file directly to an AI tool instead of pasting plain text, all of it goes with it.

This isn't theoretical. Metadata scrubbing has been a known malpractice and confidentiality risk in legal practice for over a decade around document exchange with opposing counsel, and the same exposure now applies to AI uploads. The fix is mechanical: strip metadata using your word processor's built-in inspector, or better, don't upload the file format at all. Copy the substantive text into plain text first. That single habit change closes a hole that pure text redaction cannot touch.

The actual four-step process

First, identify every category of identifier in the document: party names, addresses, account and case numbers, dates tied to a specific matter, dollar figures, and any detail unique enough to narrow the field to one person or company. Second, generalise rather than delete where deletion would strip out the legal substance you need the AI to reason about. "Client, a mid-size logistics company" survives with the analysis intact. "[REDACTED]" often doesn't, because the AI needs enough structure to give you a useful answer.

Third, run the metadata check described above before the document format leaves your machine in any form. Fourth, and this is the step Oregon's opinion makes clear most guidance skips entirely: ask whether the fact pattern, even fully anonymised, is still unique enough that someone familiar with your jurisdiction or practice area could identify the client. If yes, you're back to needing informed consent, anonymisation or not.

We've covered the consent and disclosure side of this separately in when lawyers have to tell clients they're using AI. Read that alongside this one. Anonymisation and disclosure solve different problems and a solid workflow needs both.

Open models, closed models, and why anonymisation still matters for both

Oregon's opinion draws a clear line between open AI models, which may train on what you input, and closed models that isolate and don't learn from your data. Anonymisation before using an open model without consent isn't optional under the opinion's language. But the opinion also says lawyers "may determine that it is appropriate to anonymize or redact certain information" even inside a closed model, because training exposure isn't the only risk in play. Vendor staff access, contract breach, subpoena exposure, and simple misconfiguration remain live regardless of whether the model trains on your prompt.

We go deeper on the open versus closed distinction, and which specific tools fall into each category, in what "do not train on my data" actually means in 2026. The anonymisation habit in this article should sit on top of whatever tool tier you're using, not replace the due diligence on the tool itself.

A workflow you can actually run in five minutes

Before pasting anything into an AI tool, run this sequence. Strip the file's metadata using your word processor's document inspector if you're uploading a file at all. Pull out every proper noun, number, and date that isn't legally necessary to the question you're asking. Generalise the remainder: sector instead of company name, range instead of exact figure, region instead of city. Read the anonymised version back and ask whether a colleague in your field could still guess the client. If the answer is yes, generalise further or don't send it.

None of this requires software. It requires five minutes and the discipline to do it every time, not just on the matters that feel obviously sensitive. The ones that don't feel sensitive are usually where the habit slips.

FAQ

Is it enough to just delete the client's name before pasting text into ChatGPT?

No, and this is the mistake most guides encourage. A name is one identifier among many. Dates, deal sizes, jurisdictions, unusual contract terms, and the sequence of events in a fact pattern can each narrow a matter down to one identifiable client, especially in a small practice area or a small town. Oregon's ethics committee flagged this exact problem in Formal Ethics Opinion 2011-184, warning that hypothetical framing fails when facts are unique enough to reveal identity even without a name attached.

Does anonymising a document remove the need for client consent to use AI on their matter?

Not automatically. Oregon Formal Opinion 2025-205 states that even with anonymisation or redaction in place, the client's informed consent may still be required, particularly for an open model. Treat anonymisation as a risk-reduction step you take regardless, not a substitute for the consent conversation when a matter's facts are unusual enough to still be identifiable.

Do closed, enterprise-tier AI tools still require anonymisation?

Oregon's opinion says lawyers may still determine it's appropriate to anonymise sensitive information even inside a closed model that doesn't train on inputs, because confidentiality risk isn't limited to model training. Contract breach, vendor staff access, subpoena exposure, and internal misconfiguration are all still live even when a vendor promises not to train. Treat closed-model status as one layer of protection, not the whole answer.

ShareX / TwitterLinkedIn
Was this useful?

Disclaimer · Educational content about software and productivity, not legal advice. AI tools and regulatory guidance change frequently, so always evaluate any tool against your own firm's obligations and your regulator's current guidance (e.g. the SRA in England & Wales, or your state bar / the ABA in the US) before using it with client data.

Free starter kit

Want the safe-tools shortlist as a PDF?

10 lawyer-safe AI tools, 12 ready-to-use prompts, and a client-confidentiality checklist for the SRA (UK) and ABA Rule 1.6 (US). Free, no spam.

Get the free Starter Kit →

Go deeper: The Lawyer's AI Toolkit (£29) →