Legal AI Data Security: What to Verify Before You Upload

Legal AI data security comes down to five questions: does the vendor (or the model provider behind it) train on your data, how long does it keep your prompts and files, which sub-processors can touch them, which independent audits back the controls, and where is the data hosted. None of those answers should come from a marketing page. Each one has a document that proves it: the master agreement or data processing addendum for training and retention, the published sub-processor list for the supply chain, a SOC 2 Type 2 report or ISO/IEC 27001 certificate for the controls, and the order form for hosting region. If a vendor will not produce those documents before you upload a client file, you have your answer.
We run a directory of AI legal software and we don't sell any of it. That matters here, because most of what gets written about legal AI security is written by vendors describing their own safeguards. A vendor explaining why its product is secure is not lying, but it is answering a different question from the one you need answered. You need to know how to check any vendor's claims, including the ones you already like.
This page is software evaluation, not ethics advice. Your professional obligations come from your own jurisdiction's rules, and the bar guidance quoted below is there to show what regulators expect lawyers to look at, not to tell you what your bar requires.
What happens to a document after you upload it?
It helps to picture the path a file takes, because every stop is a place where a copy can live.
- Your browser or add-in sends the file to the vendor's application.
- The vendor's application stores it (in a workspace, a vault, a matter folder) and usually breaks it into chunks and builds a search index so the AI can find relevant passages later.
- The model provider receives the prompt plus the relevant chunks. For most legal AI products, that model is run by a separate company: a foundation model lab or a cloud platform hosting one.
- Logs and monitoring systems at both the vendor and the model provider may record the request for debugging, abuse detection or billing.
- Backups keep copies of all of the above for disaster recovery.
A vendor that says "we delete your data when you delete it" is usually talking about step 2. Ask about steps 3 to 5 separately. The question that exposes the most is simple: "If I delete this file today, list every system where a copy or a derivative (index, embedding, log entry, backup) still exists, and for how long."
This is also why the State Bar of California's practical guidance on generative AI in the practice of law says confidentiality risk depends on how a system is used: an isolated prompt tool, one integrated with other applications, or an autonomous agent with ongoing data access. An agent wired into your email and document management system has a much longer data path than a chat window, and the 2026 version of that guidance says the degree of diligence must correspond to the level of access and autonomy.
Does the vendor train models on your data?
Nearly every legal AI vendor now says it does not train on customer data. The useful work is reading how that promise is scoped.
Look for three layers in the answer.
- The vendor's own models. Does the vendor fine-tune or improve its own models, classifiers or ranking systems with your prompts, documents or outputs? "We don't train foundation models on your data" can be true while your data still improves the vendor's retrieval or feedback systems.
- The model provider. The vendor's promise means little if the lab running the model can train on what it receives. The vendor should be able to say the provider is contractually barred from training on your data, and ideally that the provider runs with zero data retention for your traffic.
- Defaults versus options. Watch for "by default." Some vendors offer custom models trained on a firm's own data as an opt-in, which is fine if it is truly opt-in and isolated to you.
A worked example. As of September 2026, Harvey's public security page states that it contractually prohibits model providers from training on customer data and requires zero data retention from them. The same page says a firm can have its own data used for training, but only on explicit request, in a bespoke model for that customer. And a Harvey blog post from September 2026 disclosed that one newer third-party model it offers as an optional choice requires 30-day retention and limited safety review by the model's maker, which Harvey says it flagged publicly rather than folding in quietly.
That example is useful precisely because it is candid. It shows that "no training, zero retention" is a policy that can have exceptions per model, and that the exceptions can change as vendors add new models. So the question is not only "do you train on my data" but "how will you tell me when a model you route my data to has different retention terms, and can I switch it off?" (Harvey's listing is at /tools/harvey; check the vendor's current page, since these terms change.)
What proves it: the master services agreement or platform agreement, and the data processing addendum. A FAQ answer is not a contract.
Which certifications actually mean something?
Certifications are where marketing language does the most work, so read the verbs.
SOC 2 Type 2. A SOC 2 report is an independent CPA firm's examination of a service organization's controls relevant to security, availability, processing integrity, confidentiality or privacy, under the framework the AICPA maintains for SOC reporting. A Type 1 report looks at whether controls are designed properly at a point in time. A Type 2 report tests whether they actually operated over an observation period. For a tool holding client files, Type 2 is the one to ask for. Things to check in the report itself:
- The observation period and whether it is recent.
- The scope: which product and which trust services criteria are covered. A report on a company's older product does not cover its new AI assistant.
- Exceptions noted by the auditor, and management's responses.
- Complementary user entity controls: the things the report assumes you do, such as managing user access. These are your homework, not theirs.
- Carved-out sub-service organizations: if the cloud host or model provider is carved out, the report says nothing about their controls.
The AICPA itself has spent 2026 publishing warnings about "fast and easy" and quick-turn SOC engagements, which is a good reason to read who the auditor is and how long the observation window ran rather than trusting the badge.
ISO/IEC 27001. The current edition is ISO/IEC 27001:2022, the international requirements standard for an information security management system. Certification is issued by a certification body after an audit, so ask for the certificate, the issuing body, and the scope statement. "Aligned with ISO 27001" or "built on ISO principles" is not certification.
AI-specific and privacy extensions. ISO/IEC 27701 (privacy information management) and ISO/IEC 42001 (AI management systems) are appearing on legal AI trust pages. They are real signals of maturity, but they are newer, so treat them as supplements to SOC 2 or 27001, not replacements.
Read the verbs. On the same vendor page, "controls aligned with SOC 2, ISO, GDPR" and "undergoes annual SOC 2 Type II and ISO 27001 audits" are two very different sentences. The second is a claim you can verify with a document. The first is a description.
Where is client data stored and who can reach it?
Hosting location matters for client commitments, cross-border rules and your own outside counsel guidelines. The questions:
- Which cloud and which region hosts the application, the storage and the search index?
- Where does model processing happen? The storage can sit in one region while prompts go to a model endpoint in another. Ask whether regional processing applies to sub-processors too.
- Who at the vendor can access customer content, under what approval process, and is that access logged and reviewable by you?
- How is your data separated from other customers' data? Logical separation per tenant is common; ask how it is enforced and tested.
- What does the sub-processor list include, and how much notice do you get before a new one is added? Many vendors publish this list; a vendor that won't share it is asking you to accept an unknown supply chain.
Ownership is part of this question too. North Carolina's 2024 Formal Ethics Opinion 1 on AI in a law practice, adopted November 1, 2024, carries forward earlier vendor-selection factors that include the company's stability and whether the terms say how your information will be returned or destroyed if the company goes out of business, changes ownership or the service ends. Legal AI has had a steady run of acquisitions and rebrands (we log them in the legal AI acquisitions tracker), so the change-of-control clause is not hypothetical.
What should the contract say about security?
Everything above only binds the vendor if it is in the paper. The contract stack usually has three parts: the master agreement, a data processing addendum, and sometimes a security addendum. Check that it covers:
- No training on inputs, outputs or uploaded documents, by the vendor or any sub-processor, with the opt-in exception spelled out.
- Retention limits for content, logs and backups, with deletion on request and on termination, plus a way to confirm deletion.
- Sub-processor flow-down: the vendor passes its confidentiality and no-training obligations to model providers and other sub-processors.
- Change notice before new sub-processors or models with different data terms are used for your data.
- Breach notification within a defined window, with cooperation obligations.
- Audit rights or, at minimum, annual delivery of the current SOC 2 report.
- Data return and destruction on termination, including after an acquisition.
- Ownership: the vendor claims no rights in your content or outputs. The NC opinion quotes a proposed Florida Bar advisory opinion on exactly this point, advising lawyers to determine whether a provider retains submitted information after services end or asserts proprietary rights to it.
Security is only one part of a legal AI decision. Accuracy is the other big risk, and the court record on that is in our AI hallucination sanctions roundup.
The before-you-upload checklist
Print this and send it to the vendor before the pilot, not after. Every row names the document that settles it.
| # | Verify | A good answer | The document that proves it |
|---|---|---|---|
| 1 | No training on your data by the vendor | Inputs, outputs and files never used to train or improve any shared model; custom training opt-in only | Master agreement or platform agreement |
| 2 | No training or retention by the model provider | Provider contractually barred from training; zero data retention, or a stated short window | Data processing addendum and sub-processor terms |
| 3 | Per-model exceptions disclosed | A list of any models with different retention terms, switchable off by an admin | Vendor model or sub-processor notice |
| 4 | Retention for files, indexes, logs, backups | A number of days for each, plus deletion on request | Retention policy and DPA |
| 5 | Deletion you can confirm | Written confirmation or certificate of deletion on termination | DPA termination clause |
| 6 | Sub-processor list | Named companies, their role, their region | Published sub-processor list |
| 7 | Notice of changes | Advance notice before new sub-processors, with a right to object | DPA change clause |
| 8 | Independent audit of controls | Current SOC 2 Type 2 covering the AI product you are buying | The SOC 2 report under NDA |
| 9 | Information security certification | ISO/IEC 27001 certified, with scope covering the product | Certificate and scope statement |
| 10 | Hosting and processing region | Storage, index and model processing all in your required region | Order form or data residency addendum |
| 11 | Staff access to content | Access only on approved request, logged, visible to you | Security addendum and access policy |
| 12 | Tenant separation | Customer data logically separated and tested | SOC 2 report and pen test summary |
| 13 | Breach notification | A defined window and a named contact | Security addendum |
| 14 | Change of ownership | Data return or destruction rights survive an acquisition | Master agreement assignment clause |
| 15 | Agent permissions | Scoped access to email or DMS, with human review before anything leaves the firm | Admin documentation and configuration |
A vendor does not need a perfect row 1 to 15 to be usable. It needs written answers you can file, and gaps you have consciously accepted.
How do you compare vendors on security fairly?
Trust pages are written to be compared on their own terms, which is why side-by-side badge counts are nearly useless. A fairer method:
- Send every vendor the same questions, the checklist above, in writing. Score the answers, not the slide deck.
- Score evidence, not claims. Give full marks only when a row is backed by the document in the last column. A claim with no document scores zero, however confident it sounds.
- Weight rows to your risk. A solo practitioner drafting from public filings and a firm uploading merger data rooms should not weight hosting region the same way.
- Recheck at renewal. Vendors add models and sub-processors. The California guidance notes a user may agree to terms and a privacy policy simply by using a product, so the terms you reviewed at signing may not be the terms you run on a year later.
- Keep the file. The written answers become your evidence of the reasonable efforts every bar guidance document above asks for.
For the wider buying process (accuracy testing, pricing, ownership and pilots), use our 20 question legal AI evaluation checklist, and for contract tools specifically, the procurement notes in our AI contract review software guide. If your firm is building an AI policy, the AI governance platforms comparison covers the tools that track this across vendors.
You can browse every tool we list, with pricing lines and ownership, in the legal AI directory. Vendors: if you publish your security documentation (a trust center, a sub-processor list, a SOC 2 report available on request), say so in your free listing through submit a tool. Buyers notice.
FAQ
Is legal AI safe for confidential client documents?
It can be, but "legal AI" is a category, not a security level. A tool is reasonable for client files when it contractually rules out training, keeps retention short and defined, publishes its sub-processors, holds a current SOC 2 Type 2 or ISO/IEC 27001 certification covering the product, and hosts data where your commitments require. Bar guidance such as California's says reasonable efforts require more than generalized marketing assurances, so verify with documents before uploading. Whether a given use meets your own professional obligations is a question for your jurisdiction's rules.
Do legal AI tools train on client data?
Most established legal AI vendors say they don't, and many put it in the contract. The details to check are whether the promise covers the model provider as well as the vendor, whether it covers outputs and logs as well as uploads, and whether any optional model or feature carries different terms. Some vendors offer opt-in custom training on a firm's own data, isolated to that firm.
What is the difference between SOC 2 and ISO 27001?
SOC 2 is an attestation report by an independent CPA firm on a service organization's controls, under an AICPA framework, and a Type 2 report tests those controls over a period of time. ISO/IEC 27001:2022 is an international standard for an information security management system, and certification is issued by a certification body after an audit. SOC 2 is common with US buyers, ISO 27001 with international ones, and many legal AI vendors hold both. Ask for the report or certificate and check that its scope covers the product you're buying.
Is a general chatbot less secure than legal AI?
Not automatically, and the difference is usually the terms rather than the technology. General chatbots often carry different data terms on different plans, so the plan you are signed in on decides what happens to your input. North Carolina's 2024 opinion advised that, as of its date, lawyers should generally avoid inputting client-specific information into publicly available AI resources. A legal AI product built on the same underlying model can be safer if its contract adds no-training, zero retention and regional hosting. Judge the terms you are actually under, whichever product it is.