How to Evaluate Legal AI Tools: A Buyer's Checklist

Evaluate a legal AI tool on four axes, in this order: security (where your client data goes and who can see it), accuracy (whether the vendor will let you test its claims on your own documents), pricing (whether a real number exists before you are deep in a sales cycle), and ownership (who owns the company, and what happens to your contract if that changes). Most published evaluation frameworks cover the first two well and skip the last two entirely, because the last two are the questions vendors least want asked. Run a scoped pilot on real matters with a measurable baseline, insist on written answers to the 20 questions at the bottom of this page, and treat any refusal to answer in writing as a finding rather than a formality.
We do not sell any of these tools. Nothing here is sponsored, no vendor pays for placement in our rankings, and featured listings are labeled and never reorder results. That matters for a page like this one, because almost every legal AI buying framework in circulation was written by a company that sells legal AI, and those frameworks tend to be shaped so that the publisher's product passes.
What should you ask before any demo?
The demo is where evaluation goes wrong. A scripted demo on the vendor's own sample documents will always look good. By the time you are watching one, you should already know four things.
What specific, repeated task are you buying this for? Not "improve efficiency." A task with a volume you can count: first pass NDA review, deposition summarization, discovery response drafting, research memos on a defined body of law. If you cannot name the task and roughly how many times a month your team does it, you cannot measure whether the tool helped, and you will end up renewing on vibes.
What does your current process cost? Hours, cycle time, or write offs. You need a baseline before the pilot, not after. Nearly every disappointed legal AI buyer I have talked to skipped this step and then had no way to argue with the vendor's own success metrics at renewal.
Who has to say yes internally? IT security, the managing partner, the ethics or risk lead, and whoever owns the document management system. Bring them in before the demo, not after you have picked a favorite.
What is the tool actually built on? Ask whether it is a purpose built legal system with its own retrieval layer over primary law or a wrapper around a general model with a legal prompt. Both can be legitimate. They are not worth the same money, and the price difference between them is frequently invisible.
Then ask the vendor to run the demo on your documents. A vendor that will not do this on a redacted or synthetic set of your own material is telling you something, and it is rarely good.
How do you test accuracy claims?
Accuracy marketing in this category has almost no shared definition behind it. "99% accurate" can mean accuracy against a vendor built benchmark that the vendor also grades, and vendors increasingly publish their own named benchmarks alongside their own products.
The honest reference point is independent testing. Stanford's RegLab and Human Centered AI institute ran the best known outside evaluation of legal research assistants, and the results were sobering: in their testing, two purpose built legal research tools hallucinated more than 17% of the time and one more than 34%, counting both wrong answers and citations that did not support the proposition they were attached to. That study was published in May 2024 and the products have all shipped many versions since, so do not treat those figures as current scores. Treat them as evidence about the category: retrieval grounding reduces hallucination, it does not eliminate it, and the vendors most confident in their accuracy numbers were not the ones who came out best.
So test it yourself. A workable accuracy test looks like this:
- Assemble 20 to 30 real queries or documents from your own practice, including a few you already know the correct answer to, and a few with a false premise built in ("summarize the holding in Smith v. Jones on this issue" where no such holding exists).
- Have a lawyer who knows the matter grade the output, not the person championing the purchase.
- Score three things separately: is the answer right, is every citation real, and does each citation actually support the sentence it sits under.
- Ask the vendor to explain any failure. The quality of that explanation is itself data.
The false premise queries matter most. A system that confidently invents a case rather than saying it found nothing is a system that will eventually put a fabricated citation in front of a judge, and courts have been notably unsympathetic about who is responsible when that happens.
What security questions are non-negotiable?
Four, and they are all questions of fact rather than judgment.
Is your data used to train the vendor's models? The answer must be no, and it must be no in the contract, not in a blog post. Ask specifically about model training, fine tuning, evaluation sets, and human review of your prompts for quality assurance. Those are four different things and a vendor can truthfully say no to the first while doing the fourth.
Where does the data live and who is the subprocessor? Most legal AI runs on someone else's foundation models. Ask which provider, in which region, under what data processing terms, and whether zero retention is configured on that provider's API. A vendor that cannot answer this crisply has not thought about it.
What independent audit can they show you? SOC 2 is the practical floor. The AICPA's SOC suite of service organization controls reporting covers examinations of controls relevant to security, availability, processing integrity, confidentiality, and privacy. Ask for the actual report under NDA rather than the badge on the marketing site, and check whether it is a Type 2 covering a real observation period. For AI specific governance, the NIST AI Risk Management Framework, released in January 2023 and organized around four functions (govern, map, measure, manage), is the vocabulary most serious vendors will already know. If a vendor cannot describe its practices in those terms or an equivalent, it is early.
How does access control map to your matter structure? Confidentiality walls, matter level permissions, and conflicts separation have to survive contact with the AI layer. A tool that indexes every document in your DMS into one searchable pool has quietly dissolved your ethical walls.
Security review is also the step most likely to be the reason you should not buy. That is a good outcome, and it is cheaper before signature than after.
Why does vendor ownership matter?
This is the question that vendor authored checklists never include, and it is the one that most often turns a good purchase into a bad one two years later.
Legal AI is consolidating quickly, and the pattern is well established. Casetext went to Thomson Reuters in 2023 and became CoCounsel. Lexion went to Docusign in May 2024. Evisort went to Workday in October 2024. Gavel was acquired by Relativity in June 2026. Names change too: Leya became Legora, Latch became Ivo, Callidus became StrongSuit. We keep the running list in our legal AI acquisitions and rebrands tracker.
Three practical consequences for a buyer:
Pricing and packaging change. A tool that was standalone and self serve can become a module of a larger suite with a suite sized price. Ask what happens to your rate on renewal if the company is acquired, and get a price protection clause for the term if you can.
Roadmap and integrations change. Acquirers deprecate overlapping features. If you bought the tool for its integration with a system the acquirer competes with, that integration is at risk.
Your research goes stale silently. This is the mundane one that catches people. Reviews, benchmarks, and forum threads about a product filed under its old name still describe the product accurately, and if you only search the new name you will find nothing but the vendor's own pages. Search both.
So ask: who owns you, who funds you, when did you last raise, and what is your runway. A well run vendor answers those questions. A vendor that treats them as impertinent has told you how the relationship will go.
How do you run a fair pilot?
A pilot that only proves the tool is fun to use has proved nothing. Structure it.
Scope it to one workflow and one team. Two to five users doing the same task, for four to six weeks. Shorter than four weeks and habit has not formed, so you measure novelty. Longer than eight and the pilot becomes the deployment by default and the evaluation quietly ends.
Define success before you start. Two or three numbers, agreed in writing with the vendor: time per unit of work against your baseline, an accuracy or rework rate, and adoption (what fraction of eligible work actually ran through the tool). Adoption is the one people forget, and it is the best single predictor of whether the renewal will be worth it.
Include a skeptic. Every pilot staffed entirely with volunteers succeeds. Put one person on it who thinks the whole category is overhyped, and take their objections seriously.
Run the same work both ways where you can. A parallel sample, even a small one, is worth more than any vendor case study.
Write the exit down. What happens to your data at the end of a pilot that does not convert, and how do you get it back or confirm deletion. Ask before, not during.
When should you walk away?
Some findings are disqualifying regardless of how good the product looks:
- The vendor will not run the demo on your material.
- The training and retention answers change between the sales engineer and the contract.
- There is no independent security audit and no timeline for one, on a tool touching client confidential data.
- The tool cannot produce a source for its output, or produces sources that do not support the claim.
- The pricing structure only becomes clear after you commit to a term.
- Nobody at the vendor can tell you who owns the company.
- Success in the pilot depends entirely on one enthusiastic partner. When they get busy, usage will fall to zero.
Walking away is not a failed evaluation. The alternative is a shelf ware line item and a security review you now have to unwind.
The 20 question legal AI vendor evaluation checklist
Print this, send it to the vendor, and ask for written answers before the second call. The questions are ordered so that the cheapest disqualifiers come first.
Security and confidentiality
| # | Question | What a good answer looks like |
|---|---|---|
| 1 | Is our data used to train, fine tune, or evaluate your models? | No, stated in the contract, covering all four uses |
| 2 | Which foundation model providers do you use, and in which regions? | Named providers, named regions, zero retention configured |
| 3 | Do humans at your company or your subprocessors review our prompts or outputs? | If yes: when, under what controls, and how to opt out |
| 4 | What is your data retention period, and how do we get deletion confirmed? | A specific period and a documented deletion process |
| 5 | Can you provide a current SOC 2 Type 2 report under NDA? | The report itself, with a real observation window |
| 6 | How do matter level permissions and confidentiality walls carry into the tool? | Permissions inherited from the DMS, demonstrable in the demo |
| 7 | What happens to our data if we do not renew? | Export format, timeline, and written deletion confirmation |
Accuracy and verification
| # | Question | What a good answer looks like |
|---|---|---|
| 8 | What sources ground the output, and are they licensed primary law? | Named corpora, not "the model knows the law" |
| 9 | Does every assertion carry a citation a lawyer can click through? | Yes, linked to the source passage |
| 10 | What does the system do when it finds nothing relevant? | Says so. Does not synthesize an answer |
| 11 | How was any accuracy figure you quote measured, and by whom? | Method disclosed, ideally graded by someone other than the vendor |
| 12 | Will you run our test set of real queries during evaluation? | Yes, without conditions |
| 13 | What are the known failure modes of your product? | A candid list. A vendor with no answer has not looked |
| 14 | How are model updates tested and communicated to us? | Regression testing and advance notice, not silent swaps |
Pricing and commercial terms
| # | Question | What a good answer looks like |
|---|---|---|
| 15 | What is the price, in writing, before a term commitment? | A number or a rate card, not a discovery call |
| 16 | Is it per seat, per usage, or bundled, and what triggers overage? | Clear unit of pricing and a stated overage rate |
| 17 | Are there seat minimums, implementation fees, or training costs? | All disclosed up front |
| 18 | What is the renewal uplift cap and the exit notice period? | A capped uplift and a notice window you can actually meet |
Ownership and viability
| # | Question | What a good answer looks like |
|---|---|---|
| 19 | Who owns the company, when did you last raise, and what is your runway? | Answered directly, without deflection |
| 20 | What happens to our pricing and roadmap if you are acquired? | Contractual price protection for the term, at minimum |
Two notes on using it. First, send the questions in writing and keep the written answers. A verbal promise in a sales call does not survive an acquisition. Second, treat question 15 as diagnostic rather than pass or fail. Most of this market gates pricing behind a demo, and gating is not by itself a red flag. Of the 93 tools in our directory, only eight publish a price you can read without a sales conversation: Paxton at $499 per user per month or $2,999 per year, Genie AI with a free plan and Pro at $75 per month, Gavel with a free tier and paid at $160 per user per month, MyCase from $50 to $130 per user per month, Rev with published per minute transcription rates, and three consumer services. The rest quote. Our legal AI pricing guide explains why, and what the gated majority reportedly costs.
Frequently asked questions
How do law firms evaluate AI vendors?
The firms that do it well run a structured process rather than a demo and a gut call: define the task and a cost baseline, involve IT security and risk before selecting a favorite, test the tool on real firm material with a lawyer grading the output, run a scoped pilot with pre agreed success metrics, and put the security and data handling answers in the contract. The firms that do it badly buy on a partner's enthusiasm and discover the constraints at renewal.
What security certifications should legal AI have?
SOC 2 Type 2 is the practical baseline for any tool touching client confidential data, and you should ask for the report rather than the badge. ISO 27001 is common among vendors selling internationally, and ISO 42001 is emerging as the AI specific management system standard. Beyond certifications, ask how the vendor maps to a recognized risk framework such as the NIST AI Risk Management Framework. This is a technology diligence standard, not legal advice about your own professional obligations, which are governed by your jurisdiction's rules.
How long should a legal AI pilot run?
Four to six weeks with a small, defined group is the range that works. Under four weeks you are measuring novelty rather than habit. Past eight weeks the pilot tends to become an undeclared deployment, and the evaluation never formally concludes. Set the end date and the success metrics before day one.
What questions reveal a weak legal AI product?
Three do most of the work. Ask what the system does when it finds nothing relevant: a weak product synthesizes an answer anyway. Ask how any accuracy claim was measured and by whom: a weak product cannot describe the method. Ask to run your own test set during evaluation: a weak product's vendor finds reasons why that is not possible. A fourth, softer signal is a vendor that answers "who owns you and what is your runway" with irritation instead of an answer.
Every tool in our directory is verified active, categorized, and listed with whatever pricing the vendor actually publishes. If you are still assembling a shortlist, our comparison of every AI legal tool we index is the place to start. Then take the 20 questions above and run them against the shortlist you build at /tools, starting with the category you are actually buying for, such as legal research or contract review.
This page is software evaluation and industry analysis. It is general information about buying technology, not legal advice, and it is not guidance on your professional responsibility obligations. Consult your jurisdiction's rules and your own risk counsel.