Start with the data that exists on Tuesday
Before anyone scopes an AI project for a law firm, a corporate legal department, or an agency general counsel office, the useful first question is not what the data could be. It is what a machine can read on Tuesday morning without a six-month cleanup. In legal services the answer splits cleanly, and it splits the same way almost everywhere we look. Time and billing data is structured, dated, coded, and complete. Executed contracts are usually findable but almost never parsed. Work product in the document management system is enormous and unlabeled. Custodial email is the largest pile and the least usable until someone pays to process it.
The systems are named and knowable. Work product lives in iManage, NetDocuments, or SharePoint. Time and billing lives in Elite 3E, Aderant, or a smaller practice-management package. Signed agreements live in a contract lifecycle system or, more often, a shared drive and an inbox. Matter-specific collections live in a review platform such as Relativity, Everlaw, Reveal, or DISCO. Docket data comes from PACER, the federal judiciary's public access service, which charges ten cents a page with a three-dollar cap per document and waives balances under thirty dollars a quarter. Each of those stores has a different readiness level, and treating them as one corpus is the first mistake buyers make.
The readiness gradient matters more than the total volume. A firm with forty terabytes of unlabeled work product and eight hundred megabytes of clean billing records will get faster, more defensible value from the eight hundred megabytes. That is not an argument against the large corpus. It is an argument for sequencing.

Machine-readiness by legal data type
Editorial weighting from public sources and practitioner reading. Illustrative, not a measured statistic.
Three families of work, and mostly only three
Almost every credible legal AI use case we have scoped falls into one of three families: review at volume, contract analytics, and forecasting from billing data. A fourth, knowledge retrieval over a firm's own work product, is real but sits behind an access-control problem that most firms have not solved. Everything else tends to be a feature of one of these.
| Work family | Typical volume | Latency the work tolerates | What decides success |
|---|---|---|---|
| Review at volume | 200 GB to 5 TB collected per large matter; 300K to 3M documents after culling | Overnight batch; hours are fine | Defensible recall measurement and a documented stopping rule |
| Contract analytics | 5,000 to 100,000 executed agreements; 15 to 60 pages each | 20 to 90 seconds per agreement; overnight for the full corpus | Text-layer quality and version and amendment chaining |
| Matter and spend forecasting | 3 to 10 years of coded time entries; often under 5 GB | Nightly refresh | Coding discipline at entry and enough comparable matters |
| Work-product retrieval | 1 to 50 TB of DMS content across all matters | Sub-second search; 2 to 8 seconds for a cited answer | Ethical walls enforced in the index, not the interface |
| Intake and conflicts triage | Thousands of parties and affiliates per year | Minutes; humans are waiting | Entity resolution across name variants and corporate families |
Review at volume: the number that makes the argument
Document review is where the money is, because it is where the hours are. Take a mid-size commercial dispute with fifteen custodians and a five-year window. Individual mailboxes commonly run ten to fifty gigabytes each, so the collection lands somewhere between 150 GB and 750 GB. Deduplication, email threading, date filtering, and domain culling typically remove half or more. Call it 200 GB going to review.
Conversion from gigabytes to documents is where vendors quietly disagree. Depending on file mix, a gigabyte of email plus attachments expands to somewhere between 5,000 and 15,000 documents. Anyone who quotes a single conversion factor without looking at the file mix is guessing. Take the middle: 200 GB becomes roughly 1.6 million documents. A trained contract reviewer working at 40 to 60 documents an hour needs on the order of 30,000 hours. At rates in the $35 to $75 range, first-pass review alone runs past a million dollars before a single privilege log entry is written. Processing at $25 to $100 per gigabyte and hosting at $5 to $20 per gigabyte per month sit on top.
That is the arithmetic every legal AI conversation is really about. It is also why the technical argument long ago stopped being about whether machine classification works and became an argument about how you prove it worked.
Recall, not accuracy, is the currency
The foundational study here is still Blair and Maron in 1985. Attorneys running keyword searches on a litigation corpus believed they were retrieving about 75 percent of the relevant documents. Measured recall was closer to 20 percent. That gap is the reason technology-assisted review exists, and it is the reason a court will entertain a machine-driven process at all: the human baseline was never as good as the profession assumed.
The case law followed. Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182 (S.D.N.Y. 2012), was the first judicial approval of predictive coding. By Rio Tinto PLC v. Vale S.A., 306 F.R.D. 125 (S.D.N.Y. 2015), the same court called it black-letter law that technology-assisted review is an acceptable method. What courts have never blessed is an unmeasured process. Negotiated recall targets in the 75 to 80 percent range, a randomly sampled control set, and an elusion test on the discard pile are the mechanics that make a stopping decision defensible under the proportionality standard in Rule 26(b)(1) of the Federal Rules of Civil Procedure.
Language models change the cost side of this arithmetic sharply. A three-page document is roughly 2,000 tokens. At input pricing in the low single dollars per million tokens, reading a million documents once costs in the low thousands of dollars. Against a million-dollar review, inference cost is a rounding error. The expensive part is no longer the reading. It is the adjudicated gold set, the sampling plan, and the written protocol that lets you defend the result at a meet-and-confer.
Contract analytics: a smaller corpus and a harder truth
Contract work involves far less data and far more pain per document. A mid-market company holds perhaps 5,000 to 40,000 executed agreements. A large enterprise or a federal prime holds several times that. The documents are PDFs, many of them scans of signed originals, some of them scans of faxes of signed originals. Optical character recognition quality is the gate, and no model recovers from a text layer that lost the word "not."
What buyers actually want extracted is a short and stable list: renewal and expiration dates, auto-renewal notice windows, assignment and change-of-control triggers, limitation-of-liability caps, indemnity scope, governing law, termination-for-convenience rights, and any data protection addendum. Companies selling to the federal government add a flowdown layer, including the commercial-item clause flowdown at FAR 52.244-6 and the safeguarding and reporting obligations under DFARS 252.204-7012 for covered defense information. The measurable output is an obligation register: one row per obligation, with a date, an owner, a trigger, and a citation to the page and clause it came from.
The failure mode here is subtle and worth naming, because it is the one that damages trust fastest. Getting an extraction right is easy to check. Proving that a clause is genuinely absent from an agreement is much harder, and a confident silence is more dangerous than a wrong answer. We wrote about that specific problem at length in contract review when the stakes are legal.
Forecasting is a billing-data problem, not a language problem
The most underused asset in a legal organization is its own billing history. Time entries submitted under the LEDES 1998B standard carry a task code, an activity code, a timekeeper, a date, hours, and a rate. The Uniform Task-Based Management System behind those codes was developed in 1995 by the American Bar Association, what was then the American Corporate Counsel Association, and PricewaterhouseCoopers, and it is still how most outside counsel bills large clients. Litigation codes run L100 for case assessment through L400 for trial preparation, with separate A-codes for activities and E-codes for expenses.
That is structured, dated, coded, complete data measured in gigabytes rather than terabytes, and it supports questions that partners and general counsel care about: cost to completion on an open matter, cycle time by matter type, staffing-mix drift, budget variance by phase, and which matters are about to breach their approved budget. A defensible target to write into a statement of work is a mean absolute percentage error inside 20 to 30 percent at a 90-day horizon, on matter types with at least a couple of hundred comparable historical analogues. Below that sample size, the honest answer is that the history cannot support a forecast yet, and saying so early is worth more than a confident number.
The rules that shape the architecture
Professional-responsibility obligations are not a compliance appendix in legal AI. They are architectural requirements, and they arrive before the first design review.
Confidentiality. ABA Model Rule 1.6(c) requires reasonable efforts to prevent unauthorized disclosure of information relating to a client. Formal Opinion 477R applied that to electronic communication, and Formal Opinion 483 addressed obligations after a breach. In practice this pushes toward single-tenant or in-tenant deployment, encryption with customer-controlled keys, and contractual prohibitions on any use of client data for model training.
Generative AI specifically. ABA Formal Opinion 512, issued in July 2024, is the current reference point. It addresses competence, confidentiality including client consent before inputting client information into self-learning tools, communication with clients about the use of such tools, and fees. On fees the opinion is blunt: a lawyer cannot bill for hours the tool saved and the lawyer never worked.
Supervision of vendors. Model Rule 5.3 extends supervisory responsibility to nonlawyer assistance, and that includes outside technology providers. A firm cannot outsource the duty. The practical consequence for a build is auditability: every automated determination needs a log entry, a model and prompt version, and a traceable source citation.
Competence. Comment 8 to Model Rule 1.1 established the duty to keep abreast of the benefits and risks of relevant technology, and the large majority of states have adopted it in some form. The floor it sets is that nobody signs off on output they cannot check.
Client-imposed terms. Outside counsel guidelines increasingly restrict subprocessors, name approved data locations, and prohibit training on matter data. These are contractual, they vary by client, and they frequently bind harder than the ethics rules. Any architecture that cannot segregate one client's data under a stricter rule than another's will fail an audit.
Two more constraints show up in cross-border matters. Article 48 of the EU General Data Protection Regulation limits transfers in response to foreign court orders absent an international agreement, and blocking statutes such as French Law No. 68-678 create direct exposure for discovery conducted outside the Hague Evidence Convention. Neither is exotic. Both change where the compute has to run.
Why data quality decides the outcome
Model selection is the most discussed and least decisive variable in a legal AI build. In document workflows the gap between a well-prepared corpus and a raw one moves end-to-end accuracy further than any model swap our engineers have measured, and it does so at a fraction of the argument.
The specific defects repeat across organizations. A scanned agreement with a text layer produced by an old OCR engine drops negations and misreads dollar figures. The same counterparty appears as six different strings across the repository, so a rollup by entity is wrong before any model sees it. Unsigned drafts sit next to executed originals with no reliable flag distinguishing them. Amendments exist as separate files with no link to the parent agreement, so an extraction reports the original liability cap that a later amendment doubled. Date fields hold the scan date rather than the execution date. Email families get broken during export, so an attachment is reviewed without its parent.
Each of those is fixable, and fixing them is unglamorous engineering: re-OCR at a measured character error rate, deduplicate while preserving family integrity, resolve entities against a canonical party list, classify document type before extracting anything, normalize dates against the signature block, and chain versions and amendments into a single document lineage. Budget 60 to 70 percent of a first engagement to this work. A proposal that allocates less than half is describing a demo.
Latency and cost, concretely
Legal work has a wide tolerance for latency, which is a gift to the architecture. Interactive search over an indexed corpus should return in well under a second at the 95th percentile. A grounded answer with citations to source documents should land in two to eight seconds. A full analytical pass over a forty-page master services agreement, roughly 25,000 tokens, takes twenty to ninety seconds and costs single-digit cents at current input pricing. Nobody is waiting on a full-corpus re-read, so that runs overnight.
The cost structure surprises buyers who expect inference to dominate. Re-reading a 100,000-document contract corpus costs in the hundreds to low thousands of dollars. Hosting and storage for a large litigation collection can cost that much every month. Human adjudication of a gold set of 500 to 1,500 documents, at senior-associate or partner review rates, often costs more than a year of inference. That gold set is also the single line item most likely to get cut, and cutting it is how a project loses the ability to prove anything.
Federal legal work has its own shape
Agency general counsel offices, litigation support contractors, and records shops face the same volume arithmetic with a different overlay. The Freedom of Information Act, 5 U.S.C. § 552, generates well over a million requests government-wide per year according to the Department of Justice summary of annual FOIA reports, and the bottleneck is almost never finding responsive records. It is applying exemptions and redacting: (b)(5) for deliberative process and attorney work product, (b)(6) for personal privacy, and the (b)(7) law enforcement series. Redaction is a precision-critical task with an asymmetric error cost, which makes it a good fit for machine assistance under mandatory human review and a poor fit for autonomous action.
Records retention sits underneath all of it. Federal records obligations under 44 U.S.C. Chapter 33 and the National Archives Capstone approach to email, set out in NARA Bulletin 2013-02, determine what exists to be searched in the first place. Any interface delivered to an agency also carries Section 508 accessibility obligations under 29 U.S.C. § 794d, and any corpus containing controlled unclassified information carries the NIST SP 800-171 control set. None of this is optional, and pricing a federal legal AI engagement without it produces a bid that cannot be executed.
Scoping a first engagement that can fail cheaply
The best first engagement answers one narrow question, on a real slice of the actual corpus, against a metric agreed in writing before the work starts, on a fixed price, in under ten weeks. It ends with a number and a written go or no-go. If the number misses, the buyer has spent a defined amount to learn something true, and that is a successful outcome rather than a failure.
A ten-week first engagement
Five contract terms make the failure cheap rather than expensive. Fix the price so the downside is bounded. Name the acceptance metric and its threshold before the work begins, so nobody argues about the definition of success afterward. Keep production integration out of scope entirely, so a negative result costs nothing to unwind. Require that the code, the data-preparation steps, and the evaluation scripts are delivered regardless of outcome, so the buyer keeps the asset. And state in writing that a documented negative result is an accepted deliverable. In this market a scope of that shape usually prices between $45,000 and $90,000, and it is the cheapest way to find out whether the larger program is real.
Two things belong in the same conversation. Get a Rule 502(d) non-waiver order in place under the Federal Rules of Evidence before any privileged material moves, so an inadvertent production does not become a waiver. And agree the privilege-log approach under Rule 26(b)(5)(A) early, because log production is frequently the most expensive part of a review and the part automation helps most.
Bottom line
Legal services has more genuinely valuable AI work available than most sectors, for a plain reason: the underlying task is reading documents at volume, and the cost of doing it by hand is measured in millions of dollars per matter. What separates the projects that land from the ones that stall is rarely the model. It comes down to whether someone did the work of profiling the corpus, repairing the text layer, resolving the entities, building an adjudicated gold set, and writing down what success means before the first line of pipeline code. Our engineers build these systems the same way every time, and we would rather tell a client in week ten that the data will not support the claim than discover it in production.
Frequently asked questions
Usually forecasting from billing data, because the data is already structured, dated, and coded, and the corpus is small enough to work with in weeks rather than months. Document review has a larger dollar prize but a longer path, since the defensibility mechanics and the gold set have to be built first.
Yes, and they have since Da Silva Moore v. Publicis Groupe in 2012, with Rio Tinto PLC v. Vale S.A. in 2015 treating it as settled. What courts expect is a measured process: a negotiated recall target, a sampled control set, an elusion test on the discard set, and a documented stopping rule consistent with Rule 26(b)(1) proportionality.
That depends on Model Rule 1.6(c), ABA Formal Opinion 512, and the client's own outside counsel guidelines, which often bind more tightly than the ethics rules. The common resolution is deployment inside the firm's own cloud tenant with contractual prohibitions on training, customer-managed encryption keys, and per-client data segregation.
Sixty to seventy percent on a first engagement. Text-layer repair, deduplication with family integrity preserved, entity resolution, document-type classification, date normalization, and amendment chaining move accuracy further than any model choice, and they are the work most proposals underprice.
Eight to ten weeks and roughly $45,000 to $90,000 on a fixed price for one narrow question measured against an adjudicated gold set. The scope should exclude production integration entirely, so that a negative result is a cheap and useful answer rather than a stranded system.