Prove the review worked — with numbers.
The fight over technology-assisted review stopped being about permission years ago. It is now about proof: whether the producing party can show the result was adequate, and whether the requesting party can show it was not. Both are statistical questions, and both are answerable.
What validation actually measures
A TAR workflow is a classifier applied to a population. Its adequacy is a measurable property of the output, not of the software brand or the vendor's methodology deck. The measurements that matter are few and well understood.
The measurements a defensible protocol produces
- Recall — the share of responsive documents in the population that the process actually surfaced, reported with a confidence level and margin of error rather than as a bare number.
- Precision — the share of documents the process called responsive that in fact are, which bears on cost and review burden more than on adequacy.
- Elusion — the responsive rate within the discard pile, estimated from a random sample of what was not produced. This is the direct test of what was missed.
- Richness — the prevalence of responsive material in the collection, which determines how large a sample must be before any of the above means anything.
- Stability — whether the classifier's behavior had converged when review stopped, or whether the cutoff was driven by budget.
Where protocols fail under examination
Sample size is the most common defect. A validation sample sized for convenience rather than for the collection's richness produces a confidence interval wide enough to be consistent with both an excellent and an inadequate review, which means it proves nothing. In low-richness collections — the ordinary case — the required sample is substantially larger than practitioners expect.
The second common defect is a control set contaminated by the training process, which inflates every metric derived from it. The third is a validation performed on a population that no longer matches what was produced, because the collection was expanded, de-duplicated, or family-grouped after the sample was drawn.
Disclosure disputes
How much of its process a producing party must reveal remains contested, and the answer is usually driven by proportionality under Rule 26(b)(1) rather than by any rule specific to TAR. A workable position is that validation results and sampling methodology are disclosable because they bear on the adequacy of the production, while the training decisions themselves generally are not. Getting that line drawn early — in the ESI protocol rather than in motion practice — avoids the most expensive version of this dispute.
Challenging a production
On the requesting side, the work is narrower and often more decisive: establishing whether the reported metrics support the conclusion drawn from them. A recall estimate without a stated confidence interval, an elusion test with no sample-size justification, or a validation run before the final de-duplication pass are each grounds to require more — and each is visible from the protocol and the metrics themselves, without access to the other side's documents.
Common questions
What recall rate is required for a defensible production?
No rule sets a threshold, and any expert who quotes one without reference to the collection should be questioned. Adequacy is assessed against reasonableness and proportionality, taking into account the richness of the collection, the burden of further review, and the marginal value of what additional effort would surface. What matters is that the estimate is properly derived and honestly reported with its uncertainty.
What is elusion testing?
A random sample drawn from the documents the process did not produce, reviewed to estimate how many responsive documents were left behind. It is the most direct evidence of what a review missed, and its absence from a validation protocol is a meaningful gap.
Must a party disclose that it used TAR?
There is no rule requiring it, but the practical answer is usually yes — disclosure in the Rule 26(f) conference and the ESI protocol converts a later challenge to the methodology into a dispute the parties already framed, rather than a motion arguing that the choice was concealed.
Does TAR 2.0 or continuous active learning change the validation analysis?
It changes when validation happens, not whether. Continuous learning blurs the training/review boundary, so a static control set is less useful and the stopping decision carries more weight. The adequacy question is still answered by sampling the discard pile after review stops.
Have a matter where this is the fight?
Send the matter name, jurisdiction, and key dates. You will get a prompt conflict check and a scoping conversation — not a sales call.
Retain the expert→Not ready to retain?
Track the rules and the case law instead. Practitioner notes, twice a month.