The AI errors your law firm may never see

An accuracy score can tell you how often a system agrees with its test data. It takes more work to establish whether it deserves a role in decisions about clients.

dashboardWISE Team12 min read

Suppose a firm introduces an AI tool to screen new enquiries. After a month, the report looks encouraging: less time spent on intake, fewer unsuitable consultations, a higher proportion of appointments becoming matters.

The partners approve wider use.

One question has gone unanswered: how many of the rejected enquiries should have reached a lawyer?

The firm has measured what happened to the work it accepted. It has not examined what the system removed from consideration. Potentially worthwhile matters could be disappearing without creating an obvious error message, complaint or missed target.

This deserves as much attention as the more conspicuous failures of legal AI. A fabricated authority gives a reviewer something to check. An enquiry that never reaches the reviewer presents a different problem: the firm must first decide to look for it.

For lawyers evaluating AI, that is a useful starting point. Before asking how much work a tool saves, establish which mistakes its performance report would fail to reveal.

What does “95% accurate” leave out?

Consider a deliberately simple example. A test collection contains 1,000 documents, of which 50 are privileged. A system labels every document non-privileged.

It is 95% accurate. It has also missed every privileged document.

The arithmetic is correct because the headline figure counts all correct classifications together. It conceals complete failure on the category that required particular care.

In a document-review workflow, incorrectly flagging an ordinary document as potentially privileged may create additional review work. Missing a privileged document could lead to its disclosure if nobody catches the error. Those consequences should influence how the firm tests the system and what it permits the system to do.

Ask for the results behind the percentage: how many privileged documents were found, how many were missed, and how many ordinary documents were incorrectly flagged. Then ask what kinds of documents appeared in the test. A result obtained from clean, searchable commercial agreements does not, by itself, establish performance on handwritten annotations or poor-quality scans.

This approach is consistent with NIST’s AI Risk Management Framework, which calls for documented evaluation in conditions resembling deployment and attention to the limits of a system’s ability to generalise beyond its testing. [7]

The same scrutiny applies to intake. A high conversion rate among accepted enquiries cannot tell a partner how many promising enquiries were rejected. Reviewing only the accepted work leaves a substantial part of the decision untested.

Your past matters are not an answer key

There is a second problem beneath the accuracy figure: what the system has been taught to regard as the correct answer.

Machine-learning evaluations often compare predictions with recorded outcomes or labels. But a reliable prediction of the recorded label may still be a poor answer to the question the firm actually cares about. As Barocas, Hardt and Narayanan explain in Fairness and Machine Learning, improving prediction of a badly chosen measurement does not resolve the underlying mismatch. [3]

Imagine the intake tool was developed using the firm’s historical enquiries. Those that became matters were labelled “good”; those that did not were labelled “poor”.

The distinction looks convenient until someone examines the rejected enquiries. Some might have involved conflicts. Others might have arrived when the relevant team lacked capacity. Some prospective clients might have declined the proposed fee. None of those circumstances necessarily tells the firm whether the legal problem deserved attention.

A model that reproduces those decisions accurately could be learning the firm’s previous capacity constraints and commercial preferences, rather than the merits of an enquiry.

There is an equivalent trap in financial data. Suppose a proposed payment-risk tool learns from overdue balances. An unpaid invoice might reflect a client’s reluctance to pay. It might also reflect a dispute, an agreed payment arrangement or a payment that has not yet been reconciled. Unless those distinctions are preserved, the model could turn an administrative condition into a judgement about the client.

Before accepting a prediction, ask what event or label the model predicts, who recorded it, and what else could explain it.

Removing protected characteristics from the inputs does not settle the fairness question either. Other features can carry related information, and the legal analysis cannot safely be reduced to whether a field labelled “race” or “sex” appeared in the dataset. Research on directly discriminatory algorithms also cautions against assuming that apparently indirect mechanisms necessarily fall only within indirect-discrimination analysis. The applicable law and the actual decision mechanism matter. [4]

Nor does every improvement require a sacrifice in accuracy. In the billing example, distinguishing an unreconciled payment from genuine non-payment could improve the prediction while reducing unfair treatment. The first task is to establish whether the data describe the situation accurately.

What COMPAS actually tells us

COMPAS is a criminal-justice risk-assessment system, rather than a generative chatbot. Its controversy nevertheless illustrates a problem relevant to firms buying decision-support software: different fairness tests can produce different answers.

Critics highlighted higher false-positive rates for Black defendants. Its developer, Northpointe, pointed to predictive parity: similar recorded reoffending rates among people classified as higher-risk in each group. One measure asks how often a higher-risk classification proves correct. The other asks who is wrongly classified as higher-risk among people without the recorded outcome. [2]

Alexandra Chouldechova’s 2017 analysis explained why these results can coexist. For imperfect predictions, where observed outcome rates differ between groups, equal predictive value and equal false-positive and false-negative rates cannot generally all be achieved together. This is a constraint on particular statistical criteria, not a finding that every attempt to improve fairness must reduce accuracy. [2]

For a purchasing partner, the implication is practical. A supplier’s statement that a system has been “tested for bias” needs a second sentence: tested against which standard, on whose data, and with what remaining disparities?

The legal response to COMPAS also deserves careful reading. In State v. Loomis, the Wisconsin Supreme Court permitted consideration of COMPAS at sentencing subject to restrictions. It did not give the tool unrestricted approval. The court stated that risk scores could not determine whether someone was incarcerated or the severity of the sentence, and could not be determinative of whether someone could be supervised safely in the community. Other factors had to independently support the sentence. [1]

The decision also required cautions addressing limitations, including the proprietary nature of the assessment and its reliance on group-level information. It should not be read as a general endorsement of opaque AI systems. [1]

A firm should make an equally explicit distinction between allowing a tool into a workflow and authorising reliance on a particular output. Approval to organise enquiries does not establish that the tool is suitable for declining them. Approval to extract contractual provisions does not establish that it can advise on their enforceability.

Follow the answer back to its source

Data quality is easier to discuss when it becomes a question about a particular recommendation: what information produced this result?

For a generative legal-research tool, distinguish the material used to train the model from the sources it retrieves when answering a question. A system that searches a legal database still has to select appropriate material and use it correctly. The presence of a genuine authority in its answer does not establish that the authority supports the proposition attributed to it. [5]

A study published in the Journal of Empirical Legal Studies in 2025 examined precisely this distinction. Magesh and colleagues identified errors in which legal-research outputs cited real sources that did not properly support their answers. Their analysis also distinguishes reflecting a source accurately from identifying the law that actually governs the question. [5]

The review therefore needs to go beyond opening the link. Does the passage support the proposition? Does the authority apply in the relevant jurisdiction? Has the answer omitted a qualification that changes its meaning?

For a tool working with the firm’s own records, pursue the same question through the underlying data. In the hypothetical payment-risk example, a reviewer should be able to inspect the relevant invoices, receipts, adjustments and dates. A confident explanation of “payment behaviour” is not much help if the tool was working from an incomplete ledger.

A useful procurement request is to ask the supplier to demonstrate that process using one realistic example. Start with an output and trace the supporting records, transformations and assumptions. Establish what the reviewer can actually inspect, rather than accepting a general assurance that the system is explainable.

Commercial confidentiality need not be answered with a demand to publish every training record. Depending on the proposed use, technical documentation, controlled expert access or a suitably scoped independent audit may provide evidence without exposing confidential material.

But there still has to be enough evidence to assess the intended reliance. Where important limitations cannot be examined, the firm should narrow the permitted use—or decline that use altogether.

Human review needs time, evidence and authority

Professional guidance makes clear that introducing AI does not transfer responsibility for legal work to the supplier.

The ABA’s Formal Opinion 512, issued in July 2024, addresses lawyers’ use of generative AI under the Model Rules. It recognises that the appropriate degree of independent verification depends on the tool and task; it does not prescribe manually repeating every automated step. It also makes clear that lawyers cannot relinquish their professional judgement to the system. Applicable state rules and guidance still require separate attention. [6]

In England and Wales, the SRA’s warning notice of 17 August 2026 addresses misuse of AI, including inadequate verification, confidentiality risks and failures of supervision and governance. These are responsibilities of the regulated firms and individuals using the technology. [6]

The difficult part is designing a review process that can discharge those responsibilities.

Imagine assigning a junior lawyer to approve an AI-generated recommendation. They can see the recommendation but not the documents behind it. They are expected to clear a large queue quickly, and departures from the recommendation require a partner’s approval.

There is a human in that process. Whether there is effective scrutiny is a separate question.

NIST’s generative-AI risk guidance identifies automation bias: excessive deference to an automated system, including an unjustified assumption that its output is better than information from other sources. A review process has to allow for that risk. [7]

For a consequential decision, give the reviewer access to the relevant evidence, enough time to examine it and authority to stop or change the proposed action. Where a recommendation is disputed, preserve the reason for the final decision. Another generated explanation should not substitute for examining the source.

This does not require assuming that unaided lawyers are infallible. A sensible evaluation compares the actual human-and-tool workflow with the firm’s actual existing process. It asks whether the combined process produces better decisions, including whether reviewers catch the errors that matter.

A low override rate alone cannot answer that question. It could indicate excellent recommendations. In the hypothetical queue above, it could also indicate that disagreement has been made impractical.

Give the tool a specific job—and a review date

A firm does not need to resolve every question in AI ethics before trying a tool. It does need a bounded proposal that somebody can approve, test and reconsider.

For the intake example, an initial approval record might look like this:

Decision to documentIllustrative starting position
PurposeFlag enquiries that may need earlier lawyer assessment.
LimitsNo automatic rejection or assessment of legal merits communicated to the prospective client. Existing review and deadline safeguards remain in place.
TestingCompare the tool’s recommendations with independent lawyer assessments on enquiries it was not developed or tuned against.
Review coverageExamine both prioritised and deprioritised enquiries. Investigate disagreements rather than assuming either the tool or the first reviewer is correct.
ResponsibilityName the supervising lawyer, who can change a recommendation and who can suspend use.
ReconsiderationRecord the version tested, material limitations and changes that require further testing.

This is an example of a limited pilot, not a complete compliance checklist.

Choose test material that reflects the work the firm actually receives. For intake, that might include brief written enquiries, telephone transcripts, incomplete accounts and requests from people communicating in a second language. A test consisting only of polished submissions will leave important questions unanswered.

Where decisions may affect protected groups differently, assess how that risk can be examined lawfully and responsibly. Do not infer that a small or unrepresentative sample establishes fairness, and do not collect sensitive personal information casually in the name of auditing.

Pay particular attention to negative results. In an intake pilot, independent review of deprioritised enquiries can reveal matters that deserved further assessment. It cannot establish the eventual outcome of litigation the firm never conducted. Keep that distinction in the report.

Agree the response to serious failures before the pilot starts. Missing an urgent issue, using the wrong source records or leaving a recommendation impossible to reconstruct may justify suspending the affected use while it is investigated. A promising average score should not prevent that decision.

Continue the evaluation after launch. NIST’s framework treats monitoring and reassessment as continuing activities, rather than treating initial testing as permanent assurance. Changes to a model, its data sources or the way staff use it may change the risk. [7]

The reporting should preserve those distinctions. Track time saved, but also decisions changed on review, material errors found, unresolved disagreements and the work sampled outside the system’s preferred results.

Look at the work that disappeared

Return to the partners reviewing their intake report.

They may still decide that the tool is useful. It could direct urgent enquiries to the right lawyer sooner or spare staff repetitive administrative work. Those are worthwhile benefits to test.

Before expanding its authority, however, they should ask to see an independent review of the enquiries it deprioritised. Which deserved more attention? What information did the tool miss? Were those mistakes isolated, or did they recur in a particular type of enquiry?

Those questions connect fairness to something a firm can investigate and improve.

Hours saved belong in the report. So does the work the system decided nobody needed to see.

References

  1. Wisconsin Supreme Court. State v. Loomis, 2016 WI 68, particularly paragraphs 98–100. The decision’s restrictions and cautions are more useful than a general description of it as approving algorithmic sentencing.
  2. Alexandra Chouldechova. “Fair prediction with disparate impact: A study of bias in recidivism prediction instruments” (2017). Analysis of predictive parity, error-rate disparities and the COMPAS debate.
  3. Solon Barocas, Moritz Hardt and Arvind Narayanan. Fairness and Machine Learning (2023), particularly the chapter on datasets and measurement.
  4. Jeremias Adams-Prassl, Reuben Binns and Aislinn Kelly-Lyth. “Directly Discriminatory Algorithms,” Modern Law Review 86(1), 144–175 (2023).
  5. Varun Magesh and colleagues. “Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools,” Journal of Empirical Legal Studies 22, 216–242 (2025).
  6. American Bar Association. Formal Opinion 512, “Generative Artificial Intelligence Tools,” 29 July 2024. Solicitors Regulation Authority. “Misuse of AI—Warning notice,” 17 August 2026. These address different professional regimes and should be read accordingly.
  7. NIST. AI Risk Management Framework 1.0 (2023) and Generative Artificial Intelligence Profile, NIST AI 600-1 (2024). Practical frameworks for evaluation, monitoring and human–AI interaction risks.

Source note: The supplied Module 2 material provides the conceptual starting point, particularly its treatment of measured accuracy, outcome definitions and competing fairness standards. The law-firm scenarios and suggested approval record are illustrative applications developed for this article. External references were checked on 29 September 2026.