The unmeasured AI revolution in legal work

The adoption question is settled. Artificial intelligence is now embedded in the working life of legal departments and law firms alike — drafting, reviewing, monitoring, summarising. The question that remains open, and that most of the profession has quietly declined to answer, is a harder one: whether anyone is measuring what this technology actually does to the cost, the quality, and the reliability of legal work. On the available evidence, the answer is no.

That evidence is now beginning to accumulate. In June 2026, some sixty general counsel and senior legal-operations leaders convened at Harvard Law School to examine precisely this gap — not whether legal teams are using artificial intelligence, but whether they can demonstrate, with anything resembling rigour, its impact. A preliminary survey conducted among the participants produced findings that ought to unsettle any institution that retains counsel, employs counsel, or supervises those who do.

This article examines those findings, and what they demand of legal functions, particularly those operating in regulated financial environments, where the discipline of measuring effectiveness is not a novelty but an existing legal obligation.

I. The measurement deficit

The headline finding admits little softening. Fewer than five per cent of the legal departments surveyed measure artificial intelligence through systematic output metrics — the kind that would capture quality, cost, and error. Roughly half track an ad hoc mixture of usage and output indicators, assembled without method. Twenty-two per cent track nothing at all.

The pattern deepens on inspection of individual outcomes. Quality of legal outputs — the variable the profession would presumably rank above all others — is measured rigorously by nine per cent of departments; the overwhelming majority, eighty-six per cent, assess it informally or occasionally, which in practice means by impression. Time saved against baselines fares somewhat better, with a third of teams measuring it consistently, though the existence of a defensible baseline is itself doubtful in most cases. Reduction in external spend, the metric that ought to be the easiest to compute, is measured rigorously by fourteen per cent.

The most consequential figure sits at the bottom of the table: half of legal departments do not track error and accuracy rates at all, and only five per cent do so with rigour. A profession whose entire value proposition rests on being right has, in its majority, no mechanism for detecting when its new tooling is wrong.

II. What the vendors do not tell

If departments are not measuring, one might expect the technology providers to fill the gap. They do not, or rather, they supply the data that flatters and withhold the data that matters.

Vendors provide usage and adoption statistics readily enough: seventy per cent of departments receive at least some. But the figures thin out precisely where they become decision-relevant. Roughly two-thirds of departments receive nothing on time savings or cost per use. Seventy-two per cent receive nothing on estimated return on investment. Eighty-one per cent receive nothing on error and accuracy rates, and eighty-two per cent nothing on the quality of legal outputs. The asymmetry is not accidental. Usage data demonstrates that a product is being consumed; impact data would demonstrate whether it is worth consuming. Only one of those propositions is commercially safe for the seller.

The practical consequence is that legal departments negotiating renewals, expansions, or new procurement are doing so on evidence supplied almost entirely by the counterparty, covering almost exclusively the dimensions on which the counterparty cannot lose. Any institution that applied this standard to the selection of a correspondent bank or a screening provider would expect to answer for it.

III. The say–do gap with outside counsel

The survey's most revealing finding concerns the relationship between clients and their external law firms. Eighty-six per cent of the general counsel surveyed describe quality assurance and accuracy from outside firms as extremely or very important. Yet forty-nine per cent require their external firms to report nothing whatsoever about their use of artificial intelligence, and only thirteen per cent ask for systematic reporting across all engagements.

The obstacle, notably, is not the one usually cited. When asked what blocks measurement, respondents ranked confidentiality and data-privacy concerns last, at twenty-one per cent. The dominant answers were structural: sixty-one per cent have no baseline data against which to compare, forty-five per cent consider their own adoption too early, and forty-five per cent have no internal owner of the question. Indeed, the most common answer to who owns the measurement of artificial intelligence's impact within the legal function is that no one does.

That vacuum will not persist. Nearly half of the departments surveyed intend, within twelve months, to introduce business-outcome metrics; forty-four per cent plan return-on-investment measurement, and the same proportion intend to link lawyer productivity to artificial intelligence. The direction of travel is unmistakable: the informal tolerance now extended to external firms is a transitional condition, not a settled one. Law firms that cannot answer, with data, how they use these tools, what they save, and how they control for error, will find the question arriving in the next request for proposals rather than the next friendly conversation. It is worth noting one omission in those same roadmaps: tracking of artificial-intelligence-specific risks and incidents was the least planned measure of all, at a third of respondents — a gap that supervisors, if not clients, can be expected to notice.

IV. Governance is what makes speed safe

None of this argues for retreat. The convening's working sessions, drawing on implementations inside several major corporations, pointed consistently in the opposite direction: the departments extracting genuine value are those that have built structure around velocity, not those that have restrained it. The analogy is the motorway engineered for high speed — the absence of a speed limit is tolerable only because the road, the rules, and the drivers are held to exacting standards. Speed without that structure is not efficiency; it is exposure.

The elements of that structure recur across successful implementations.

  1. Data governance precedes tool governance: teams that fed ungoverned data into their systems learned quickly that defective inputs produce defective outputs, and the more disciplined among them responded by restricting their tools to curated, controlled sources — in at least one case through a standing content council with named owners for each automated agent. That discipline has generated new roles within legal teams, from content curation to agent oversight, a quiet restructuring of what legal-department work consists of.

  2. Human review is treated as a control, not a courtesy. Mandatory review of machine output before use, anti-fabrication rules, and structured testing appeared wherever implementations were mature. Approximately half of the departments surveyed rely on senior-lawyer review as their principal safeguard — an entirely defensible control, provided it is designed and documented as one, rather than assumed.

  3. Mandate matters. Departments operating under a clear direction from executive leadership integrated the technology more successfully, strategically and culturally, than those left to improvise. Scepticism within teams was managed not by decree but by demonstration — concrete wins, shown internally.

  4. The economics are being renegotiated in the open. Where first drafts that once consumed associate hours are produced in minutes, and contract and compliance reviews compress thousands of hours into a fraction of that, clients are moving work in-house, scrutinising where external firms genuinely add value, and shifting toward alternative fee arrangements. The division of labour between in-house teams and external counsel — and the basis on which the latter are paid — is an active negotiation, not a settled convention.

V. A familiar lesson

For readers of these pages, the shape of this problem should be recognisable. It is the distinction, drilled into every financial institution over the past decade, between technical compliance and effectiveness. A bank may hold every policy, deploy every system, and tick every procedural box, and still fail the only question that matters: does the framework actually work? Supervisors long ago stopped accepting the existence of controls as evidence of their efficacy. They demand measurement, testing, documented outcomes, and named ownership.

The legal profession now faces the same reckoning with its own tooling, and is at present failing it. Adoption is the technical-compliance stage: the tools exist, the licences are paid, the usage statistics accumulate. Effectiveness — measured quality, tracked error rates, demonstrated value, accountable ownership — is the stage the survey shows the profession has barely entered. That fewer than one in twenty departments measures systematically, and half do not track accuracy at all, is the legal function's equivalent of an institution that has never tested its own transaction monitoring.

The remedy is equally familiar. Establish baselines before the absence of baselines becomes the permanent excuse. Assign ownership, because a control without an owner is a decoration. Measure the outcomes that matter — error, quality, cost against a defensible reference point — rather than the outcomes that are convenient. Require reporting from external providers, both vendors and firms, on the dimensions that bear on the client's risk. And treat human review as a designed, documented control with defined triggers and escalation, not as an ambient assumption that someone senior will catch what the machine misses.

Artificial intelligence will restructure how legal services are delivered, who delivers them, and what the exercise of professional judgement consists of. That much is no longer in serious dispute. What remains within the profession's control is whether that restructuring happens under measurement and governance, or under the pleasant fog of unexamined efficiency. Financial institutions were not permitted to choose the fog. There is no reason their lawyers should be.

Sections of this article were generated with the assistance of AI for purposes of linguistic expression and idea formulation.
Next
Next

AML Compliance in Lebanon (2024 - 2026)