In the sixth and final article of a series on safely adopting AI, Ed Molyneux turns from where data goes to whether the answer is right, and sets out the evidence a firm should demand before it relies on an AI’s output.
Every article in this series has been about where data goes and whether a firm can prove it: accountability, access control, an audit trail that writes itself. All of it matters, but none of it touches the question a firm relying on AI still has to answer separately. Is the output actually correct?
They are different questions. A system can be immaculately governed, every access logged and every permission enforced, and still be confidently, fluently wrong. Governance tells you the AI saw only what it should and left a record of it. It says nothing about whether what came back is right. Conflating the two is the most consequential mistake in the whole subject, because a well-governed wrong answer is more dangerous than an ungoverned one. It arrives wrapped in the credibility of a controlled process.
It is also the point the SRA’s warning notice hits hardest, in saying that reliance on an AI output would not be a suitable defence.
Start with why sounding right is worth nothing. Fluency reads as competence. The output is well written, appropriately hedged and structurally sensible, but none of that is evidence that it is correct. These models are trained to produce plausible text, and plausibility is precisely what makes a wrong answer hard to catch. In conveyancing the danger is specific, because a model is often most confident, and most quietly wrong, on exactly the details that decide the outcome: a rate that changed at the last Budget, a lender’s current requirement, whether a covenant still bites.
A fee-earner reading a clean, confident paragraph cannot tell a researched answer from an invented one. That is not a failing of the reader but a property of the medium, and it is why the AI said so cannot be the basis for anything.
So rank the kinds of backing an answer can have. The weakest is that it sounds right, which is no evidence at all. Better is an answer that cites a source you can check, though only if the citation is real and says what the answer claims; a checkable authority at least lets a professional verify rather than trust. Stronger again is an answer produced by a validated tool rather than free-form generation, because a purpose-built calculator or checker tested against known-correct results does not go stale the way a model’s training does, and fails in more predictable ways. At the top is an answer measured against expert judgement, at scale, with a known error rate. Not that an expert looked at it once, but that the system was run across a large, representative set of real cases, that domain experts ruled on the hard ones, and that someone counted how often the two agreed.
Most AI offerings sit on the bottom two rungs and are described as though they were on the top one. The gap between we are confident it is accurate and here is our measured agreement with expert judgement is the gap between a claim and evidence.
One further subtlety separates real accuracy evidence from a flattering headline. A single overall accuracy figure can hide exactly the failures that hurt: if a system is nearly always right about the routine and quietly unreliable on the rare, serious issue, the continuing covenant breach, the un-consented alteration, the lease term that fails a lender, then a high average is worse than useless, because it manufactures confidence in precisely the place caution was needed.
The honest measure is per severity. How well does the system agree with expert judgement on the high-consequence minority, rather than across the easy majority? Ask for that breakdown and treat its absence as an answer in itself. Measured how, and against whose judgement, are the follow-ups, and the willingness to answer plainly is the tell.
Two things follow. The first is to match reliance to evidence: an answer that merely sounds right is a draft to be checked, a validated or source-cited answer can carry more weight, and only measured, per-severity accuracy justifies leaning on a system for the serious calls. Calibrating how much to trust which output is the actual skill of using AI well.
The second is that the professional stays responsible. The first article named a duty AI cannot be allowed to erode: judgement that stays human, which is the point the SRA’s notice underlines. Accuracy evidence does not remove that duty, it informs it. Knowing a system’s measured error rate, and where the errors concentrate, is what lets a professional decide how far to rely and where to look harder. It is an input to judgement, never a substitute, and a firm that treats a good accuracy number as permission to stop thinking has misunderstood what the number is for.
Which is why accuracy and governance belong in the same conversation. Governance lets a firm prove what its AI did; accuracy evidence lets it know how far to trust what the AI said. A firm needs both, from any tool it adopts and from itself, because the client, the regulator and the insurer will each ask about one sooner or later. The firms that can answer with evidence rather than assurance are the ones for which AI becomes something to stand behind rather than quietly hope about.
About the author
Ed Molyneux is co-founder and CTO of Moverly and the original author of the Property Data Trust Framework (PDTF), the open standard for machine-readable property data now being adopted across the industry. Ed writes about AI, property data infrastructure, and the future of conveyancing.
The views expressed in this article are those of the author and not necessarily those of Today’s Conveyancer. This article is general information, not legal advice.
















