What my doubt cost me

What my doubt cost me

Posted on: 26 August 2026

The Horizon system went into Post Office branches in 1999 and began producing shortfalls that were not there. Fujitsu built it, the Post Office prosecuted on the strength of it and five hundred and fifty-five subpostmasters eventually took the matter to the High Court, where in December 2019 Mr Justice Fraser found the software had contained bugs, errors and defects capable of producing exactly the discrepancies people had been convicted over. The supplier's reputational exposure was real. It also arrived roughly twenty years late, spread across a portfolio of public contracts, and it never at any point sat in the same room as the person whose name was on the branch accounts. That asymmetry is the whole of what people mean when they say a machine has nothing at stake. The observation is correct and it does not take you very far, because the interesting question is not who carries the loss but what you can actually do about carrying it.

The obvious answer is to verify. If the exposure is mine then I check the figures, the dates, the sources, the load-bearing steps in an argument, and I check them with some discipline because I learn something in the process. That is the part everyone is happy to describe. It is also the part where I have discovered I had built a confidence I could not support.

I got there working on a diagnostic framework, in a project that has been running for months. Twenty exchanges or so, several changes of analytical direction, each one checked as it happened. Every step held. Every objection produced an answer that stood up. Then I applied the thing to an actual case and the case took it apart in an afternoon, because the foundation assumed something that was not true and everything built on top of it was perfectly coherent with a false premise. Twenty times.

The error rate is beside the point. Continuous verification checks the step and not the direction, which is to say it certifies local coherence, and local coherence is precisely the property that a foundational error carries through intact. A language model builds each sentence as a plausible continuation of the one before, so a model that begins from a wrong premise goes on being plausible all the way to the end. There is no moment where the reasoning stumbles and tells you it has stumbled. Anyone who has supervised a junior analyst knows that human error correlates with the difficulty of the task, so the easy things come back right, the hard ones come back visibly strained, and after a few weeks you know where to look. That correlation is absent here, which is why the cost of checking never falls: you never learn where to look, so you look everywhere, forever.

A word on the word. "Hallucination" is a poor metaphor for any of this, because it suggests an isolated perceptual episode, a departure from otherwise healthy functioning. There is no departure. There is normal functioning producing a wrong result without any change of register, which is a good deal more unsettling and a good deal harder to put in a press release. Whoever settled on the term had a different problem in mind.

Back to the framework, where what interests me is not the error but the manner in which I let it run. There was a signal. In those twenty exchanges at least two positions cut against things I know, against things I have watched happen in places where I worked, and I decided to entertain them rather than dismiss them. Not through inattention. I chose it deliberately, in order to avoid the most embarrassing posture available to a man of sixty with four decades of trade behind him, the expert who discards anything that does not suit him on the grounds that he studied and the machine did not. I wanted to stay open, which is what one is supposed to do. The result was that a defence against a real bias switched off the only detector I had.

This is the mechanism worth looking at, because it does not resolve cleanly. Suspending your own judgement in the face of an analysis that contradicts what you know is correct when the machine is seeing something you cannot see, and it is fatal when the machine is saying something foolish with the same fluency it would bring to something true. The two cases present identically. No rule distinguishes them in advance, because the only difference lies in the content and the content is exactly what you have just decided not to trust yourself about.

There is something disagreeable in this for anyone who has read enough about cognitive bias, which is that awareness of a bias does not correct it. It opens a hole somewhere else. The decision-maker who knows he over-trusts his own experience builds a procedure to neutralise the tendency, and the procedure becomes the new blind spot. Boards have been demonstrating this for years, where the practices introduced to counter groupthink have tended to produce a groupthink that is better argued and considerably better minuted.

A practical consequence follows about what can be delegated, and it is not the one usually drawn. The right question is not how good the instrument is, because the quality of the instrument makes the problem worse rather than better. A collaborator who fails clumsily costs almost nothing to supervise, since the failures announce themselves. One who fails twice in a hundred, plausibly, obliges you to check the other ninety-eight, and the cost per error found becomes unsupportable well before you notice it happening. At which point you stop, and you stop precisely when the instrument has become good enough to make you stop.

The useful dividing line is the asymmetry between doing and checking. Where verification costs a fraction of execution, delegation works even at a high error rate, because the code runs or it does not, the source exists or it does not, the figure reconciles or it does not. Where verification costs what execution costs, which is to say in judgement, in the choice of angle, in deciding whether an argument holds, delegation has negative value at any error rate whatsoever, since the checking consumes the whole of the time the delegation was meant to release.

The diagnostic framework sat squarely in the second category and I was treating it as though it belonged in the first. What saved me was not a check. The real case resolved in an afternoon what twenty exchanges of careful verification had not resolved, and there is no consoling reading of that in favour of pragmatism. The case worked because it was the only thing in the entire chain that was not participating in the conversation. It had nothing to confirm.

What I am left with, and have not sorted out, is that the only reliable check on a long piece of reasoning comes from outside, while everything we do to protect ourselves happens inside. I go on verifying step by step, partly from habit and partly because I learn, aware that the verification covers a risk other than the one that worries me.


© 2026 Rolando "Rollo" Alberti - All rights reserved
About Privacy Policy Cookie Policy