The big fear I'm hearing from chief executives isn't whether AI will get things wrong, or whether it'll cost too much. It's that their people will keep getting the right answers without any idea how those answers were arrived at.
One put it to me plainly. He's already watched it happen with ordinary digital tools - the ones nobody thinks twice about. When a system goes down, the business stops...
Not because the work is impossible by hand, but because there aren’t that many people left who remember how it used to be done before all this new-fangled technology stuff turned up. And if that wasn’t bad enough already, now he's looking at technology that does judgement rather than processing, and he’s asking the obvious next question. If the answers stop arriving, does anyone here still know how to reach the right conclusion?
Forty years ago I wasn't allowed a calculator in O-level maths, on the grounds that I wouldn't always have one to hand. On that, my teachers were spectacularly wrong. It turns out the trick was never whether we've got a calculator in our hands - it's whether we can put the bloody thing down for five minutes.
But they were also chasing the wrong skill, and it took me decades to see which one. All that grinding through complicated long division sums trained me to follow a process - precisely the thing the machine now does better than I ever will. What nobody trained was the other thing: the ability to glance at an answer somebody else produced and think hang on, that can't be right. Not producing the number, estimating it. Knowing the area of a circle with a 10cm radius is comfortably north of 300, so that when something tells you 31.4 you stop.
Which is precisely what my chief executive is worried about. Not the sums. The sense-check.
But notice what actually got handed over there, because it's the same in every example I could give you. The calculator does the arithmetic. It has never once decided which sum was worth doing, or what to do with the answer once it arrived. Same with the spellchecker, the spreadsheet, the sat-nav. Every one of them took the what. You and I kept the why and the how.
Until now.
These days my agents advise me on my tax returns. They help me decide what to cook for my family. I use them on investment decisions and - most absurdly - to spec the computer hardware that will eventually become their own physical form, which makes me a sort of digital Dr Frankenstein with a comparable electricity bill.
And in every one of those cases, the checking changed shape. Not less of it - but different, depending on what rides on the answer. Deciding what's for dinner gets a glance: does that look about right, could that reasonably be true? Anything I'm going to stand behind in public, or hand to a client, or build an argument on, gets taken apart properly. Sources followed to source. Numbers recalculated. The whole thing pulled through by hand.
Which sounds like a decent system, and mostly is. The trouble is what happens to everything in between - the great mass of ordinary work that doesn't announce which category it belongs to. Because a working paper published in January describes what people do there with uncomfortable precision. Researchers at Wharton ran an experiment where participants could consult an AI. Most did, on most trials. Against a no-AI baseline, their accuracy rose 25 percentage points when the machine was right - and fell 15 points when it was wrong. Read that second number again. When the machine was wrong, people did worse than those working without it at all. Their own judgement didn't act as a brake, because it was never engaged. What the experiment captured wasn't people being helped or hindered by a tool. It was people trading their own accuracy for the machine's - and mostly getting a good deal, because the machine is mostly right. But it's the same trade either way. You're no longer as good as you are. You're as good as it is.
The distinction they draw is the one that matters. Cognitive offloading is when you delegate but stay in the driving seat - the calculator, the map app, the template. Cognitive surrender is when you adopt the machine's judgement as your own and the deliberate part of your thinking simply never engages. The loss is double: more wrong answers get through, and you lose the skill of checking.
Which is what a plausibility check catches and what it misses. A wrong answer that looks wrong gets stopped. A wrong answer that looks entirely reasonable sails straight through - and the better these tools get, the more of the second kind there are. Which is fine when it's dinner. It's rather less fine when the thing that looked reasonable was the basis of an important decision.
Daniel Kahneman spent a career showing that we don't have one thinking system, we have two. There's the fast one - instinctive, pattern-matching, the thing that lets a goalkeeper glove a puck at 98mph or lets me drive a familiar route while thinking about something else entirely. And there's the slow one, the deliberate, effortful sort of thinking you do when weighing up a decision that matters.
But there's a third layer sitting above both, and it's the one we don't talk about enough: metacognition. The supervisor part of our brain that monitors our own decisions and overrides a habit when the habit is about to let us down.
Which reframes the whole thing for me. Cognitive surrender isn't some new system arriving from outside. It's the supervisor quietly stopping turning up for work.
There's corroboration around it. Microsoft and Carnegie Mellon surveyed 319 knowledge workers and found reduced critical engagement, sharpest in routine, lower-stakes tasks where people simply deferred. And remember that finding from the London Taskforce I quoted a fortnight ago - on legal tasks, the quality improvements landed disproportionately on the lower performers, while some of the strongest saw their quality drop. That second half is surrender, showing up in exactly the people you'd least suspect.
So what do we actually do about it? For anyone running an organisation, the first move is to stop treating this as a question of virtue and start treating it as one of resilience.
And notice what the sense-check is actually for, because it's easy to reduce it to fact-checking. Spotting a wrong number matters, obviously. But the more valuable version is spotting a right answer to the wrong question - work that's accurate, well-presented, entirely plausible, and solving something nobody needed solving. You can only catch that if you understand how the answer was built, which is precisely the part we've stopped holding onto. That's the drift my chief executive is worried about, and it's the harder one to catch, because nothing about the output looks incorrect.
Which is why the question to stop asking is "are our people using AI?" (You’d better hope that they are!) Instead, start asking "are our people still thinking when they use it?" Adoption metrics can't tell those apart, which is why almost nobody is measuring the thing that matters.
Then borrow shamelessly from aviation regulation, the one industry that has already been all the way through this. Nobody wants to know whether pilots are using autopilot. They want to know whether the pilot can still fly - and then they go and check. The FAA warned back in 2013 that continuous use of autoflight systems degrades a pilot's ability to recover an aircraft from an undesired state, and simply doesn't reinforce manual skills. In Canada, recurrent training must include mandatory modules where the pilot flies manually through simulated abnormal scenarios, hand-flies an approach and landing, and is graded on it.
Nobody calls that nostalgia. Nobody argues pilots should stop using autopilot. The industry concluded that if you depend on automation you must deliberately practise without it - and then made the practice mandatory and assessed.
That's a playbook, and it's four things:
- Know which of your processes have no human fallback left.
- Build deliberate unassisted work into normal operations rather than treating it as a fire drill.
- Assess it, because unassessed practice doesn't happen.
- And create real error signals, because without them the default is acceptance.
And there's a personal version of all this, which the Wharton team put better than I could. Form your own answer first - instinct, then deliberation. Then let the machine challenge it, refine it, or tear it apart. Same tools, same speed, completely different relationship. The order is the whole thing.
I should point out that there’s a catch in that first point, and my favourite illustration of it comes from retail. There was a time when every till had a click-clack machine - a little metal contraption that took your credit card and embossed its raised numbers onto a paper slip. System down? Fine. Take the imprint, take the signature, carry on trading.
Cards don't have raised numbers anymore. Nobody voted to remove the fallback - it stopped being useful for anything else and quietly disappeared, which is how fallbacks always go. Not abolished, just designed out while nobody was looking. Some of that is fine. But it means the first job on that list is a choice rather than an audit: which fallbacks are you deliberately keeping alive, and what will you pay to keep them?
And here's the part that ought to concentrate the mind. If embossing turns out to matter, a card manufacturer can reintroduce it inside a product cycle. Human judgement doesn't work like that. It takes fifteen years to grow and you cannot procure it in a hurry - which is exactly what I was arguing about the on-ramp a fortnight ago. Of all the fallbacks quietly rotting in your organisation, judgement is the only one you can't buy back.
This causes me to revisit my oldest complaint, and one I've written about here before - that we still pay people for their time rather than what they achieve with it. This is the same argument wearing different clothes. Reward people for how fast the answer arrived and you'll get surrender. Reward them for whether it was right and you get judgement. We are still, overwhelmingly, paying for the first.
Which brings me back to where I started, and why I think that fear is the right one to have. It isn't a training problem, and it isn't a technology problem. It's a resilience problem, and it belongs on a risk register alongside every other single point of failure a chief executive is already accountable for.
What makes it different is that it degrades silently. No system goes down. No alarm sounds. The work keeps arriving, on time and looking entirely plausible - right up until the day it was solving the wrong problem and the person who would have noticed had quietly stopped looking.
I'm not giving up my agents. I have no intention of going back to doing long division by hand, or to filing my own tax return unaided. But the machines used to take the arithmetic and leave me the thinking. Now they offer to do both, and they do it well enough that I don't always notice which one I've handed over.
So the job is to keep asking. Not can it do this - but did I stop thinking when it did?