The Sledgehammer Problem
Nobody keeps one tool in the shed. There’s a hammer for the picture hooks, a mallet for the tent pegs, and somewhere at the back, a sledgehammer for the one job a year that actually needs it.
As tempting as it is to always reach for the sledgehammer first, experience tells you that different jobs require different levels of force.
Maybe we just haven’t built that experience with AI yet. That might explain why we reach for the most powerful, most expensive model by default - constantly, and mostly without noticing - without stopping to ask whether we’re about to crack the nut or smash it.
Last week’s question wasn’t whether anything was worth building - it was whether we’d quietly assumed away the need to keep paying attention to it, just because building got cheap and fast. This week, the same trap shows up one layer down. It isn’t just applications we’ve stopped paying deliberate attention to. It’s the model doing the work underneath them.
I watched a finance team route every incoming supplier email - most of them a one-line status update - through the same frontier model they used for their hardest contract negotiations. It worked. It also cost roughly the same to process “delivery confirmed, thanks” as it did to interrogate a genuinely ambiguous clause. Nobody had decided to do it that way. It’s just what “using AI” had quietly come to mean by default.
That’s the sledgehammer problem. It isn’t laziness, exactly - it’s habit. Reaching for the biggest, most capable tool because it’s a single click away, rather than asking what the job in front of you actually needs.
Most of what crosses an organisation’s desk in a day - sorting, extracting, summarising, routing - doesn’t need frontier reasoning, and doesn’t need frontier caution either. A small, self-hosted, open-weight model handles it for a fraction of the cost, run entirely in-house. The genuinely hard stuff - the ambiguous, the high-stakes, the judgement calls - still deserves the expensive model, because that’s where the expense earns its keep.
Two questions decide which is which: how hard is the task, and how sensitive is the data behind it. Plot any piece of work against those two and the routing mostly does itself. Easy and low-stakes, handle it locally and cheaply. Hard and sensitive, that’s where the frontier model and the extra governance both belong. Most work sits somewhere in between, and that’s fine - the point isn’t a perfect four-box diagram, it’s to stop defaulting to the top corner for everything.
There’s a second decision organisations need to make alongside this one, and it might matter just as much: don’t build your strategy around one model, the way you once built it around one operating system or one cloud provider.
I don’t think this market’s reached that point, and I’m not sure it will for a while. Office may or may not be the best collaboration suite in the world. It doesn’t need to be. It’s good enough that everyone can agree to use it and the IT guys can get back to the conversations that actually move the business forward. That’s what the operating system did too, and the cloud platform after it - not so much lock us in as let us out. Once the plumbing cleared the bar, we could stop fettling with it and put our attention back where it belonged.
The model layer hasn’t reached that point. “Good enough” needs the field to slow down long enough for consensus to catch up with it, and right now it isn’t slowing down. What was frontier six months ago is mid-tier today. What’s expensive this quarter is commodity next year. There’s no plateau to converge on yet - which means the moment that let organisations stop arguing about which operating system or cloud provider they were going to double down on hasn’t arrived here. Not yet, not for a while and potentially not ever.
Add a third kind of risk to sit alongside cost and performance: availability isn’t just commercial, it’s regulatory. A model can become unusable overnight for reasons that have nothing to do with your contract, your data, or how well it did the job - a national jurisdiction, not a market, deciding you no longer get to use it. That’s not a hypothetical. It’s already happened this year, more than once, to more than one provider. An architecture that assumes uninterrupted access to any single model isn’t just a cost risk. It’s a continuity risk.
So don’t pick a winner. Not to avoid getting locked in - there’s no “good enough” consensus yet to let you stop paying attention to the choice, and committing early just means defending a decision the market’s about to overtake anyway. Build the architecture so you can route work to whichever model does the job best today, and swap it out without rebuilding everything when a better or cheaper option turns up next month or next year.
Put the two decisions together and the real discipline isn’t just matching model to task. It’s building an AI architecture flexible enough to keep doing that as the ground underneath it keeps moving - because it will. Every task over-served by a model you’re locked into is money and attention not spent on the work that actually needed it.
What’s the sledgehammer doing in your organisation right now - and how tied are you to the hand holding it?