Critical Thinking
← Opinion & Analysis Opinion & Analysis

AI Strategy Is Only Precise Where You Can Be Proven Wrong

The vocabulary of transformers and agents is easy to borrow. A commitment a skeptic could falsify by Friday is not — and that line is what separates a strategy from a wish on a star.

By Igor Nesterenko · 7 July 2026

A brass sextant resting on a wooden case beneath a single distant star.

The strategy lead had twenty minutes and a good deck. Every board member could repeat the headline by the second slide — become an AI-first organisation, put AI into every product line, lead the sector by 2027. When I asked which sentence in the document a skeptical outsider could prove wrong by the end of the quarter, the room went quiet — not because the answer was hard, but because there wasn't one.

The deck was fluent in the vocabulary. Transformers, agents, retrieval, compute budgets, an evaluation workstream. The commitments underneath were not fluent in anything, because there were no commitments underneath — only a direction and a date.

That gap, between sounding technical and being falsifiable, is where most AI strategy lives.

Precision Is Falsifiability, Not Vocabulary

I read AI strategies for a living — corporate ones, national ones, the sort that arrive as a forty-page PDF with a foreword. The first thing I look for is not ambition. It is whether any single claim in the document could turn out to be false.

That test has a name. A statement is falsifiable when reality is allowed to contradict it — when you can describe, in advance, the observation that would prove it wrong. Karl Popper used exactly this line to separate science from everything that merely sounds like it: a theory compatible with every possible observation, he argued, tells you nothing. If that phrasing is new, the version worth keeping is that a claim you cannot be wrong about carries no information. Falsifiable does not mean pessimistic or likely-to-fail. It means the world has been given a way to answer back.

Measured against that, most AI strategies are compatible with every possible outcome. "Become an AI-first organisation" is equally true whether you ship one model or a hundred, whether revenue moves or doesn't. It cannot be contradicted, so it cannot inform — a heading, not a commitment.

This is the difference between a bearing and a star. You can point at a star all night and never be wrong, because pointing commits you to nothing you could check. A bearing is different. It names a direction you can verify against the world, correct when you drift off it, and be caught getting wrong. A precise strategy sets bearings. A hopeful one points at a star and calls the gesture rigour.

The Fluency Trap — Sounding Technical Without Committing

There is a reason the vocabulary runs ahead of the commitments. The money is enormous and the fear of missing it is real. When that much capital is chasing AI, every organisation needs to be seen to have a strategy — and being seen to have one rewards fluency far more than falsifiability.

Governments sit in the same position. The OECD's policy repository now tracks more than 1,300 AI initiatives across over 80 jurisdictions; by 2023, 51 countries had filed a national AI strategy, up from a handful in 2017.

What has not kept pace is the machinery to check any of it. The same OECD review notes that monitoring-and-evaluation mechanisms — the part that would tell you whether a strategy is working — remain scarce. The destinations are documented in triplicate; the instruments to confirm arrival mostly aren't built.

So the pattern repeats at every scale. Ambition is abundant and cheap to write; verification is scarce and expensive to sign.

The strategy that sounds most technical is often the one that has committed to the least — the vocabulary quietly doing the job a checkable claim is supposed to do.

What a Precise Strategy Commits To

If unfalsifiable ambition is the problem, the answer is not more ambition. It is a commitment with coordinates. When I find a claim in a strategy that holds still long enough to be tested, it carries four of them:

The clearest worked example is not a corporate deck; it's a regulation. The EU AI Act's conformity assessment — plainly, a third party checking whether a high-risk system meets defined requirements before it ships — has all four coordinates and refuses to hide them. The metric is a specified set of obligations: risk management, data quality, activity logging, human oversight, documented robustness. The owner is a notified body. The threshold is the requirement itself. The consequence is real — the body can refuse the certificate in writing, with reasons, and a substantial change to the system forces the whole assessment again.

Agree with the Act or not, it describes a strategy you could prove wrong. That is what most "AI-first" decks are missing. Not vocabulary, and not ambition — any sentence a third party could stand in front of and say no to.

Where a "Precise" Strategy Still Drifts — a Stress Test

Suppose a strategy does set coordinates. It still drifts, and the drift is worth naming, because it's where careful teams lose the plot rather than lazy ones. Take a plausible case: a company commits to "AI-first by 2027," picks a metric, assigns an owner, sets a threshold. Three failures tend to follow it.

The first is ordinary. The metric that gets chosen is the one that's easy to move, not the one that matters. "Models shipped" and "GPU hours consumed" are precise, ownable, and thresholdable — and they measure activity, not the capability the strategy claimed. A team can hit every number and deliver nothing a customer would notice. The bearing was checkable; it just pointed at the wrong star.

The second is about incentives. A falsifiable target creates the one thing most organisations quietly avoid: a named person who can be shown to have failed.

Vagueness protects careers; precision exposes them.

So the pressure runs always toward the softer metric, the movable threshold, the consequence that never quite triggers. This isn't laziness — it's the reward structure working as designed, paying for the confident claim and taxing the checkable one.

The third is drift in the destination itself. "AI-first" means one thing the day the strategy is written and another two model generations later. The roadmap outgrows the metric, nobody re-baselines, and by the time the deadline arrives the target has slid to wherever the work happened to land.

None of this is exotic. RAND's study of why AI projects fail put the failure rate above 80% — roughly twice that of ordinary IT projects — and the leading cause was not a technical wall but leadership never clearly defining what success meant, so systems were built and optimised for the wrong thing. In the language of the four coordinates, the metric and threshold were missing before the work even started. It's a qualitative study, so read that as "the large majority fail," not a precise figure.

A separate cause was a technology-first mentality: reaching for the newest tool instead of the actual problem. Both are the disease this piece keeps circling — a strategy that names a star and never sets a bearing anyone can check.

Verifiability Before Ambition

There is a discipline that treats this as a solvable problem rather than a mood. Researchers call it technical AI governance — in the Oxford sense, the analysis and tools that close the gap between what a policy wants and what can actually be checked. The useful claim inside that phrase is that the gap is partly an engineering problem, not only a question of values: much of what we say we cannot verify, we have simply not built the means to verify yet. Oxford's own work argues that verifying serious AI commitments — even between states that don't trust each other — looks feasible within a few years if the instruments get built.

Which reframes the whole precise-or-hype question. The problem was never that AI strategies aim high; aiming high is the point. The problem is aiming at a star with no instrument on board — spending real capital and real credibility on a destination while having no way to tell, on an ordinary Tuesday, whether you're getting closer or just moving. When the star isn't reached, an unfalsifiable strategy can't even run the post-mortem, because nothing was ever measurable enough to explain what went wrong.

So before the next AI strategy ships — yours, a vendor's, a government's — find the one sentence a skeptic could prove wrong by Friday, and name its metric, its owner, its threshold, its consequence. If you can't find that sentence, you don't yet have a strategy. You have a heading and a hope.

Which line in your current AI strategy could an outsider falsify this quarter — and if the honest answer is none, which coordinate is missing first: the metric, the owner, the threshold, or the consequence?

Take the RADAR reading More analysis →