Your AI programme looks better from where you are sitting

Hero image for Your AI programme looks better from where you are sitting

Second day of the AI summit in Barcelona, in a room with maybe two hundred people in it, a governance speaker asked how many of us had the authority to stop an AI system we oversee, not merely flag a concern about it. Nobody raised a hand. Then she sharpened it: "human in the loop" has become a checkbox, and the question worth asking is which human, with what authority.

Every organization in that room would answer yes to a survey question about human oversight. The documentation exists. Someone is named in it. And in the room, with no consequences attached to putting your hand up, the real answer was different. Both facts are true at the same time, and only one of them ends up in a report that reaches a board.

And the gap between them leans the same way every time: the further you sit from the work, the better it looks.

Thirty-five points apart, same companies, same week

Valtech surveyed 1,003 professionals who already use AI at work in May 2026, about 60% of them in the EU and most of that in DACH. Among directors, 60.1% described AI as integrated across workflows, tools and teams. Among C-level respondents it was 55.9%, and across the whole workforce sampled, 25.2%. Same companies, same instrument, same week, thirty-five points apart. Valtech's own framing for this is a leadership visibility gap: leaders see pockets of integration that the wider organization does not experience.

Measured from the other end, it looks like impatience. In the ServiceNow and ThoughtLab maturity index for 2026, 57% of US employees said leadership is not keeping up with fast-changing trends. In Japan, six in ten employees said leadership was "not up to the task." I would not put much weight on either number as a fact about leadership. They are useful as evidence that the two groups are describing different objects and calling them by the same name.

Two opinions disagreeing is still two opinions. What makes the direction interesting is that the more senior source is also the more confident one about things it can verify least.

Deloitte's 2026 Global Technology Leadership Study surveyed 662 technology leaders, 87% of them C-suite. 81% said they were confident in their organization's ability to deploy and govern AI capabilities at scale. In the same study, 75% agreed their operating models and processes have to change within twelve to eighteen months to get more value out of AI. Deloitte titled the slide "Confidence is ahead of readiness." When a Big Four firm puts that label on its own respondents, it is reporting what those respondents said about themselves: in one answer, a capability they are confident about; in the next, the machinery underneath it that they know does not work yet.

Nobody is lying in any of this. The mechanism is simpler than dishonesty.

The answer depends on which question you ask

Deloitte's State of AI in the Enterprise, January 2026: 42% rate their organization highly prepared on strategy. Ask the same population about risk and governance and it drops to 30%. McKinsey's State of Organizations 2026, with 10,018 respondents, reports that 86% of leaders feel their organizations are not very prepared to adopt AI in day-to-day operations. More than half of that sample are middle managers rather than executives, which makes the figure harder to dismiss: it is bad news from the people closest to the work.

The closer you sit to a thing, the more accurately you can rate it. A strategy document sits on your desk: you wrote it or approved it, so you can judge it honestly. Day-to-day operations are hundreds of small handoffs happening three floors down, and nobody senior is watching them, because that is what seniority means.

Cloudera's Data Readiness Index, 1,270 IT leaders, has the cleanest version of this. 85% say their data strategy is clearly defined and tied to business objectives. 84% are confident in the accuracy and completeness of their data. Then: 18% say their data is fully governed, and 30% say their sources are fully integrated. The high numbers describe the plan, the low numbers describe the plumbing, and both sets came from the same person in the same sitting.

I had lunch in Barcelona with a former professor of mine, now running regional IT for Western Europe at a multinational in healthcare. He told me two things an hour apart without connecting them. First, his company uses very little AI and moves slowly, although he sees it as a structural change. Second, that inside a large organization, even when he is certain his proposal is right, he sometimes has to take the long way round: ask instead of tell, let someone else own the idea, and accept that being right early is worth nothing on its own. Nobody surveys that. If you asked his organization about AI strategy maturity, you would get an answer about the strategy. The bottleneck: convincing costs more than building, and the person who knows what to do spends his week on the long route.

The one time somebody used a clock

METR ran a randomized controlled trial with 16 experienced open-source developers on 246 real tasks in repositories they had worked in for about five years on average, between February and June 2025. Before starting, the developers forecast a 24% speedup. After finishing the work, having done it themselves, they estimated a 20% speedup. The measured result was a 19% slowdown.

The study has not replicated cleanly since, and I would not build a policy on the sign of that number. What survives is the spread inside those 16 people: perception and clock pointed in opposite directions, and the second estimate came after first-hand experience. Doing the work did not correct the error. If domain experts cannot self-report a speedup on their own keyboards, a director three levels up is not going to produce a reliable integration figure for a company of four thousand.

Survey against logs, and what your AI programme actually reports

Every figure so far comes from asking people. That is the limit of all of them, including the thirty-five-point gap this article opened with.

Survey against survey gives you a disagreement. Survey against logs gives you a measurement.

The comparison is available in almost every organization and almost nobody runs it. Compare seats assigned against seats with activity in the last thirty days. Count the tools that were opened once and never again. Then take the workflows a steering committee listed as automated and count how many tickets still arrive for them through the manual channel. Anthropic's own December 2025 study on how AI is changing work states plainly that it cannot determine where reported time savings were reinvested — and that is the vendor with the best telemetry in the market declining to claim it.

The size of the gap between the reported number and the logged number is itself the finding. It tells you how far your reporting line has drifted from the work, it is specific to your company rather than to a benchmark population, and it is the one measure you can compute about yourself without asking anyone a question.

I argued in an earlier piece that the EU AI Act assumes a logging and oversight floor most companies never built. This is the companion problem: from the top of the organization, you have no way to see whether any of that logging exists, because the instrument reporting on it was assembled by people who have never watched a single ticket move.

Before signing the next round of AI investment, find out which instrument produced the case for the last one. If it was a survey, a maturity self-rating or a slide in a steering committee, the number was measured in the wrong place. Replace it with one observed number. Take a request that actually arrived last week (a customer onboarding, a quote, a support escalation) and follow it end to end with timestamps: when it landed, when each person touched it, when it left. Compare the elapsed clock time to what the slide says it takes. Two weeks of that gives you a baseline nobody in the building can argue with, including you.