What Should Companies Measure to Know If AI Is Working?

Most companies measure what AI helps them produce: adoption rates, hours saved, output volume, return on the licence. The organizations that pull ahead measure something else as well — what their people can still do without it. That second number is the one that predicts whether an AI investment compounds or quietly hollows out the capability it was meant to extend.

Almost every measurement problem in enterprise AI is treated as a problem of finding better metrics for the tool. It is not. It is a problem of knowing what the organization would look like if the tool were removed — and most measurement systems are built in a way that makes that question unanswerable.

What are companies actually measuring today?

Assisted output, almost without exception. Seats deployed, prompts run, tickets closed, documents produced, hours reclaimed, and — in the more disciplined organizations — financial return against spend.

This is not a soft list. Tying AI programs to margin and revenue rather than to usage counts is a genuine improvement over adoption dashboards, and the firms doing it are ahead of the firms that are not. But every one of those numbers describes the same object: the combined output of a person and a system working together. None of them can tell you how much of that combination is yours.

Why does measuring only assisted output corrupt the signal?

Because assisted output is capability plus compensation, and from the outside the two are indistinguishable.

A team whose underlying judgment is improving and a team whose underlying judgment is eroding can produce the same rising output curve, because the tool absorbs the difference. The number goes up in both cases. The organization reads the same signal and reaches opposite realities.

Everything downstream then runs on that signal. Which functions to scale. Which roles to stop hiring for. Which processes to automate next. Which verification steps are now redundant. These are the decisions that compound, and they are exactly the decisions being made on a figure that cannot distinguish “we got better” from “the system is covering.” This is the mechanism behind asymmetric adaptation: the divergence between firms is not created by the tool, which everyone has, but by whether the organization can still see itself clearly enough to decide well about it.

What does the research show about capability and compensation?

Two lines of evidence, neither of them speculative.

The first is direct. In a prospective observational study of 180 intensive-care physicians and nurses working with an AI agent in a simulated clinical setting, Bienefeld and colleagues found that drawing on the AI agent was positively associated with generating new hypotheses and with speaking up — but only in higher-performing teams. The same agent, the same task environment, and an effect that reversed depending on the state of the team it entered. The tool was not the variable. The team was.

The second is theoretical and older. Organizations run on what Wegner called a transactive memory system: a working directory of who knows what, which lets people rely on each other without duplicating expertise. Argote and Ren argued that this directory is a microfoundation of dynamic capability — the thing that lets a firm reconfigure itself when conditions change. An AI system enters that directory as a new node. It is consulted, weighted, and relied upon like any other. But it is a node whose competence cannot be inspected, and its arrival changes how often the human nodes get consulted at all.

An organization that no longer knows which competence is human and which is borrowed has not lost a metric. It has lost the directory that adaptation runs on.

One honest limitation: the clinical study is observational and simulated. It establishes that the effect is conditional. It does not establish how often, or in which industries.

What does it look like when a company actually measures this?

Concretely, and it is not a maturity assessment:

  • A periodic sample of unassisted work on live problems, not test scenarios — a real brief, a real analysis, produced without tooling
  • Reviewed for judgment, not speed — what was noticed, what was questioned, what was rejected and why
  • Held at team or function level, never as individual appraisal. The moment it becomes a performance review, people optimize for the measurement and the signal dies
  • Attached to a stopping rule. If unassisted quality falls below a defined floor, expansion in that function pauses until it recovers. A measurement with no decision attached is theatre
  • Owned by the function that does the work, not by HR and not by the technology group

This is a practice, not a detection method. The separate question of how to tell whether erosion is already underway in a specific team has its own instruments and its own thresholds.

Stated as a proposition, so that it can be wrong: organizations that hold a current measurement of unassisted capability adapt better to AI than comparable organizations that do not, independent of spend, adoption speed, and talent density. If a large share of high-performing adopters turn out to hold no such measurement, the claim fails. That test is worth running rather than assuming, and it is the next thing this research program will report.

Why is this harder in some markets than others?

The constraint differs by region, and it is rarely technical.

In the United States, the binding constraint is velocity: measurement competes with scale, and because scaling decisions compound fastest there, a corrupted signal also becomes expensive fastest. In the Gulf, AI programs are frequently board-mandated and highly visible, which makes an unassisted baseline that lands below expectation politically costly to produce — the difficulty is reputational, not methodological. In Central and Eastern Europe, AI often substitutes for capacity that was already thin after years of cost discipline, so the measurement tends to reveal a dependency with no fallback behind it. Different obstacle, same avoidance.

Frequently asked questions

Isn’t measuring unassisted work a step backwards?

No, because nothing is being taken away. The measurement is a control, not an operating policy. Aviation and medicine both retain unassisted proficiency checks without abandoning their instruments; the check exists precisely so that reliance can be extended safely.

Doesn’t this just mean testing employees?

It should not, and it fails if it does. The unit of analysis is the team or the function. Individual scoring converts a diagnostic into an appraisal, and people will optimize against it within one cycle.

How often does this need to happen?

Frequency should scale with the cost of verifying the work. Where errors surface immediately and cheaply, rarely. Where they surface late, expensively, or not at all, considerably more often.

Is there evidence that this actually predicts advantage?

Not yet at organizational scale — and that distinction matters. The supporting evidence is conditional and correlational: the effect of AI depends on the state of the team receiving it. The stronger claim, that the measurement itself predicts adaptive advantage, is a proposition currently being tested, and it will be reported with the data rather than asserted ahead of it.