How Do You Know If AI Is Eroding Your Team’s Expertise?

You measure it directly: unassisted performance on a fixed task, at a fixed interval, in the few places where it matters. No indirect signal will tell you, because every metric you currently review measures the person and the tool as one system, and the tool holds the number up while the capability underneath declines. But a narrower question comes first. Most capability erosion is harmless and always has been. Only one class of it is worth detecting at all.

Organizations have been trading skills for tools for two centuries and were usually right to. Nobody mourns mental arithmetic, celestial navigation, or the ability to draft by hand. So “AI is eroding expertise” is not by itself a finding — it is a description of what adoption does. The executive question is which erosions are recoverable and which are not.

Why doesn’t expertise erosion show up in performance data?

Because performance data measures assisted output, and assisted output is exactly where erosion hides. A team using AI produces work of comparable or better quality while its underlying capability declines; the tool absorbs the difference. Cycle time, throughput, error rate, client satisfaction — every one of these is measured on the combined human-plus-tool system, and none decomposes it.

This is the fourth and slowest channel of the Missing Down-Slope, and the only one that conceals itself by operating. Verification, correction, and rework are costs of the current work and eventually show up somewhere. Capability decay is a cost imposed on future work, and it degrades the very controls that catch the other three.

The practical consequence is that the absence of a warning sign is not evidence of health. It is evidence that nothing is watching.

Which capabilities actually need protecting?

Very few — and identifying them is the first decision, before any measurement. Two tests separate the cases.

The first is the cost of verification. Where checking AI output is cheap, decay is inconsequential: errors surface on their own. Arithmetic skill decayed with the calculator and nothing was lost, because an arithmetic error is cheap to detect. Where checking is expensive — a valuation assumption, a legal position, a clinical judgment, a strategic recommendation — the capacity to check is the control. Its decay does not just reduce skill; it removes the mechanism that catches the tool’s errors.

The second is repurchasability. Can the capability be rehired at acceptable cost and speed? Generic professional expertise usually can. Firm-specific tacit knowledge cannot: how this client’s business actually works, why a legacy restructuring left the ledger in its current shape, which supplier assurances have historically been reliable. That knowledge exists on no market at any price, and it is precisely the knowledge that was previously acquired by doing the routine work AI now performs.

Cross the two tests and only one class demands action: capabilities where verification is expensive and the underlying knowledge is firm-specific. Where verification is cheap, let the skill go. Where verification is expensive but the expertise is generic, the exposure is a procurement problem with a known price — manageable, and not a reason to slow adoption.

Regional structure changes the second test more than executives expect. In the United States, a liquid senior market means most capability is repurchasable, which makes erosion survivable and therefore easy to ignore until the firm-specific layer is the part that has gone. In the Gulf, rapid growth and heavy reliance on experienced expatriate hiring produce the same masking effect more strongly: an organization can appear to be deepening expertise while its internally grown, context-specific capability thins, because the depth is bought rather than built. In Central and Eastern Europe, where cost discipline keeps benches thin by design and the external senior market is shallower, the same erosion converts into operational exposure considerably faster.

How fast does capability decline once it starts?

Faster than seniority intuitions suggest, and tenure offers less protection than most succession plans assume.

The clearest evidence comes from clinical practice, where a 2025 multicenter study measured detection performance among endoscopists before and after routine AI assistance was introduced. Their performance on unassisted procedures fell significantly within months. The detail that matters for targeting is not the magnitude but the population: every participant had performed more than two thousand procedures. Deep experience did not insulate them.

That finding invalidates the proxy most organizations rely on. Tenure, seniority, and past performance ratings all describe capability as it was demonstrated under prior conditions. None of them tracks capability under current ones. If experienced practitioners can decline measurably within a quarter, the assumption that senior staff constitute a stable capability reserve is an assumption, not an observation.

The study is observational and drawn from one specialty; commentators noted plausible confounders, and no single clinical setting generalizes automatically to legal, financial, or engineering work. But the direction and the clock are the operative facts, and both point the same way.

What do you measure, and where do you install it?

Three measurements make capability visible, and the criterion above determines where to install them — in the expensive-to-verify, firm-specific work only. Applied everywhere they become an assessment bureaucracy that will be resented and quietly abandoned.

  • A periodic unassisted sample. A fixed, representative task performed without AI at a set cadence. What is assessed is judgment quality, not speed: does the analysis identify the same risks, ask the same questions, reach the same conclusion?
  • Time-to-detect on seeded error. Insert a plausible but incorrect figure or a subtly flawed argument into AI-generated material entering normal review. How long until someone catches it, and at what stage? Detection latency is the cleanest available proxy for whether review remains substantive.
  • Escalation quality. When a case exceeds what the tool handles, what does the escalation contain? A team retaining expertise escalates with a diagnosis. A team that has lost it escalates with a description.

Aviation provides the working precedent. As cockpit automation expanded, the industry recognized that pilots who primarily monitor systems lose manual flying proficiency, and that the loss surfaces only when automation disengages. The response was not to remove automation but to require periodic manual operation and to treat proficiency as something demanding active maintenance. A US Department of Transportation Inspector General audit criticized the regulator on the measurement point specifically: it could not determine how often pilots actually flew manually. That audit finding describes the position most organizations occupy today.

Why is the junior pipeline the wrong place to look first?

Because hiring data reports a condition that has already existed for years. Payroll research has documented employment for early-career workers in the most AI-exposed occupations declining relative to older colleagues in the same fields, with the adjustment running through hiring rather than layoffs — which is why it is easy to miss. Nobody is dismissed. The on-ramp narrows quietly and aggregate headcount looks stable.

At the level of an individual firm this is a rear-view metric, and it is also mis-targeted. Headcount counts people, not the firm-specific knowledge those people would have accumulated. A company can rebuild its junior intake and still have lost the layer that mattered, because what the junior work produced was context, not capacity. Treating pipeline data as an early warning inverts the sequence twice over.

Frequently asked questions

Is expertise erosion the same as deskilling?

Deskilling usually describes the removal of skill requirements from a job by design — the role itself is simplified. Expertise erosion describes something different: the job still requires the judgment, the person remains nominally responsible for exercising it, and the capacity to exercise it is declining. The requirement and the capability separate, which is why it stays invisible until the tool is unavailable.

How do we set a threshold with no historical baseline?

You cannot, initially, and pretending otherwise produces arbitrary targets. The first measurement establishes the baseline rather than testing against one; the threshold is set on the trend across the following two or three cycles. This is an argument for starting sooner with a rough instrument rather than later with a precise one — every quarter without a baseline is a quarter of change that becomes permanently unobservable.

Should we restrict AI use in areas where erosion is a risk?

Rarely, and never as a first move. Restriction sacrifices a measured gain to address an unquantified risk. Where the criterion identifies genuine exposure, the usual remedy is not restriction but reallocation — protecting a defined portion of the work as unassisted practice while the rest proceeds normally. That preserves the gain and the control simultaneously.

Who should own this measurement — HR or the business line?

The business line. HR can hold the method and the cadence, but the threshold is an operating decision about acceptable capability risk in a specific function, and it belongs to whoever is accountable for that function’s output. Measurement owned solely by HR becomes a reporting exercise disconnected from the decisions it should trigger.