How Do You Know What a New Hire Can Do Without AI?

Most hiring processes now measure a joint output: the candidate and whatever tools they used. What the receiving team needs is a different number — what this person can produce when the assistance is not there, and where their judgment is not yet reliable. Organizations that never establish it build the team’s division of labor around an estimate no one has tested.

The failure does not announce itself as a weak hire. A new colleague who leans heavily on tooling still produces acceptable work. It announces itself months later, as a gap in what the team assumed someone was covering.

What does a hiring process actually measure now?

Assisted output, on both sides of the table. The candidate prepares with a model; the employer screens with one. Everything observed before the offer is a joint product, and the joint product is what gets converted into an expectation about the person.

The field’s response has been detection and policy. Interview technique that probes past a rehearsed answer, tiered rules on what candidates may and may not use, disclosure requirements. These are reasonable, and they answer a different question. They establish whether assistance was present. They do not establish what remains when it is absent — and that second quantity is the one the team will actually rely on.

Why does the team make this worse rather than better?

Because a new member does not arrive into an empty structure. They arrive into an existing division of cognitive labor and are immediately assigned a position in it.

Lewis, Belliveau, Herndon and Keller demonstrated the consequence in 2007. Groups that replaced one member performed worse on a knowledge-intensive task than groups that stayed intact — and worse than groups reconstituted entirely. The mechanism was not the newcomer’s ability. Incumbents carried forward the specialization they had built with the departed member and expected the replacement to occupy the same slot. Partial change was more damaging than total change, because total change forces a team to rebuild the map, while partial change leaves an obsolete map in place.

Add an inflated entry estimate to that finding and the effect compounds. The team assigns a competence slot on day one, inherited from a predecessor and calibrated on an assisted sample. Nobody checks it, because nothing looks wrong. Work is getting done. This is the practical form of what an organization can still do without its tools, asked at the one moment when the answer is genuinely unknown.

What does the evidence say about assistance hiding differences in skill?

That it hides them most precisely where a new hire sits.

Brynjolfsson, Li and Raymond studied the staged rollout of a generative AI assistant across 5,179 customer-support agents. Average productivity rose about 15 percent. The distribution is the finding: novice and lower-skilled agents improved by roughly 34 percent, while the most experienced and highest-skilled saw little gain and, on some quality measures, slight decline. The tool did not lift everyone. It lifted the bottom of the distribution toward the top.

For hiring, that is the whole problem stated as a coefficient. Assistance compresses observable differences exactly at the low-experience end — and every new hire is at that end relative to the work in front of them, however senior their record. Whether that compression persists depends on the work — it holds where good performance follows a transferable pattern, and reverses where judgment itself is the task. The signal is weakest where the organization needs resolution most.

One honest limitation: this was structured support work with a transferable pattern, in a single firm. Where the work is genuinely ambiguous, the compression is smaller — but the misreading is more expensive, because ambiguous work is where an incorrect directory entry stays hidden longest.

Why does this matter more now than two years ago?

Because the entry point is where the labor market is actually changing. In an analysis revised in August 2026, Brynjolfsson, Chandar and Chen found employment for 22-to-25-year-olds in AI-exposed occupations running roughly 19 percent below comparable workers in less exposed occupations, with the adjustment operating primarily through reduced hiring rather than separations.

The organizational consequence is not the headline one. Firms are making fewer entry-level hires, which means each one carries more structural weight, and the long observation period that used to reveal capability slowly is shorter. Fewer entries, higher stakes per entry, weaker signal at the entry. That combination is new.

What should the first weeks do differently?

Treat it as placement, not assessment. The distinction is not cosmetic — it determines whether you get an honest sample.

  • One early piece of real work produced without tooling, chosen by the receiving team rather than by HR, on a live problem rather than an exercise
  • State the purpose to the person openly: to place them accurately in the team, not to test whether they should have been hired. A sample gathered under evaluative pressure measures composure, not capability
  • Read the interaction, not the artifact. Where did they ask? Where did they stop? What did they flag as uncertain, and was the flag in the right place?
  • Record what the person should not yet be relied on for. The negative entry is the one that prevents mis-assignment, and it is the one nobody writes down
  • Owned by the team lead, kept out of the probation file. The moment it touches employment decisions, it becomes an appraisal and the signal dies within one cycle

The obstacle differs by market. In the United States the binding constraint is that fewer junior hires means less tolerance for a slow placement. In the Gulf, senior hiring is fast and status-sensitive, and asking an experienced arrival to work unassisted reads as doubt unless it is framed explicitly as placement — which is also where the largest inherited slots and the least scrutiny sit. In Central and Eastern Europe, teams are thin enough that a newcomer becomes load-bearing within weeks, leaving no slack to absorb a wrong entry.

Frequently asked questions

Isn’t this just testing new hires?

No, and it fails if it becomes that. A test asks whether to keep someone. This asks what to rely on them for. Different question, different owner, different record — and nothing that enters a performance file.

Should we restrict AI use during onboarding?

No. The point is a sample, not a policy. A general restriction produces resentment and theatre; one deliberate unassisted piece of work produces information.

Does this apply to senior hires?

More, not less. Senior arrivals inherit the largest assumed competence and attract the least verification, so an inaccurate entry propagates further before anything surfaces.

How is this different from a probation review?

Probation evaluates the hiring decision after the fact. This calibrates the team’s map at the point where the map is being drawn. Doing the second well usually means the first has nothing to correct.