Auditing AI adoption on the AL0-AL5 scale without fooling yourself
Anthropic published a number on September 17 that people will be quoting at each other for months. Claude now leads 26% of the company's AI research and development work, which is level AL4 on the AL0 to AL5 scale from Epoch AI. In February the figure was under 1%. More than 90% of the work sits at AL3 or above (the report calls that AI collaborating with people), no area is fully autonomous, and in August roughly 30,000 agents were running at one moment on the most used internal platform.
Impressive, if you take it at face value. I'd rather look at how they got it, because the method is the reusable part, and also the place where the number goes soft.
How they measured, and where it bends
As I read the coverage, an agent lists tasks pulled from Slack and documents, a second Claude acts as judge and rates how automated each task was, and every week a 20% sample of people is included. So a model extracts the tasks, a model grades them, and humans are sampled. It is self-reported, and the judge belongs to the same family as the thing being judged. Anthropic itself frames the result as a trajectory signal rather than an independently verified productivity benchmark, and I'd read it that way too. A curve from under 1% to 26% in seven months tells you direction. It can't tell you whether the work was worth doing.
I'd watch three bends. First, the denominator: 26% of tasks, of hours, or of value? Small tasks are easier to hand over, so share of tasks flatters. Second, extraction bias. Work that lives in Slack and documents is visible, whereas a hallway conversation or a decision made in someone's head never becomes a task. Third, a judge grading its own kind may be generous, and nobody outside the lab can check.
Running the audit on your own team
You don't need a lab to copy the skeleton. Pick four weeks. Each week, sample a fixed fraction of your engineers, pull the tasks they closed, and label every task on the scale. I can't source definitions for every level from what I read (the coverage has AL2 as AI writing code, AL3 as collaboration, AL4 as AI leading the task), so write your own one-line anchor for each level, with a real example from your repository, before anyone labels anything. Vague anchors are where audits die.
Then do the step that a self-reported pipeline skips. Have two humans label the same thirty or so tasks independently and compare them with whatever automated judge you use. If two people disagree on a third of the tasks, your scale is the problem, not your team.
Sample size deserves the same honesty. The numbers here are my assumptions, not data. Take 40 engineers, a 20% weekly sample (8 people) and about 15 closed tasks each per week, so 120 labelled tasks. If the true AL4 share were 26%, the standard error is sqrt(0.26 × 0.74 / 120) = 0.040, which makes a 95% interval about plus or minus 7.8 points. Your "26%" would mean anything from 18 to 34. And tasks cluster by person, so one enthusiast can drag the whole estimate, which widens the real interval further. Report the interval or don't report the number.
An adoption percentage without a denominator and an error bar is a mood, not a metric.
What to put next to the headline share
I'd report three figures side by side: the share of tasks at each level, the share of estimated effort at each level, and the fraction of AL4 tasks that got through human review without rework. The third is the one a CFO will care about, and as far as I know nobody has published it. If the AL4 share climbs while rework climbs with it, you've automated typing, not delivery, and you're back at the constraint I described in verification is the new bottleneck.
Then run it again in a quarter with the same anchors and, ideally, some of the same labellers. One measurement is a number. Two measurements with an identical rubric are the beginning of a trend. And keep one person on the labelling team whose job is to doubt the labels.