← all posts
// economics · economics

54 percent of customer-comms tasks stick. Legal research doesn't.

Most adoption numbers are useless because they count logins. OpenAI's second Work at the Frontier report, published on 16 September, counts something better: stickiness, meaning whether a task that an employee tried with AI keeps showing up in their monthly workflow afterwards. I'm reading it through a secondary summary and haven't dug through the full methodology, so treat the exact definitions as something to check at the source.

The headline figures are these. Customer communication sticks at 54 percent. Promotional writing sticks at 44 percent. Specialized work, with legal research as the example, sticks less. The report also says that when people step outside their core expertise, their prompts get shorter and more direct.

What the pattern says

Look at what the two sticky categories have in common. Both are high-volume, both have an obvious acceptable output (a reply to a customer, a blurb for a product), and in both the person prompting can judge the result in seconds. You know if the email sounds right. Failure is cheap and visible, so you keep going.

Legal research is the reverse. Output is long, citations need checking, and a confident wrong answer costs far more than an awkward paragraph. If you have to re-verify every line by hand, the time saved shrinks, and the habit doesn't form. I wrote about that dynamic in verification as the bottleneck: the task sticks when checking is cheaper than doing.

That is my interpretation, not a finding of the report. The report gives rates; the why is my inference, and other explanations fit as well. Maybe legal work simply has fewer repeatable tasks per person per month, which would lower any stickiness metric regardless of model quality.

Adoption holds where the user can grade the answer faster than they could have written it.

The short-prompt detail

The prompt-length finding is the part I keep thinking about. Outside their own field, people write shorter, more direct prompts. It sounds harmless, but it is the opposite of what you'd want. Inside your expertise you know which constraints matter, so you can write a long, specific prompt. Outside it, you don't know what to specify, and you also can't spot a plausible but wrong answer. Short prompt plus weak judgment is the combination where a model's confident tone does the most damage.

For a team lead that cuts two ways. It is a real benefit that a marketer can now draft a first-pass contract summary or a developer can write decent customer copy. It is also exactly where the review step should be heaviest, and exactly where it is lightest in practice, because nobody in the room feels qualified to push back.

Planning around it

If you run AI rollouts, the practical reading is to stop measuring seats and start measuring which tasks survive a month. A rough way to do it: pick ten recurring tasks, log who tried AI on each in week one, and check in week five. Anything under about half probably needs a better workflow (templates, a checker, a narrower scope), not more training slides.

There's a macro layer that I'd only add if it fits, and I think it loosely does. Anthropic's Econ Scenario Explorer, which I haven't gone through in depth, lays out three paths for the US economy by 2030: a modest scenario at +1.6 percent GDP, a substantial one at +8.3 percent with a labor share of 56.1 percent, and an extreme one at +32.4 percent with labor share at 45.2 percent. Those are scenarios, not forecasts. The link to stickiness is that every one of them depends on tasks actually staying automated inside real workflows, and the OpenAI data is an early, small look at how much of that happens.

So the useful number for your own team is not a GDP scenario. It is the percentage of tried tasks that are still running in month two, broken down by how easy they are to verify.

I'd bet your own split looks a lot like OpenAI's. It costs a spreadsheet to find out.

#economics#workflow#analysis#policy