When agents write 80% of the code, CI breaks first: Anthropic's test impact analysis and a 25x plan
Anthropic's engineering team published a post about scaling their test impact analysis service, and the opening numbers are the interesting part. Their engineers ship about 8x more code per quarter than in the 2021 to 2025 period. Claude writes roughly 80% of it and plays a large role in reviewing and approving pull requests too. CI jobs grow exponentially with the number of agents per engineer, and the advice is to plan for 25x the load within two quarters.
I read that as a warning from the team furthest along. Something is going to break at your company as well, and my bet is that it will be the pipeline, not the model.
I've only seen a summary of the post, not every detail, so the arithmetic below is mine and illustrative.
Why CI takes the hit first
A person opens a couple of PRs a day, and each PR runs the suite a few times as they push fixes. An agent does not get tired, does not wait for lunch, and retries by reflex. Put several agents next to one engineer and each iterates on its own branch, and the number of CI runs is no longer a linear function of headcount. It is headcount times agents per head times pushes per task.
Here is a small model with invented numbers. Say 100 engineers open 2 PRs a day, with 3 CI runs each: 600 runs. Give each engineer 4 agents, each producing 5 PRs a day at 4 runs a PR (agents push more often): 100 x 4 x 5 x 4 = 8,000 runs. That is 13x, and I have not yet counted the review agents that trigger their own checks. Twenty-five times is not exotic once you add them.
If a run costs 10 minutes on a runner and you pay 0.02 dollars a minute, 600 runs is 120 dollars a day, and 8,000 runs is 1,600. Roughly 440,000 dollars more per year for a mid-sized team, before latency, which usually hurts sooner. Queue times stretch, agents wait on red checks, and the expensive model sits idle behind a cheap runner.
The model got faster and the pipeline did not, so all the speed you bought is now waiting in a queue.
What test impact analysis does
The idea is old and simple. A change touches some files, some tests depend on those files, and the rest of the suite is noise for that PR. Test impact analysis (TIA) maps changed code to affected tests and runs only those.
There are two common ways to build the map. Static analysis follows imports and build-graph edges: cheap, deterministic, and blind to reflection, config and dynamic dispatch. Coverage-based mapping records which tests execute which lines during a full run, then looks up your diff: more accurate, but you have to refresh the data and it costs a full run to build.
Most real systems combine them and add a safety net. Selected tests run on the PR, and the full suite runs on merge or nightly, so a miss gets caught within hours rather than never. That safety net is the part people skip and the part I would insist on. A TIA that silently skips a test it should have run is worse than a slow suite, because you learn about it in production.
Anthropic's post is about scaling their own service, and I can't tell you from a summary how they handle those misses. That is the first question I would ask. (Also: what happens to a selected-tests run when the map itself is stale after a big refactor?)
How to size your own exposure
Start with three numbers from your CI provider: runs per day, median minutes per run, and PRs per day. Then count how many of those PRs come from agents today, and ask each team how many agents per engineer they intend to have in six months. Multiply. If the answer is above 5x, you need a plan now, because the slow parts (coverage data, dependency graphs, flaky test cleanup) take months.
Then work in this order. Fix flaky tests first, since retries multiply with agent volume and a 2% flake rate at 8,000 runs is 160 spurious failures a day. Add caching for build artifacts. Turn on TIA for the biggest suite, keep the full run on merge, and track the miss rate. Cap concurrent agent runs per repo so one runaway loop does not eat the budget. The review side of the same queue is in verification is the new bottleneck, and if your agents run inside CI themselves, headless Claude Code in CI covers the setup.
The uncomfortable question is who owns the pipeline when agents write 80% of the code and also approve it. Somebody has to keep the gate honest, and it may not be the ones opening the PRs.