← all posts
// tooling · tooling

Alibaba open-sourced its code reviewer, and half of it isn't an LLM

The most interesting line in the open-code-review README is an architectural one: it is a hybrid, a deterministic pipeline plus LLM agents. Alibaba open-sourced it on 16 September and says it is battle-tested in their own production. I haven't run it on a repo, so this is a read of the design as described, not a verdict on the results.

But the design alone is worth a post, because it disagrees with how most of us bolt AI onto pull requests.

What the deterministic half is for

The typical setup is: take the diff, paste it into a model, ask for review comments. It works well enough to be seductive, and it fails in a very particular way. The model is asked to do two jobs at once. One is pattern detection (is this value dereferenced without a null check, is this string concatenated into SQL, is this shared map mutated from two threads). The other is judgment (is this abstraction right, does this name lie, is this change going to hurt us in six months).

The first job has right answers. A null pointer path either exists or it doesn't. Doing it with a probabilistic text generator means it will sometimes miss the path on Tuesday and find it on Wednesday, same code, same prompt. That is a strange property for something in your merge gate.

Open-code-review ships built-in, multi-language rules for exactly these classes: NPE, thread-safety, XSS and SQL injection. Rules like that are what static analysis has done for twenty years. They are cheap, they run in milliseconds, they give the same answer every time, and when they're wrong you can open the rule and fix it. Nobody has to re-prompt anything.

A reviewer that can't repeat itself on identical input is a suggestion engine, and you shouldn't wire a suggestion engine to a required check.

What the model still has to do

So why add agents at all? Because rules only catch what somebody thought to write a rule for. They have no opinion on whether a function does two unrelated things, or whether the test you added actually exercises the bug. The brief I worked from describes the system as producing line-level review, which I read as findings anchored to specific lines of the change rather than a wall of prose at the top of the PR. That matters more than it sounds: a comment attached to line 84 can be resolved or dismissed. A paragraph of vibes cannot.

My guess at the division of labour, and it is a guess from the architecture rather than from reading the code: deterministic stages decide what is definitely wrong and what context to gather, the LLM agents handle interpretation, explanation and the fuzzy stuff. That ordering has a nice side effect. The model gets a smaller, pre-filtered problem, so it burns fewer tokens and has less room to invent an issue that isn't there.

It also ships with native compatibility for both OpenAI and Anthropic models, which means the model is a swappable part and the pipeline is the product. I like that. Models change every few weeks (check the benchmark page if you doubt me), and a review process you have to rebuild each time is a bad investment.

Numbers I don't have, and the ones you should collect

Here is the gap. The project claims production use, but I don't have false-positive rates, recall against a labelled bug set, or cost per review. Without those, “hybrid is better” is a hypothesis I find plausible, not a fact. Plausible, because the failure modes of pure-LLM review (inconsistency, invented findings, nothing to audit) are the exact things deterministic rules fix. Unproven, because hybrids have their own tax: two systems to maintain, rule sets that rot, and findings that get reported twice, once by each half.

If you want to test this on your own code, the experiment is small. Take twenty merged PRs where a bug later slipped through and twenty that were clean. Run a pure-LLM reviewer and the hybrid on all forty. Count caught bugs, count comments that somebody on your team would call noise, and record the spread when you run each one three times. That last number is the one vendors never show you.

I made a similar argument when I wrote about reviewing AI-heavy changes: the tool's job is a first pass that never skips a PR, and the human still owns the merge. The hybrid pattern makes that first pass more trustworthy, because the part of it you can predict is actually predictable. There is a related idea in graph-based review context, where structure from the code graph narrows what the model has to read.

What I'd steal for an IDE plugin

If I were building this into a JetBrains plugin, I'd copy the shape and skip the scale. Run PSI-based inspections first, they're already deterministic and already know the language. Feed only the flagged regions plus their callers to a model. Show the rule hit and the model's explanation side by side, with the source of each labelled, so a developer can tell at a glance which claim is a fact and which is an opinion.

That labelling is the whole trick, really. Most of the distrust people feel toward AI review comes from not knowing which kind of claim they're reading. Two channels, clearly marked, would fix a lot of it.

The open question for me is maintenance. Who owns the rule set when your codebase grows a new framework, and does the LLM half quietly start covering for rules nobody updated?

#tooling#workflow#security#architecture