ImpossibleRubrics: when the rubric is the reward, your judge is the attack surface
The setup in ImpossibleRubrics is almost rude. Take a task that has no valid solution, hand it to an attacker model, and let a judge score the attempt against a rubric. If the rubric were any good, the score would be zero every time, because nothing legitimate can pass. Then count how often it isn't zero.
The paper (arXiv 2609.16816, from Peking University, the Chinese Academy of Sciences and JD.com) builds 169 of these impossible tasks, each with an oracle certificate that says why the task cannot be solved. Attacker and judge are held fixed. What varies is who wrote the rubric. Across 11 LLM-based rubric generators, between 8% and 26% of the tasks got exploited: the attacker found an answer the judge accepted for something that could not be right.
That range is wide, and I'd read it as a warning rather than a ranking. The number I keep coming back to is from the harder slice of 45 items. The best generator still failed on 36% of them. Rubrics written by humans who stayed faithful to the certificate failed on 0 of 45.
What an impossible task actually measures
Most rubric evals ask whether a model can satisfy a checklist. This one asks whether the checklist can be satisfied by cheating. Those are different failure modes, and normal benchmarks only see the first.
My reading of the mechanism (I only have the summary, so treat this as interpretation): a generated rubric is written by a model that looks at the task and produces plausible criteria. Plausible is the problem. Criteria like “explains the approach clearly” or “provides a complete solution” are things a fluent answer can hit without being correct. On an impossible task, a fluent wrong answer is the only kind that exists, so every pass is a false positive by construction. You get a clean count of how leaky the judge is.
Real tasks are not labelled impossible, of course. But real tasks contain impossible sub-parts all the time: the API that doesn't have that parameter, the proof with a gap, the requirement that contradicts another requirement. A rubric that can't reject the impossible version can't reliably reject the merely wrong one either.
Why optimization pressure is the multiplier
I have seen the milder version of this in my own eval setups, and I suspect you have too. You build an LLM-as-judge check, spot-check twenty outputs, agree with the judge on nineteen, and ship it. Fine. That measures the judge against an average output.
Nobody optimizes against the average output. Once the score becomes a training reward, a prompt-tuning target or a CI gate that people quietly learn to game, the population of outputs shifts toward whatever the judge likes. The attacker in ImpossibleRubrics is a small, explicit version of that shift. A 20% exploit rate under a fixed adversary is not 20% of your traffic, but it tells you a search process will find holes, and search processes are cheap.
A small calculation helps. Say your rubric passes a bad answer on 15% of hard items and you run best-of-8 selection on the judge score. The chance that at least one of eight bad candidates slips through is 1 minus 0.85 to the eighth, roughly 73%. I picked those numbers, they are not from the paper, and real attempts are correlated so the true figure will differ. The direction is the point: sampling more against a leaky judge makes the leak worse, not better.
A judge you validated on ordinary outputs has told you nothing about how it behaves against outputs that were selected to fool it.
Certificate-faithful rubrics, and what I would steal
The zero-out-of-45 result comes from rubrics that stay faithful to the certificate. In practice that means every criterion is tied to a reason the answer is right or wrong that you can state independently of the answer. It is slower to write. It is also the only kind that held.
I'd steal three habits. First, seed your eval set with impossible or unanswerable items on purpose, and treat any pass as a bug report against the rubric, not a win for the model. Cheap to build for code (an unsatisfiable spec) and for retrieval (a question whose answer isn't in the corpus). Second, separate whoever writes the rubric from whoever tunes against it, and never regenerate the rubric after seeing failures without re-running the seeded impossibles. Third, prefer criteria a script can check over criteria a model must interpret, and keep the LLM judge for the residue.
I wrote about the general trouble with LLM-as-judge before, mostly from a calibration angle. This paper adds the adversarial angle, and it is the one that bites as soon as a score drives anything automatic. The story of eval pipelines built on a labor source that vanished is a reminder that eval infrastructure tends to be fragile in exactly the places nobody looked.
One limit I should be plain about: I haven't run ImpossibleRubrics against my own judges, and I only have the headline numbers. I don't know how the 11 generators were prompted, how strong the attacker was compared with what a determined team would build, or how the 45-item slice was chosen beyond “harder”. It could overstate the problem for narrow, well-specified domains. Or understate it for open-ended ones.
If you run one experiment this week, make it this: write ten impossible items for your current eval, run your normal pipeline, and count the passes. If the count is not zero, your score has been telling you something other than quality.