[research] · · 3 min read
MARCH Code Judge Fails 78% of Comparisons Without Evidence
A new arXiv study shows that multi-agent code judges often lack the specific evidence needed to distinguish between two solutions, leading to high rates of non-discrimination.
By ByteBulletin Editor · Editor

AI-generated illustration · Z-Image-Turbo, self-hosted
The Confidence Gap in Code Judging
According to the arXiv paper, when a language model judges whether another model's code is correct, it rarely reports the absence of evidence. Instead, it returns a confident verdict with reasoning attached, making it indistinguishable from a verdict based on solid grounds. This behavior creates a critical blind spot in automated code evaluation, particularly in multi-agent verification systems that decompose judgments into checkable claims. The paper argues that these methods require two specific conditions of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. While the second condition holds automatically for retrieved documents, it frequently fails in code judging scenarios.
MARCH Framework and Benchmark Results
The study evaluates MARCH, a published multi-agent verification framework, by running it unmodified over 80 condition-by-cell measurements on two code judging benchmarks. The results reveal a significant performance deficit. MARCH declares both solutions equally good on 78 to 95% of comparisons. In instances where it does attempt a distinction, its accuracy drops to 4.4%, whereas the same model asked directly reaches 43.7% accuracy. The paper notes that neither easier problems nor a larger judge model changes this outcome, suggesting the issue is structural rather than a matter of model scale or problem difficulty.
Label-Free Measurements
The authors propose two measurements taken from the pipeline's own logs to explain this failure mode without needing ground-truth labels. These measurements assess whether the evidence provided to the judge actually differentiates between the two code candidates. When the evidence does not differ between the candidates, the judge has no basis for preferring one over the other, yet standard frameworks still force a verdict. By gating on one of these measurements, the pipeline can decline comparisons it cannot make. This approach raises accuracy from 20.7% to 36.9% while still answering half of all comparisons. The core contribution is not a more accurate judge, but a label-free way to identify when a judge has no basis for its answer.
Implications for Multi-Agent Systems
This finding challenges the assumption that decomposing a judgment into checkable claims inherently improves reliability. In document-based retrieval, the evidence is static and distinct for each query, ensuring the second condition is met. In code judging, however, the evidence is often the code itself or generic best practices, which may not differ between two functionally similar solutions. This structural limitation means that multi-agent verification can amplify confidence in the absence of discriminative evidence. Developers relying on these systems for automated code review or competitive programming evaluation should be aware that a "tie" or a low-confidence verdict may be the more accurate reflection of the judge's state, rather than a forced choice.
What to Watch
- Adoption of Gating Mechanisms: Whether major AI coding tool vendors will implement similar label-free gating to decline low-confidence judgments.
- Benchmark Standardization: If new benchmarks emerge that specifically test a judge's ability to recognize when evidence is insufficient.
- Hybrid Approaches: The development of systems that combine multi-agent verification with explicit uncertainty reporting to developers.
This research highlights a fundamental limitation in current AI code evaluation pipelines. By identifying the specific condition under which judges fail, the paper provides a practical path toward more honest and reliable automated code assessment.
Get the signal, not the noise.
One short email when it matters. No recaps of recaps.
SHARE
RELATED

[tooling] ·
Anthropic launches Claude Code Projects for multi-agent workflows

[research] ·
OpenAI agents brute-force UN site after API limits

[research] ·
OpenAI agents hacked government databases and leaked user images

[research] ·
OpenAI pauses training after model exploits sandbox loophole

[research] ·
Irregular testing errors sent AI agents to real targets

[research] ·