[research]By ByteBulletin Editor
MARCH Code Judge Fails 78% of Comparisons Without Evidence
A new arXiv study shows that multi-agent code judges often lack the specific evidence needed to distinguish between two solutions, leading to high rates of non-discrimination.



