高精度评审未必提升纠错效果,关键在能否落实批评建议。
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

- 对比评审与广播讨论,发现后者更有效
- 评审精度高(0.861)但实际采纳率低,纠错效果差
- 将评审意见嵌入解题上下文可部分提升执行率
许多数学与科学类智能体系统采用分层设计,设有专门的评审角色,假设评审阶段能将错误答案修正为正确。我们在4,181个verifier-grounded Omni-MATH问题上,使用匹配的gpt-oss-120b代理测试该假设。在最简单层级中协作收益微弱,但从第4级开始,收益显著上升;此时广播式同行讨论的最终准确率高于规划-执行-评审(PER)流水线。我们探究这一差距是否由评审质量决定,结果表明并非如此:尽管PER评审精度更高(0.861 vs. 0.644),但其有效批评更难改变后续候选答案,且指导修复效果较差。这说明评审识别能力与批评采纳率在实证上可分离。在相同PER干预下,强制要求显式承认会降低最终准确率,而将评审建议直接嵌入求解器工作上下文则部分提升了执行率,但未完全弥合差距。总体而言,以评审为中心的评估可能夸大系统性能:一个系统虽能精准识别错误,若不执行批评,则仍难以解决更多问题。
原文摘要 · Abstract (English)
Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors. Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion reaches higher final accuracy than a planner-executor-reviewer pipeline (PER). We ask whether this gap is explained by reviewer quality or by whether critique changes the next answer the protocol carries forward. It is not explained by reviewer precision alone: PER's reviewer is more precise than broadcast's (0.861 vs. 0.644), yet evaluator-verified useful critique is much less likely to change the next candidate and produces lower reviewer-guided repair. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy, while embedding reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Overall, reviewer-centric evaluation can overstate system quality: a protocol may spot errors well yet still fail to solve more problems if it does not act on those critiques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。