arXiv:2605.27483cs.CLcs.AI2026-05被引 1

辩论让弱裁判更好识别强模型,关键在裁判能验证批评者主张。

Debate Helps Weak Judges Reward Stronger Models

  • 裁判需将批评视为待验证主张而非陈述,且批评者能力须强于裁判。
  • 在五组配对中三组辩论显著优于基准,仅当批评者胜过裁判时有效。
  • 无需多轮反驳,单次独立批评即可发挥辩论主要优势,成本更低。

尽管辩论作为可扩展的监督协议理论上具有潜力,但实证结果参差不齐:某些场景下有提升,其他则无效果,尤其当裁判未掌握隐藏信息时。本文在程序可验证的代码与逻辑任务中研究了提议者-批评者辩论,在强辩手/弱裁判设定下发现,当批评者具备可利用的优势时,辩论能帮助裁判超越咨询基线。即批评者的分类能力必须超过裁判,且裁判需将批评内容视为待验证的主张而非总结性陈述。在五组配对中,有三组满足条件,辩论效果显著优于咨询基线,且这些配对均为表现最强的模型组合。另外两组不满足条件的配对中,辩论无显著效果,裁判验证率下降数十个百分点。此时批评者与裁判的二分类能力相近,批评分歧被误读为证词而非待验证主张。消融实验表明,移除反驳回合对裁判表现无明显影响:单次独立批评已可恢复辩论的大部分收益,推理成本更低。这些发现揭示了一种训练无关的可扩展监督廉价范式(答案、批评、裁判),并提出部署前审计标准:批评者是否优于裁判?裁判能否验证其主张?该标准可预测辩论是否有效。

原文摘要 · Abstract (English)

Despite theoretical promise, debate as a scalable oversight protocol has produced mixed empirical results: gains in some settings, and null effects in others, especially when the judge does not have information hidden from it. We study proposer-critic debate in a stronger-debater/weaker-judge setting on programmatically verifiable code and logic tasks. Debate helps the judge over a consultancy baseline when the critic provides a usable advantage: the critic's classification ability must exceed the judge's, and the judge must treat critic speeches as claims to verify rather than testimony to summarize. On the three of five pairings where the condition holds, proposer-critic debate's gains are statistically significant over consultancy, and these pairings are the most capable model pairings. On the two non-responder pairings in our set, debate produces null effects, and judge verification rates drop by tens of percentage points once a critic enters the transcript. In these cases the critic's binary-classification ability and the judge's are within noise of each other, and the critic's disagreement is parsed as testimony rather than a claim to check. Ablating rebuttal rounds from debate produces no measurable change in judge performance: a single independent critique recovers the bulk of debate's benefit at lower inference cost. These findings suggest a cheaper primitive for training-free scalable oversight in verifiable domains (answer, critique, judge) and a pre-deployment audit (does the critic beat the judge, and will the judge verify it?) that predicts when debate will help.

辩论机制模型评估可验证任务监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。