让AI协作找共识,比对抗辩论更准识真。
Collaborative Disagreement Resolution for Scalable Oversight

- AI不再互辩立场,而是共同找分歧点并查证据
- 非专家模型判别准确率达62.1%,超传统辩论49.2%
- 适合需要可信决策的自动化系统,如内容审核
辩论作为可扩展监督的关键方法,面临根本矛盾:模型为说服裁判而表达,未必追求真理。本文提出一种新范式——分歧解决,将对抗性辩论转为合作求真。借鉴人类调解机制,设计自动化流程,引导模型识别分歧、查验证据,达成共识或定位核心争议点。实验显示,该方法使非专家模型判断准确率达到62.1%,显著高于标准辩论的49.2%。结果表明,从对抗说服转向协作求真,是提升可扩展监督可信度的有效路径。
原文摘要 · Abstract (English)
Debate, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight. However, debate faces a fundamental tension: models are incentivized to be persuasive to the judge, which may not always align with epistemic honesty. In this work, we propose an alternative paradigm: disagreement resolution, which reframes the interaction mechanism from adversarial debate to collaborative truth seeking. Drawing on principles from human mediation and conflict resolution, where mediators facilitate dialogue to help disputing parties reach consensus rather than adjudicating between them, we design an automated pipeline that adapts these strategies to AI oversight. Unlike standard debate where models argue for fixed positions, our pipeline directs models to collaboratively identify points of disagreement, examine the evidence for conflicting claims, and converge toward consensus or isolate the specific ''crux'' of their disagreement. We find that Disagreement Resolution consistently helps non-expert models identify the truth, achieving 62.1% judging accuracy compared to 49.2% for standard debate. Our results provide encouraging empirical evidence for rethinking the scalable oversight protocol from adversarial persuasion to collaborative truth-seeking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。