用AI辩论帮人摆脱偏见,更准判断新冠与气候争议事实。
AI Debate Aids Assessment of Controversial Claims
- 让两个AI分别辩护对立观点,引导有偏见的人类判断趋近真相。
- 人类在辩论下准确率提升4%-10%,主流信念者最高增15.2%。
- 具人格化的AI裁判比人类更准(78.5%),适合监督前沿AI。
随着AI能力增强,其对世界认知的影响日益深远,但也可能放大错误信息并加剧社会分歧,尤其在影响福祉的关键议题上。为实现可扩展的可信监督,需确保AI系统在能力超越评估者时仍保持真实。但人类评估者自身信念与偏见会影响判断。本文研究AI辩论能否引导有偏见的裁判走向真相,聚焦新冠和气候变化等公众持强烈先入之见的争议性事实。两组实验显示:在人类裁判中,辩论协议相比单方咨询,准确率普遍提升4%-10%,主流信念者提升最多(新冠案达+15.2%),而质疑派也向正确观点靠拢(+4.7%)。在使用具人格化特征的AI裁判时,准确率达78.5%,高于人类裁判(70.1%)及无个性的默认AI裁判(69.8%)。结果表明,AI辩论是应对争议领域偏见、实现可扩展可信监督的可行路径。
原文摘要 · Abstract (English)
As AI grows more powerful, it will increasingly shape how we understand the world. But with this influence comes the risk of amplifying misinformation and deepening social divides-especially on consequential topics where factual accuracy directly impacts well-being. Scalable Oversight aims to ensure AI systems remain truthful even when their capabilities exceed those of their evaluators. Yet when humans serve as evaluators, their own beliefs and biases can impair judgment. We study whether AI debate can guide biased judges toward the truth by having two AI systems debate opposing sides of controversial factuality claims on COVID-19 and climate change where people hold strong prior beliefs. We conduct two studies. Study I recruits human judges with either mainstream or skeptical beliefs who evaluate claims through two protocols: debate (interaction with two AI advisors arguing opposing sides) or consultancy (interaction with a single AI advisor). Study II uses AI judges with and without human-like personas to evaluate the same protocols. In Study I, debate consistently improves human judgment accuracy and confidence calibration, outperforming consultancy by 4-10% across COVID-19 and climate change claims. The improvement is most significant for judges with mainstream beliefs (up to +15.2% accuracy on COVID-19 claims), though debate also helps skeptical judges who initially misjudge claims move toward accurate views (+4.7% accuracy). In Study II, AI judges with human-like personas achieve even higher accuracy (78.5%) than human judges (70.1%) and default AI judges without personas (69.8%), suggesting their potential for supervising frontier AI models. These findings highlight AI debate as a promising path toward scalable, bias-resilient oversight in contested domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。