十款顶尖大模型在辩论中普遍高估胜率,越辩越自信,自信心严重失准。
When Two LLMs Debate, Both Think They'll Win
- 让大模型进行多轮对抗辩论,实时评估其胜率信心变化。
- 模型平均胜率信心从72.9%升至83%,61.7%的对决双方同时超75%自信。
- 即使被告知胜率仅50%,模型仍持续提升信心,显示自我认知偏差。
大语言模型能否在面对反对意见时准确调整自身信心?基于此前对静态事实问答任务的校准研究,本文在动态对抗辩论场景下评估了大语言模型的表现,首次结合两个现实因素:(a) 多轮交互需模型随新信息更新信念;(b) 零和结构控制任务不确定性,因双方高信心声称暗示系统性过度自信。我们组织了10个顶尖大模型参与的60场三轮政策辩论,每轮后模型私密评分其胜率(0-100)。发现五种令人担忧的现象:(1) 系统性过度自信:初始平均信心为72.9%,远高于理性基准50%;(2) 信心上升:辩论进程中心情升至最终轮平均83%;(3) 双方同时高估:61.7%的辩论中双方均宣称≥75%胜率,逻辑上不可能;(4) 自我辩论偏见:与自身副本辩论时信心从64.1%升至75.2%;即使明确告知胜率恰好50%,信心仍升至57.1%;(5) 私密推理不一致:模型内部思考与公开信心评分存在偏差,质疑链式推理的可信度。结果表明,大模型在动态多轮任务中缺乏准确自我评估或信念更新能力,这在它们日益被用于助理和智能体角色的背景下构成重大风险。实验代码已开源。
原文摘要 · Abstract (English)
Can LLMs accurately adjust their confidence when facing opposition? Building on previous studies measuring calibration on static fact-based question-answering tasks, we evaluate Large Language Models (LLMs) in a dynamic, adversarial debate setting, uniquely combining two realistic factors: (a) a multi-turn format requiring models to update beliefs as new information emerges, and (b) a zero-sum structure to control for task-related uncertainty, since mutual high-confidence claims imply systematic overconfidence. We organized 60 three-round policy debates among ten state-of-the-art LLMs, with models privately rating their confidence (0-100) in winning after each round. We observed five concerning patterns: (1) Systematic overconfidence: models began debates with average initial confidence of 72.9% vs. a rational 50% baseline. (2) Confidence escalation: rather than reducing confidence as debates progressed, debaters increased their win probabilities, averaging 83% by the final round. (3) Mutual overestimation: in 61.7% of debates, both sides simultaneously claimed >=75% probability of victory, a logical impossibility. (4) Persistent self-debate bias: models debating identical copies increased confidence from 64.1% to 75.2%; even when explicitly informed their chance of winning was exactly 50%, confidence still rose (from 50.0% to 57.1%). (5) Misaligned private reasoning: models' private scratchpad thoughts sometimes differed from their public confidence ratings, raising concerns about faithfulness of chain-of-thought reasoning. These results suggest LLMs lack the ability to accurately self-assess or update their beliefs in dynamic, multi-turn tasks; a major concern as LLMs are now increasingly deployed without careful review in assistant and agentic roles. Code for our experiments is available at https://github.com/pradyuprasad/llms_overconfidence
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。