用大模型辩论模拟人类集体求真,提升推理准确性。
Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate

- 让多个大模型通过对抗辩论协作,模拟社会性求真过程。
- 即使单个模型表现一般,集体辩论也能显著提升答题正确率。
- 首次提出用辩论动态评估模型内在特性,超越传统静态测试。
人类推理长期被认为具有社会性,而非孤立的个体认知,这一观点称为论证式推理理论(ATR)。ATR将真理视为社会认识论的产物:在辩论对抗压力下,由有缺陷的个体推理不断修正而成。该分布式集体智能机制推动人类不断突破认知边界,并支撑所有民主制度的基础。本文首次通过大语言模型多代理辩论(LLM-MAD)模拟ATR。经严格实证分析,当模型具备认知多样性时,LLM-MAD可在问卷类任务中显著提升求真性能,即使单个参与者表现有限。此外,我们提供有力证据表明,性能提升源于ATR核心机制,说明集体推理普遍优于个体推理,非生物或进化偶然。最后,基于辩论动态分析,提出一种新基准方法,利用LLM-MAD测量模型内在属性(如幻觉倾向),实现当前静态评测无法支持的模型对比。
原文摘要 · Abstract (English)
Human reasoning has long been theorised to operate socially, not through isolated individual cognition, but through collective adversarial discourse, a framework known as the Argumentative Theory of Reasoning (ATR). Rather than relying on individual "intellectualist reasoners" as the primary vehicle for truth-seeking, ATR reconceptualises truth as an emergent property of social epistemology: the product of imperfect individual reasoning refined under the adversarial pressure of debate. This distributed method of collective intelligence has guided humanity to ever-greater epistemic heights and underpins the foundational principles of all democratic systems. This thesis breaks new ground by, for the first time, simulating ATR through the multi-agent debate (MAD) of large language models (LLMs). With rigorous empirical analysis, we demonstrate that, when correctly engineering an epistemically diverse set of models, LLM-MAD can significantly improve truth-seeking performance on questionnaire-based tasks, even when individual debate participants exhibit limited standalone performance. Furthermore, we present strong empirical evidence that this performance gain is mechanistically grounded in the central principles of ATR, suggesting that collective reasoning may be universally favourable over individualist reasoning, rather than a quirk in biology or evolution. Finally, drawing on our analysis of debate dynamics, we propose a novel benchmarking methodology that leverages LLM-MAD to measure intrinsic model properties (such as hallucination propensity) in order to compare models in ways that current static benchmarking approaches cannot support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。