推理模型通过模拟思想社会实现更优解题,而非单纯延长思考过程。
Reasoning Models Generate Societies of Thought
- 将内部思维建模为具不同性格与专长的多主体互动群体。
- 相比普通模型,推理模型在思考中展现更高视角多样性与观点冲突。
- 适合关注认知机制、人机协作与群体智能的研究者阅读。
大型语言模型在多个领域表现卓越,但其复杂推理机制仍不清晰。近期推理模型在高阶认知任务上优于同类指令微调模型,传统归因于更长的思维链。本文发现,推理能力提升并非仅源于计算延展,而是源于模拟多主体互动——即“思想社会”,使内部认知视角具备多样性和辩论性,表现为不同性格特质与专业领域的差异化激活。通过对 DeepSeek-R1 与 QwQ-32B 等模型的推理轨迹进行量化分析与可解释性研究,发现其在推理过程中产生的视角多样性显著高于指令微调模型,且在思维中表现出更强的异质性特征冲突。这种多主体结构体现为问答行为、视角切换及观点调和等对话特征,并伴随强烈的社会情感互动,共同促成任务准确率优势。受控强化学习实验显示,仅以推理准确率为奖励信号时,基础模型会自发增加对话行为;而采用对话结构进行微调的模型则比基础模型更快提升推理能力。这些结果表明,思想的社会化组织能有效探索解空间。我们提出,推理模型建立了一种类比人类群体集体智慧的计算机制,当多样性被系统化组织时,可实现更优问题解决,为新型智能体架构设计提供新思路。
原文摘要 · Abstract (English)
Large language models have achieved remarkable capabilities across domains, yet mechanisms underlying sophisticated reasoning remain elusive. Recent reasoning models outperform comparable instruction-tuned models on complex cognitive tasks, attributed to extended computation through longer chains of thought. Here we show that enhanced reasoning emerges not from extended computation alone, but from simulating multi-agent-like interactions -- a society of thought -- which enables diversification and debate among internal cognitive perspectives characterized by distinct personality traits and domain expertise. Through quantitative analysis and mechanistic interpretability methods applied to reasoning traces, we find that reasoning models like DeepSeek-R1 and QwQ-32B exhibit much greater perspective diversity than instruction-tuned models, activating broader conflict between heterogeneous personality- and expertise-related features during reasoning. This multi-agent structure manifests in conversational behaviors, including question-answering, perspective shifts, and the reconciliation of conflicting views, and in socio-emotional roles that characterize sharp back-and-forth conversations, together accounting for the accuracy advantage in reasoning tasks. Controlled reinforcement learning experiments reveal that base models increase conversational behaviors when rewarded solely for reasoning accuracy, and fine-tuning models with conversational scaffolding accelerates reasoning improvement over base models. These findings indicate that the social organization of thought enables effective exploration of solution spaces. We suggest that reasoning models establish a computational parallel to collective intelligence in human groups, where diversity enables superior problem-solving when systematically structured, which suggests new opportunities for agent organization to harness the wisdom of crowds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。