arXiv:2606.19494cs.AI2026-06被引 1

揭示多智能体大模型辩论中的隐藏信念锚点,解释为何推理能超越初始观点。

Hidden Anchors in Multi-Agent LLM Deliberation

论文配图:Hidden Anchors in Multi-Agent LLM Deliberation
图 1 · 摘自论文原文
  • 将多智能体辩论建模为带隐藏信念锚点的闭环动力系统
  • 锚点可从辩论过程恢复,且使信心突破初始信念凸包
  • 适用于评估模型是否真正依赖内部信念,非所有模型都需完整闭环

多智能体大模型辩论通过多轮交互与修正答案提升推理准确性,但其机制尚不清晰。该研究将其建模为闭环动力系统,每个代理携带一个隐藏的内部信念(锚点),持续拉拽其观点,不受邻居影响。研究表明,锚点可仅从辩论过程中恢复,并解释了经典共识模型无法涵盖的现象:代理对正确答案的信心可超过所有初始观点,突破初始信念的凸包。检验恢复锚点在未见运行中是否有效,可作为判断模型是否真实依赖锚点的简单测试。在三个开源模型族中,此现象呈连续谱而非二元状态;各锚点影响力相近,但位置不同;仅当锚点远离初始观点时,辩论才需完整闭环模型来突破凸包。

原文摘要 · Abstract (English)

Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and accuracy, yet how and why it works is rarely modelled. Such deliberation mirrors how humans reach decisions. As social animals we are pulled both by the group, the herd effect that classical opinion-dynamics models such as DeGroot and Friedkin--Johnsen capture, and by our own internal belief, which they do not. We model multi-agent deliberation as a closed-loop dynamical system in which each agent carries a hidden internal belief, its anchor, that continually pulls its opinion regardless of its neighbours. We show this anchor can be recovered from the deliberation alone, and that it explains a behaviour classical consensus rules forbid: an agent's confidence in the correct answer can climb past where any agent started, escaping the space (convexhull) formed by the initial beliefs. Checking whether the recovered anchor also predicts held-out runs (generalizes) gives a simple test for when a model is truly driven bysuch an anchor. Across three open-weight model families this is a spectrum, not all-or-nothing. All anchors' influence are about equally strongly, but they differ in where the anchor sits, and only when it sits far from the initial opinions does deliberation escape the hull and need the full closed-loop model.

多智能体大模型推理信念建模动态系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。