让推理图主动思考,比让大模型读图更有效。
Don't Make the LLM Read the Graph: Make the Graph Think

- 用图结构控制行动选择,而非仅作提示上下文
- 强模型在二阶心理理论任务中准确率从20%提至100%
- 适合研究多智能体协作与大模型决策偏差的学者
我们探究显式信念图是否能提升大模型在合作式多智能体推理中的表现。在汉诺伊牌类游戏中,针对四个大模型家族开展3000余次受控实验,得出四项结论:其一,集成架构决定图的价值——作为提示上下文时,仅弱模型受益(二阶心理理论任务中80% vs 10%,p<0.0001,OR=36.0);当图通过排序清单控制行动选择时,即使强模型也需依赖该结构(100% vs 20%,p<0.001)。其二,发现“规划者违抗”现象:模型在部分能力水平下会忽略正确规划建议(90%违抗,重复验证N=20),Gemini近乎无违抗,而Llama 70B达90%;模型区分事实信息(服从)与建议(拒绝)。其三,完整对局证据显示,跨智能体惯例(+128%,p=0.003)优于所有单智能体干预,且各图组件需组合使用才有效。其四,初步扩展分析(每组N=10,探索性)表明图深度收益递减:浅层图性价比最优,深层心理理论图在五人场景下反而有害(-1.5分,p=0.029)。
原文摘要 · Abstract (English)
We investigate whether explicit belief graphs improve LLM performance in cooperative multi-agent reasoning. Through 3,000+ controlled trials across four LLM families in the cooperative card game Hanabi, we establish four findings. First, integration architecture determines whether belief graphs provide value: as prompt context, graphs are decorative for strong models and beneficial only for weak models on 2nd-order Theory of Mind (80% vs 10%, p<0.0001, OR=36.0); when graphs gate action selection through ranked shortlists, they become structurally essential even for strong models (100% vs 20% on 2nd-order ToM, p<0.001). Second, we identify "Planner Defiance," a model-family-specific failure where LLMs override correct planner recommendations at partial competence (90% override, replicated N=20); Gemini models show near-zero defiance while Llama 70B shows 90%, and models distinguish factual context (deferred to) from advisory recommendations (overridden). Third, full-game evidence confirms inter-agent conventions (+128% over baseline, p=0.003) outperform all single-agent interventions, and individual belief-graph components must be combined to produce gains. Fourth, preliminary scaling analysis (N=10/cell, exploratory) suggests graph depth has diminishing returns: shallow graphs provide the best cost-benefit ratio, while deeper ToM graphs appear harmful at larger player counts (-1.5 pts at 5-player, p=0.029).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。