让大模型通过简短对话提升策略行为的稳定性。
Communication Enhances LLMs' Stability in Strategic Thinking
- 用低成本预沟通模拟廉价谈话,减少模型策略波动。
- 多数模型-上下文组合中,合作轨迹噪声显著降低。
- 适合需要可靠多智能体决策的场景,尤其对高波动模型有效。
大型语言模型在需要策略思考的任务中常表现出明显的上下文依赖性变异,影响多智能体行为的可预测性。针对7至90亿参数的模型,在十轮重复囚徒困境任务中,我们评估了短时、零成本的预沟通(模拟廉价谈话)是否影响策略稳定性。通过模拟层面的自助抽样与非参数推断,比较消息与无消息条件下经LOWESS回归拟合的合作轨迹。结果显示,大多数模型-上下文组合中轨迹噪声显著下降。该稳定效应在多种提示变体和解码方式下持续存在,但强度受模型选择与上下文框架影响,基线波动更高的模型获益最大。尽管沟通极少引发有害不稳定,但存在少数特定上下文下的反例,并识别出通信损害稳定性的有限范围。研究结果表明,廉价谈话式沟通是提升多智能体大模型策略行为可预测性与可靠性的一种低成本实用工具。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often exhibit pronounced context-dependent variability that undermines predictable multi-agent behavior in tasks requiring strategic thinking. Focusing on models that range from 7 to 9 billion parameters in size engaged in a ten-round repeated Prisoner's Dilemma, we evaluate whether short, costless pre-play messages emulating the cheap-talk paradigm affect strategic stability. Our analysis uses simulation-level bootstrap resampling and nonparametric inference to compare cooperation trajectories fitted with LOWESS regression across both the messaging and the no-messaging condition. We demonstrate consistent reductions in trajectory noise across a majority of the model-context pairings being studied. The stabilizing effect persists across multiple prompt variants and decoding regimes, though its magnitude depends on model choice and contextual framing, with models displaying higher baseline volatility gaining the most. While communication rarely produces harmful instability, we document a few context-specific exceptions and identify the limited domains in which communication harms stability. These findings position cheap-talk style communication as a low-cost, practical tool for improving the predictability and reliability of strategic behavior in multi-agent LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。