让多个大模型协作更稳定,通过熵正则化选最优协同方式。
DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination

- 引入熵正则化均衡机制,动态选择最佳协作模式。
- 在11个任务中平均提升4.3~8.5个百分点准确率。
- 适合需要多模型稳定协作的推理与规划场景。
多智能体大语言模型系统常无法稳定超越单个强模型加最佳N采样的表现。我们指出其核心问题是协调均衡选择不当:现有系统规定信息共享方式,但未明确应采用何种协作规范。我们将此类系统形式化为折扣不完全信息马尔可夫博弈,并证明两种常见病态——不同规范间的振荡与漂移——均会导致学习不稳定和线性贝叶斯后悔。为此提出异质量化响应均衡(HQRE),一种具有智能体与状态相关温度的熵正则化均衡概念。在单调性条件下,HQRE唯一且支持镜像更新线性收敛,实现有界贝叶斯后悔;同一条件还导出可回溯测量的稳定性诊断。我们在两种算法中实现该目标:DICE-PC通过提示控制协调冻结模型,DICE-FT则进行参数高效镜像微调。在四个领域的十一个基准上,DICE在准确率-成本权衡上优于同类强基线;在推理与规划任务中,DICE-PC平均提升4.3个百分点,DICE-FT提升8.5个百分点。
原文摘要 · Abstract (English)
Multi-agent large language model (LLM) systems often fail to reliably outperform a single strong model equipped with best-of-N sampling. We argue that a core source of this instability is ill-posed equilibrium selection: current systems specify what information agents share, but not which coordination convention should be selected. We formalize a broad class of such systems as discounted incomplete-information Markov games and show that two common pathologies, oscillation between competing conventions and drift across them, can both induce unstable learning and linear Bayesian regret. To obtain a well-posed target, we introduce the Heterogeneous Quantal Response Equilibrium (HQRE), an entropy-regularized equilibrium concept with agent- and state-dependent temperatures. Under a monotonicity condition, HQRE is unique, admits linearly convergent mirror updates, and yields bounded Bayesian regret; the same condition yields rollout-measurable stability diagnostics. We instantiate this objective in two algorithms: DICE-PC, which coordinates frozen models through prompt-control actions, and DICE-FT, which performs parameter-efficient mirror fine-tuning. Across eleven benchmarks in four domains, DICE improves accuracy-cost trade-offs over strong within-class baselines; on reasoning and planning tasks, DICE-PC improves by 4.3 percentage points on average and DICE-FT by 8.5 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。