arXiv:2605.29042cs.AIcs.LG2026-05

通过可微信念动态,让智能体自然影响对手认知,提升协作效率。

Differentiable Belief-based Opponent Shaping

论文配图:Differentiable Belief-based Opponent Shaping
图 1 · 摘自论文原文
  • 将对手信念视为可优化状态,通过梯度反向传播更新策略。
  • 在隐角色游戏中超越PPO和BBM,混合动机场景下提升显著。
  • 无需预设欺骗或合作目标,适合复杂多智能体协作任务。

人类协作常依赖通过策略行为影响他人信念。多智能体强化学习中的对手塑造试图复现此能力,但现有方法通常作用于对手的参数、策略或价值空间。而隐藏角色游戏中的信念操纵技术常依赖硬编码目标,如欺骗或信念饱和。本文提出可微信念驱动的对手塑造(D-BOS),一种一阶方法,将每个观察者的信念视为被塑造的对手状态,并对k步softmax-Bayes信念动态进行微分。不显式奖励欺骗或合作行为,而是以信念状态为目标进行塑造,使最优策略由环境奖励结构自然涌现。该信念空间公式通过反向传播对手信念更新生成塑造信号,且可通过聚合各观察者独立推断的信念轨迹自然扩展至多观察者场景。实验表明,D-BOS在隐角色游戏中优于PPO和BBM,尤其在混合动机设置中表现最佳。

原文摘要 · Abstract (English)

Human coordination often relies on the ability to influence the beliefs of others through strategic action. In multi-agent reinforcement learning, opponent shaping attempts to replicate this influence, though existing methods typically operate within an opponent's parameter, policy, or value space. Meanwhile, belief-manipulation techniques in hidden-role games often rely on hard-coded objectives, such as deception or belief saturation. We propose Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats each observer's belief as the shaped opponent state and differentiates through $k$-step softmax-Bayes belief dynamics. Rather than explicitly rewarding deceptive or cooperative behavior, our method treats the belief state as the target for shaping. This allows the optimal strategy to emerge naturally from the environment's reward structure. This belief-space formulation provides an opponent-shaping signal by differentiating through opponent belief updates, and naturally extends to multiple observers by aggregating gradients over their individual inferred belief trajectories. Empirically, D-BOS outperforms PPO and BBM in hidden-role games, with the largest gains in mixed-motive settings.

多智能体信念塑造强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。