提出自适应意图建模框架,让智能体更精准理解对手策略。
Generalized Intention Modeling in Multi-Agent Reinforcement Learning

- 根据任务动态选择多种意图表征的混合方式
- 新表征与自身未来收益互信息最大,更贴近实际效果
- 在多类任务中超越现有方法,适合竞争性多智能体场景
在非合作、竞争性和一般和的多智能体强化学习中,建模对手意图对有效决策至关重要。现有方法使用预设的轨迹信息(如对手下一步动作或未来环境状态)生成意图嵌入来引导主体行为,但这些信息并非普遍适用于所有任务。我们实证发现意图具有任务和环境依赖性。为此,提出一种任务自适应的对手建模框架,学习多个意图表征的性能驱动混合。进一步引入一种最大化与主体未来回报互信息的新意图表征,以捕捉最直接影响表现的对手信息。该方法在多种任务中持续达到或超过当前最优基线性能,并揭示不同建模策略成功的情境与原因。
原文摘要 · Abstract (English)
Modeling an opponent's intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learning. Existing opponent modeling methods encode intent using an embedding derived from episode information chosen a priori, such as the opponent's next action or a future environment state, and use this to guide the ego-agent's behavior. These approaches assume that the chosen information is universally representative of intent; however, we show empirically that this is not the case as intentions are often task- and environment-dependent. To address this, we introduce a task-adaptive opponent modeling framework that learns a performance-driven mixture of multiple intent representations. We further introduce a new intention representation that maximizes mutual information with the ego-agent's future returns, thereby capturing opponent information that is most directly relevant to performance. Our approach consistently matches or exceeds the performance of state-of-the-art baselines across diverse tasks and yields insights into when and why different opponent modeling strategies succeed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。