用反事实方法让多智能体探索更高效,避免重复动作。
Counterfactual Conditional Likelihood Rewards for Multiagent Exploration
- 通过反事实条件似然评估每个智能体的独特贡献。
- 在稀疏奖励任务中加速学习,尤其适合强协作场景。
- 适合需要多智能体紧密配合的复杂环境研究者。
高效探索对多智能体系统发现协同策略至关重要,尤其在搜救或行星探测等开放域中。然而,仅从个体层面鼓励探索常导致冗余,因智能体缺乏对队友探索状态的认知。本文提出反事实条件似然(CCL)奖励,通过分离每个智能体对团队探索的独特贡献来评分。与以往仅奖励个体观测新颖性的方法不同,CCL强调对团队联合探索具有信息量的观测。在连续多智能体领域实验表明,CCL奖励能显著加速稀疏团队奖励环境中的学习,其中多数联合动作回报为零,且在要求智能体间紧密协调的任务中表现尤为出色。
原文摘要 · Abstract (English)
Efficient exploration is critical for multiagent systems to discover coordinated strategies, particularly in open-ended domains such as search and rescue or planetary surveying. However, when exploration is encouraged only at the individual agent level, it often leads to redundancy, as agents act without awareness of how their teammates are exploring. In this work, we introduce Counterfactual Conditional Likelihood (CCL) rewards, which score each agent's exploration by isolating its unique contribution to team exploration. Unlike prior methods that reward agents solely for the novelty of their individual observations, CCL emphasizes observations that are informative with respect to the joint exploration of the team. Experiments in continuous multiagent domains show that CCL rewards accelerate learning for domains with sparse team rewards, where most joint actions yield zero rewards, and are particularly effective in tasks that require tight coordination among agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。