让智能体像人一样预判他人行为,提升多智能体协作效率。
Episodic Future Thinking Mechanism for Multi-agent Reinforcement Learning
- 基于动物认知机制设计未来思维模型,通过多样化策略捕捉不同行为偏好。
- 准确推断其他智能体性格后,预测其动作并模拟未来场景,获得更高奖励。
- 适用于驾驶、粒子环境等复杂多智能体场景,尤其适合性格差异大的社会。
理解多智能体交互中的认知过程是认知科学的核心目标,可引导人工智能向具有社会决策能力的方向发展,涵盖个体异质性带来的不确定性。本文提出一种受动物认知启发的事件式未来思维(Episodic Future Thinking, EFT)机制,用于强化学习智能体。为实现未来思维功能,我们首先构建多角色策略,通过异构策略集合捕捉不同行为特征;智能体的“角色”定义为对奖励分量权重的不同组合,代表不同偏好。未来思维智能体收集目标智能体的观测-动作轨迹,并利用预训练的多角色策略推断其角色。一旦角色被推断,智能体即可预测目标智能体的下一步行动,并模拟潜在未来情景,从而自适应选择最优动作。在包含多样驾驶风格的多智能体自动驾驶场景和多个粒子环境中的实验表明,该机制在准确角色推断下比现有方法获得更高奖励,且在不同角色多样性水平的社会中均保持显著性能提升。
原文摘要 · Abstract (English)
Understanding cognitive processes in multi-agent interactions is a primary goal in cognitive science. It can guide the direction of artificial intelligence (AI) research toward social decision-making in multi-agent systems, which includes uncertainty from character heterogeneity. In this paper, we introduce an episodic future thinking (EFT) mechanism for a reinforcement learning (RL) agent, inspired by cognitive processes observed in animals. To enable future thinking functionality, we first develop a multi-character policy that captures diverse characters with an ensemble of heterogeneous policies. Here, the character of an agent is defined as a different weight combination on reward components, representing distinct behavioral preferences. The future thinking agent collects observation-action trajectories of the target agents and uses the pre-trained multi-character policy to infer their characters. Once the character is inferred, the agent predicts the upcoming actions of target agents and simulates the potential future scenario. This capability allows the agent to adaptively select the optimal action, considering the predicted future scenario in multi-agent interactions. To evaluate the proposed mechanism, we consider the multi-agent autonomous driving scenario with diverse driving traits and multiple particle environments. Simulation results demonstrate that the EFT mechanism with accurate character inference leads to a higher reward than existing multi-agent solutions. We also confirm that the effect of reward improvement remains valid across societies with different levels of character diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。