多机器人社交队形导航中引入内在动机,提升协作探索效率。
Intrinsic-Motivation Multi-Robot Social Formation Navigation with Coordinated Exploration
- 设计自学习内在奖励机制,缓解策略保守性。
- 双采样模式配合两时间尺度更新,提升导航与奖励表征。
- 在社交导航基准上超越现有最优方法,适合多机协同场景。
本文研究强化学习在多机器人社交队形导航中的应用,这是实现人机无缝共存的关键能力。尽管强化学习前景广阔,但行人行为的不可预测性和非合作性给机器人协作探索效率带来巨大挑战。为此,我们提出一种新型的协调探索多机器人强化学习算法,引入内在动机探索机制。其核心是自学习的内在奖励机制,旨在共同缓解策略保守性。此外,该算法在集中训练、分散执行框架内引入双采样模式,增强导航策略与内在奖励的表征能力,并采用两时间尺度更新规则解耦参数更新过程。在社交队形导航基准上的实验证明,所提算法在关键指标上显著优于现有最先进方法。代码与视频演示见:https://github.com/czxhunzi/CEMRRL。
原文摘要 · Abstract (English)
This paper investigates the application of reinforcement learning (RL) to multi-robot social formation navigation, a critical capability for enabling seamless human-robot coexistence. While RL offers a promising paradigm, the inherent unpredictability and often uncooperative dynamics of pedestrian behavior pose substantial challenges, particularly concerning the efficiency of coordinated exploration among robots. To address this, we propose a novel coordinated-exploration multi-robot RL algorithm introducing an intrinsic motivation exploration. Its core component is a self-learning intrinsic reward mechanism designed to collectively alleviate policy conservatism. Moreover, this algorithm incorporates a dual-sampling mode within the centralized training and decentralized execution framework to enhance the representation of both the navigation policy and the intrinsic reward, leveraging a two-time-scale update rule to decouple parameter updates. Empirical results on social formation navigation benchmarks demonstrate the proposed algorithm's superior performance over existing state-of-the-art methods across crucial metrics. Our code and video demos are available at: https://github.com/czxhunzi/CEMRRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。