arXiv:2410.12232cs.ROcs.AI2024-10ICRA

用信息论提升强化学习机器人避障多样性,更好应对未知人群行为。

Improving the Generalization of Unseen Crowd Behaviors for Reinforcement Learning based Local Motion Planners

  • 单策略内通过信息最大化增强智能体行为多样性。
  • 在未知人群场景中碰撞率更低,且不增加时间和路程。
  • 适合需要安全应对复杂人流的移动机器人部署场景。

在有人群的环境中部署安全移动机器人面临挑战,因行人行为难以预测。现有基于强化学习的运动规划器依赖单一策略模拟行人,易出现过拟合。而将避障问题建模为多智能体框架时,由于智能体同质性,可能导致与真实行人冲突。为此,本文提出一种高效方法:在单个策略内通过最大化信息论目标增强智能体多样性,丰富经验,提升对未见人群行为的适应能力。评估时,设计了受行人行为启发的多样化场景。所提出的条件化行为策略在这些挑战性场景中表现更优,显著降低潜在碰撞,且无需额外时间或行程。

原文摘要 · Abstract (English)

Deploying a safe mobile robot policy in scenarios with human pedestrians is challenging due to their unpredictable movements. Current Reinforcement Learning-based motion planners rely on a single policy to simulate pedestrian movements and could suffer from the over-fitting issue. Alternatively, framing the collision avoidance problem as a multi-agent framework, where agents generate dynamic movements while learning to reach their goals, can lead to conflicts with human pedestrians due to their homogeneity. To tackle this problem, we introduce an efficient method that enhances agent diversity within a single policy by maximizing an information-theoretic objective. This diversity enriches each agent's experiences, improving its adaptability to unseen crowd behaviors. In assessing an agent's robustness against unseen crowds, we propose diverse scenarios inspired by pedestrian crowd behaviors. Our behavior-conditioned policies outperform existing works in these challenging scenes, reducing potential collisions without additional time or travel.

强化学习机器人避障群体行为多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。