arXiv:2510.12215cs.RO2025-10

用正负示范与规则结合训练机器人安全自适应导航

Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications

  • 融合正负示范学习密度奖励,再叠加避障和抵达目标的规则约束
  • 在仿真与真实场景中,成功率与效率均优于基线方法
  • 适合需要实时性与安全性的移动机器人导航应用

动态人类环境中移动机器人导航需在适应多样行为与遵守安全约束之间取得平衡。我们提出将数据驱动奖励与规则化目标结合,以实现更优的适应性与安全性平衡。具体方法是:从正负示范中学习密度基础奖励,并加入避障与到达目标的规则约束;通过采样式前瞻控制器生成安全且自适应的监督动作,再将其压缩为轻量级学生策略,支持实时运行并提供不确定性估计。在合成环境及电梯共乘仿真中,该方法在成功率与时间效率上持续优于基线;真实人类参与者测试进一步验证了其部署可行性。视频展示见项目主页。

原文摘要 · Abstract (English)

Mobile robot navigation in dynamic human environments requires policies that balance adaptability to diverse behaviors with compliance to safety constraints. We hypothesize that integrating data-driven rewards with rule-based objectives enables navigation policies to achieve a more effective balance of adaptability and safety. To this end, we develop a framework that learns a density-based reward from positive and negative demonstrations and augments it with rule-based objectives for obstacle avoidance and goal reaching. A sampling-based lookahead controller produces supervisory actions that are both safe and adaptive, which are subsequently distilled into a compact student policy suitable for real-time operation with uncertainty estimates. Experiments in synthetic and elevator co-boarding simulations show consistent gains in success rate and time efficiency over baselines, and real-world demonstrations with human participants confirm the practicality of deployment. A video illustrating this work can be found on our project page https://chanwookim971024.github.io/PioneeR/.

机器人导航强化学习规则约束安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。