用社会偏好学习自动优化机器人导航行为,无需人工设计奖励函数。
SPLC: Social Preference Learning for Crowd Robot Navigation

- 通过社会偏好反馈机制自动生成人类行为偏好数据
- 在多个基准测试中显著优于现有方法,提升导航合规性
- 适合需要与人群自然共处的机器人应用开发
离线强化学习在人机共存场景下的群体机器人导航中具有巨大潜力。然而,行人运动的固有复杂性使得设计有效奖励函数以促进符合社会规范的机器人行为始终是一项挑战。本文提出社交偏好学习(SPLC)算法,消除对精细奖励设计的需求。其核心创新在于引入社会偏好反馈机制,通过合理的评估标准自动生成偏好数据。该流程显式考虑行人动态的复杂性,缓解奖励偏差,并系统量化广泛的社会规范,从而促进符合社会规范的行为。大量实验表明,将SPLC与离线强化学习结合,在标准性能指标上持续优于现有最先进方法。此外,基于TurtleBot4的真实世界实验进一步验证了SPLC在实际人机共存场景中的有效性。代码与视频演示见https://github.com/sklus949/SPLC。
原文摘要 · Abstract (English)
Offline reinforcement learning (RL) holds significant potential for crowd robot navigation in human-robot coexistence applications. However, the inherent complexity of pedestrian motion renders the design of effective reward functions for promoting socially compliant robot behaviors a persistent challenge. This paper proposes a Social Preference Learning for Crowd Robot Navigation (SPLC) algorithm to eliminate the need for detailed reward design. Its core innovation lies in the introduction of a social preference feedback mechanism to automatically generate preference data through principled preference evaluation criteria. By explicitly accounting for the intricacies of pedestrian dynamics, the pipeline mitigates the reward bias and facilitates the systematic quantification of broad social norms, thereby fostering socially compliant behaviors. Extensive experiments integrating SPLC with offline RL methods demonstrate consistent improvements over state-of-the-art baselines across standard performance metrics. Furthermore, real-world experiments on the TurtleBot4 further validate the effectiveness of SPLC in practical human-robot coexistence settings. Our code and video demos are available at https://github.com/sklus949/SPLC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。