arXiv:2508.01539cs.RO2025-08被引 9

让机器人学会像人一样导航,用离线数据训练视觉奖励模型。

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

  • 从人类偏好中学习导航奖励,通过偏好排序优化视觉奖励函数。
  • 实测在未知环境和硬件上表现优异,成功率提升33.3%,轨迹更短更顺。
  • 适用于各类导航框架,适合想提升机器人自主导航能力的研究者。

本文提出HALO,一种新型离线奖励学习算法,将人类导航直觉转化为基于视觉的机器人导航奖励函数。HALO从移动机器人采集的专家轨迹中学习奖励模型,在训练中对参考动作周围动作进行均匀采样,并利用以优选动作为中心的Boltzmann分布生成偏好分数,结合二元用户反馈对直观导航问题进行奖励塑造。采用Plackett-Luce损失训练奖励模型以对齐偏好排序。为验证其有效性,我们将该奖励模型应用于两个下游任务:(i) 直接基于HALO奖励训练的离线策略;(ii) 将HALO奖励作为额外代价项的模型预测控制(MPC)规划器。结果表明,HALO在学习型与经典导航框架中均具高度通用性。我们在Clearpath Husky机器人上进行真实场景部署,证明使用HALO训练的策略能有效泛化至未见环境与硬件配置。相比现有视觉导航方法,HALO实现至少33.3%的成功率提升,轨迹长度降低12.9%,与人类专家轨迹的Frechet距离减少26.6%。

原文摘要 · Abstract (English)

In this paper, we introduce HALO, a novel Offline Reward Learning algorithm that quantifies human intuition in navigation into a vision-based reward function for robot navigation. HALO learns a reward model from offline data, leveraging expert trajectories collected from mobile robots. During training, actions are uniformly sampled around a reference action and ranked using preference scores derived from a Boltzmann distribution centered on the preferred action, and shaped based on binary user feedback to intuitive navigation queries. The reward model is trained via the Plackett-Luce loss to align with these ranked preferences. To demonstrate the effectiveness of HALO, we deploy its reward model in two downstream applications: (i) an offline learned policy trained directly on the HALO-derived rewards, and (ii) a model-predictive-control (MPC) based planner that incorporates the HALO reward as an additional cost term. This showcases the versatility of HALO across both learning-based and classical navigation frameworks. Our real-world deployments on a Clearpath Husky across diverse scenarios demonstrate that policies trained with HALO generalize effectively to unseen environments and hardware setups not present in the training data. HALO outperforms state-of-the-art vision-based navigation methods, achieving at least a 33.3% improvement in success rate, a 12.9% reduction in normalized trajectory length, and a 26.6% reduction in Frechet distance compared to human expert trajectories.

机器人导航离线学习人类偏好奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。