用人类对虚拟路径的偏好,让机器人更安全避障。
CHOP: Counterfactual Human Preference Labels Improve Obstacle Avoidance in Visuomotor Navigation Policies
- 通过对比虚拟路径让人类标注避障偏好,指导机器人学习
- 实测避障失误减少49.7%,路径偏差降低45.0%
- 适合做安全导航的机器人系统,尤其重视人机协同
视觉-运动导航策略在具身智能体中展现出强大的感知-动作耦合能力,但在复杂真实环境中的安全导航与动态障碍物规避方面仍面临挑战。本文提出CHOP,一种利用反事实人类偏好标签来对齐视觉-运动导航策略与人类安全直觉的方法。在CHOP中,针对每组视觉观测,将机器人实际执行的轨迹与其他可能轨迹(反事实轨迹)并列,由人类标注员基于碰撞风险、路径效率等预期结果进行成对偏好判断。聚合后的偏好标签用于微调导航策略,使其行为更贴近人类偏好。在SCAND数据集上的实验表明,使用CHOP微调的策略相比预训练基线,近碰撞事件减少49.7%,路径偏离人类偏好评分下降45.0%,平均障碍物净空距离提升19.8%。该方法在Ghost Robotics Vision60四足机器人上实现真实部署,平均目标成功率提高24.4%,最小障碍物净空增加6.8%,碰撞与干预事件减少45.7%,路径完成度提升38.6%。结果表明,反事实偏好监督在弥合大规模视觉-运动策略与人类对齐的安全导航之间具有显著价值。
原文摘要 · Abstract (English)
Visuomotor navigation policies have shown strong perception-action coupling for embodied agents, yet they often struggle with safe navigation and dynamic obstacle avoidance in complex real-world environments. We introduce CHOP, a novel approach that leverages Counterfactual Human Preference Labels to align visuomotor navigation policies towards human intuition of safety and obstacle avoidance in navigation. In CHOP, for each visual observation, the robot's executed trajectory is included among a set of counterfactual navigation trajectories: alternative trajectories the robot could have followed under identical conditions. Human annotators provide pairwise preference labels over these trajectories based on anticipated outcomes such as collision risk and path efficiency. These aggregated preferences are then used to fine-tune visuomotor navigation policies, aligning their behavior with human preferences in navigation. Experiments on the SCAND dataset show that visuomotor navigation policies fine-tuned with CHOP reduce near-collision events by 49.7%, decrease deviation from human-preferred trajectories by 45.0%, and increase average obstacle clearance by 19.8% on average across multiple state-of-the-art models, compared to their pretrained baselines. These improvements transfer to real-world deployments on a Ghost Robotics Vision60 quadruped, where CHOP-aligned policies improve average goal success rates by 24.4%, increase minimum obstacle clearance by 6.8%, reduce collision and intervention events by 45.7%, and improve normalized path completion by 38.6% on average across navigation scenarios, compared to their pretrained baselines. Our results highlight the value of counterfactual preference supervision in bridging the gap between large-scale visuomotor policies and human-aligned, safety-aware embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。