用单目摄像头实现复杂人行道长期导航,提升避障与社交合规性。
From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation

- 以流动匹配作为动作表征,从大规模数据预训练导航策略。
- 在仿真中达42%成功率、66%路线完成率,真实场景误拦率降40%。
- 结合人类偏好反馈微调,增强复杂场景下的推理与社会行为适应性。
自主长时序人行道导航对微型移动应用(如机器人送餐、辅助电动轮椅)至关重要。与道路自动驾驶不同,人行道导航需在不可预测的地形和行人环境中实现精确操控,且感知系统极简,仅需单目RGB相机。尽管模仿学习(IL)提供了实用方案,但生成的自动导航策略常出现误差累积、缺乏社交合规性,且在复杂情境下缺乏反事实推理能力。为此,我们提出FlowPilot——一种无需地图的导航策略,仅依赖单目相机实现稳健高效的长时序导航。首先,采用锚定流匹配作为动作表征,在大规模机器人车队数据上进行策略预训练,以捕捉人行道导航行为的多样性和多模态分布。为进一步弥合模仿与对齐之间的差距,设计了一种人机协同偏好学习机制,利用少量人工干预数据微调模型,强化其反事实推理能力和人行道社交合规性。我们在多样化的仿真与真实环境进行了评估,结果显示:FlowPilot在仿真中达成42%的成功率与66%的路线完成率;而引入人类偏好后的FlowPilot-HP显著提升真实场景鲁棒性与社交合规性,相比基线模型,红外检测误拦率(IR)降低40.0%,非红外误拦率(NIR)降低52.1%。
原文摘要 · Abstract (English)
Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs. Unlike autonomous driving on the road, long-horizon sidewalk navigation requires precise maneuvering through unpredictable sidewalk terrains and pedestrians, with a lightweight perception stack as minimal as a single monocular RGB camera. While imitation learning (IL) from demonstrations offers a practical solution, the resulting autopilot policy often suffers from compounding errors, a lack of social compliance on sidewalks, and deficiencies in counterfactual reasoning to handle complex situations. To address these challenges, we introduce FlowPilot, a mapless navigation policy that achieves robust and efficient long-horizon navigation performance using only a monocular RGB camera. We first propose to use anchored flow matching as an action representation for policy pre-training on large-scale robot fleet data and to capture the diverse, complex, multimodal distribution of sidewalk navigation behaviors. To bridge the gap between imitation and alignment, we further design a human-in-the-loop preference learning scheme to tune the policy on a small amount of human intervention data. It strengthens the model's counterfactual reasoning and social compliance on sidewalks. We evaluate FlowPilot through extensive simulation and real-world experiments in diverse sidewalk environments. FlowPilot achieves 42% success rate and 66% route completion in simulation, while FlowPilot-HP further improves real-world robustness and social compliance, reducing IR by 40.0% and NIR by 52.1% relative to the base model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。