用人类偏好反馈让无人机在河流上更安全地自主导航。
Deployable Vision-driven UAV River Navigation via Human-in-the-loop Preference Alignment
- 通过人类实时纠正与偏好对比,动态优化飞行策略。
- 仅5次人类干预就达到最高奖励且结果最稳定。
- 适合需要小样本安全适应的无人机实际应用。
河流是环境监测与灾害响应的关键通道,视觉驱动的无人机可快速低成本覆盖。但部署中模拟训练的策略面临分布偏移和安全风险,需高效利用有限的人类干预进行适应。本文研究人机协同学习,引入保守监督者对不安全或低效动作进行否决,并通过比较智能体提案与修正操作提供逐状态偏好。提出状态级混合偏好对齐方法(SPAR-H),融合直接偏好优化与基于奖励的路径:从相同偏好中训练即时奖励估计器,并用信任区域代理更新策略。仅需五次人机协同轨迹采集,SPAR-H 在最终每回合奖励与不同初始条件下方差方面均优于对比方法。学习到的奖励模型与人类偏好一致,能提升未干预动作质量,支持改进的稳定传播。在真实世界验证了持续偏好对齐在河岸追踪任务中的可行性。双状态级偏好实证表明其为数据高效在线适应的有效路径。
原文摘要 · Abstract (English)
Rivers are critical corridors for environmental monitoring and disaster response, where Unmanned Aerial Vehicles (UAVs) guided by vision-driven policies can provide fast, low-cost coverage. However, deployment exposes simulation-trained policies with distribution shift and safety risks and requires efficient adaptation from limited human interventions. We study human-in-the-loop (HITL) learning with a conservative overseer who vetoes unsafe or inefficient actions and provides statewise preferences by comparing the agent's proposal with a corrective override. We introduce Statewise Hybrid Preference Alignment for Robotics (SPAR-H), which fuses direct preference optimization on policy logits with a reward-based pathway that trains an immediate-reward estimator from the same preferences and updates the policy using a trust-region surrogate. With five HITL rollouts collected from a fixed novice policy, SPAR-H achieves the highest final episodic reward and the lowest variance across initial conditions among tested methods. The learned reward model aligns with human-preferred actions and elevates nearby non-intervened choices, supporting stable propagation of improvements. We benchmark SPAR-H against imitation learning (IL), direct preference variants, and evaluative reinforcement learning (RL) in the HITL setting, and demonstrate real-world feasibility of continual preference alignment for UAV river following. Overall, dual statewise preferences empirically provide a practical route to data-efficient online adaptation in riverine navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。