用奖励塑形与安全函数,让无人机零样本高效又安全飞行
Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions

- 结合奖励塑形与安全函数,让无人机在复杂环境自适应导航
- 任务时间减少显著,复杂场景下仍保持高成功率
- 无需重新训练,适合对安全与效率要求高的实际应用
自主导航与避障仍是现代无人飞行器(UAV)的核心挑战。传统控制方法难以应对环境的复杂性和多变性,而强化学习(RL)虽能通过环境交互学习自适应行为,但现有研究往往以任务成功为首要目标,忽视任务耗时与飞行安全。本研究将基于势能的奖励塑形(PBRS)与控制李雅普诺夫函数(CLF)及控制屏障函数(CBF)相结合,同时优化任务时间并提供形式化安全保证。一个RL模型在通用简化环境中训练后,可直接应用于复杂场景,仅需搭配CLF-CBF-QP滤波器,无需进一步训练。仿真环境中的实验结果表明,任务时间显著缩短,且在复杂环境中表现出优异性能。
原文摘要 · Abstract (English)
Autonomous navigation and obstacle avoidance remain a core challenge of modern Unmanned Aerial Vehicles (UAVs). While traditional control methods struggle with the complexity and variability of the environment, reinforcement learning (RL) enables UAVs to learn adaptive behaviors through interaction with the environment. Existing research with RL prioritizes the mission success at the expense of mission time and safety of UAVs. This study integrates Potential Based Reward Shaping (PBRS) with Control Lyapunov Functions (CLF) and Control Barrier Functions (CBF) to simultaneously optimize mission time and ensure formal safety guarantees. An RL model is trained in a generalized simple environment, then used in complex scenarios incorporating a CLF-CBF-QP filter without further training. Experimental results in simulated environments demonstrate a significant reduction in mission time and outstanding performance in complex environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。