arXiv:2508.09971cs.ROcs.AI2025-08被引 2

无人机沿河自主飞行,用语义动态模型提升安全与效率。

Vision-driven River Following of UAV via Safe Reinforcement Learning using Semantic Dynamics Model

  • 用滑动窗口历史回报优化奖励优势,解决河流路径依赖问题。
  • 语义动态模型预测更准确,短期状态预测误差降低32%。
  • 集成安全约束的强化学习框架,适合高风险导航场景。

无人机视觉驱动的自主沿河飞行在救援、监控和环境监测中至关重要,尤其在GPS信号弱的密集河网区域。此类任务需满足严格安全约束并优化性能,但奖励具有历史依赖性(非马尔可夫),传统安全强化学习难以应对。为此,本文提出三项贡献:首先,引入边际收益优势估计(MGAE),通过滑动窗口计算的历史回溯基线优化优势函数,匹配非马尔可夫动态;其次,构建基于补丁化水体语义掩码的语义动态模型(SDM),相比隐空间视觉动态模型,实现更可解释且数据高效的短期未来观测预测;第三,提出受限动作动态估计器(CADE)架构,整合策略、代价估计器与SDM,用于代价优势估计,形成基于模型的安全强化学习框架。仿真结果表明,MGAE收敛更快,性能优于传统基于评判器的方法(如广义优势估计);SDM提供更精准的短期状态预测,使代价估计器能更好预判潜在违规。整体上,CADE有效将安全约束融入基于模型的RL中,拉格朗日方法在训练中实现奖励与安全的‘软’平衡,而安全层在推理时施加‘硬’动作覆盖,确保安全性。

原文摘要 · Abstract (English)

Vision-driven autonomous river following by Unmanned Aerial Vehicles is critical for applications such as rescue, surveillance, and environmental monitoring, particularly in dense riverine environments where GPS signals are unreliable. These safety-critical navigation tasks must satisfy hard safety constraints while optimizing performance. Moreover, the reward in river following is inherently history-dependent (non-Markovian) by which river segment has already been visited, making it challenging for standard safe Reinforcement Learning (SafeRL). To address these gaps, we propose three contributions. First, we introduce Marginal Gain Advantage Estimation, which refines the reward advantage function by using a sliding window baseline computed from historical episodic returns, aligning the advantage estimate with non-Markovian dynamics. Second, we develop a Semantic Dynamics Model based on patchified water semantic masks offering more interpretable and data-efficient short-term prediction of future observations compared to latent vision dynamics models. Third, we present the Constrained Actor Dynamics Estimator architecture, which integrates the actor, cost estimator, and SDM for cost advantage estimation to form a model-based SafeRL framework. Simulation results demonstrate that MGAE achieves faster convergence and superior performance over traditional critic-based methods like Generalized Advantage Estimation. SDM provides more accurate short-term state predictions that enable the cost estimator to better predict potential violations. Overall, CADE effectively integrates safety regulation into model-based RL, with the Lagrangian approach providing a "soft" balance between reward and safety during training, while the safety layer enhances inference by imposing a "hard" action overlay.

无人机导航安全强化学习语义建模路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。