arXiv:2609.05569cs.RO2026-09

用强化学习提升无人机寻气精度,兼顾安全与信息获取。

Information-Guided Safe Reinforcement Learning for Autonomous Gas Source Localization using sUAS

论文配图:Information-Guided Safe Reinforcement Learning for Autonomous Gas Source Localization using sUAS
图 1 · 摘自论文原文
  • 结合可观测性矩阵与深度强化学习,动态规划高信息探测路径。
  • 在复杂移动气源场景下实现近80%定位成功率,远超传统方法(约30%)。
  • 适合需要高安全性与自主探测能力的工业巡检场景。

利用小型无人机系统(sUAS)自主定位泄漏气体属于根本性病态逆问题。在湍流大气边界层中,高度间歇性的标量浓度场违背经典梯度导航假设,导致数据驱动估计算法受严重噪声和虚假局部极小值影响。为此,我们提出一种信息引导的安全强化学习框架,基于自定义的GPU加速三维仿真环境,耦合欧拉风场求解器与拉格朗日烟团扩散模型。我们发现确定性信息探索规划器存在关键缺陷——由格拉姆矩阵偏差引发的贪婪行为,使早期错误估计导致空间多样性不足。为系统打破此退化,架构融合经典经验可观测性格拉姆矩阵(EMGR)规划器与学习型软演员-评论家(SAC)探索策略。一个确定性元监督器通过KL散度实时监测估计算法可靠性,动态融合确定性利用与学习型探索,引导sUAS进入高信息区域。通过渐进式课程训练并受严格鲁棒控制屏障函数(RCBF)保护,该强化学习框架在复杂移动源上实现近80%的定位成功率,显著优于经典基线(约30%),且零安全违规。

原文摘要 · Abstract (English)

The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes a fundamentally ill-posed inverse problem. In turbulent atmospheric boundary layers, highly intermittent scalar concentration fields violate the assumptions of classical gradient-based navigation, causing data-driven estimators to suffer from severe noise and spurious local minima. To address these challenges, we introduce an Information-Guided Safe Reinforcement Learning framework evaluated within a custom, GPU-accelerated 3D simulation environment coupling an Eulerian wind solver with a Lagrangian puff dispersion model. We identify a critical vulnerability in deterministic information-seeking planners - a Gramian bias where agents act greedily upon flawed early estimates, starving the estimator of spatial diversity. To systematically break this degeneracy, our architecture integrates a classical empirical observability Gramian (EMGR) planner with a learned Soft Actor-Critic (SAC) exploratory policy. A deterministic meta-supervisor actively monitors estimator reliability via Kullback-Leibler (KL) divergence, dynamically blending deterministic exploitation with learned exploration to steer the sUAS into high-information zones. Trained via a progressive curriculum and safeguarded by a strictly enforced Robust Control Barrier Function (RCBF), our RL framework achieves nearly 80% localization success on complex, mobile sources - drastically outperforming classical baselines (~30%) - while ensuring zero safety violations.

无人机巡检强化学习安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。