arXiv:2608.30673cs.RO2026-08

用好奇心驱动探索,提升复杂环境下气体源定位的准确性与效率。

CIG-RL: Curiosity-Driven Information-Guided Reinforcement Learning for Source Term Estimation in Uncertain Environments

论文配图:CIG-RL: Curiosity-Driven Information-Guided Reinforcement Learning for Source Term Estimation in Uncertain Environments
图 1 · 摘自论文原文
  • 引入好奇心机制,主动探索未充分学习的环境状态变化
  • 在高噪声条件下实现90%以上的源位置估计准确率
  • 适合需要实时决策的危险气体监测场景

源项估计(STE)旨在确定气体源的关键属性,对识别有害气体泄漏至关重要。信息论方法因在噪声环境中具有鲁棒性,被用于移动传感器的自主STE,但其在线动作选择计算成本高。深度强化学习(DRL)提供了快速决策的潜力。在基于DRL的STE中,智能体根据从噪声测量序列更新的源项信念状态选择动作。然而,现有方法依赖随机探索或仅基于信念不确定性降低,缺乏有效的探索策略,限制了在噪声环境中的策略鲁棒性。为此,我们提出一种好奇心驱动的信息引导强化学习(CIG-RL),以实现稳健且高效的STE。该方法促进对训练期间未充分探索的新信念状态转移的主动探索,并引入自适应不确定性感知的主动感知奖励,以在不确定性下高效搜索源。在高噪声条件下的仿真和真实世界实验均验证了所提框架的鲁棒性与可行性,凸显其在实际STE问题中的应用潜力。

原文摘要 · Abstract (English)

Source term estimation (STE), which aims to estimate key properties of the gas source, is essential for identifying hazardous gas releases. Information-theoretic approaches have been adopted for autonomous STE using mobile sensors due to robustness in noisy environments, yet their online action selection incurs substantial computational cost. Deep reinforcement learning (DRL) provides a promising alternative with its fast decision-making capability. In DRL-based STE, the agent selects actions based on belief states of the source term updated from noisy measurement sequences. However, existing methods rely on random exploration or solely on belief uncertainty reduction without an effective exploration strategy in DRL, which can limit policy robustness in noisy environments. To address this, we propose a curiosity-driven information-guided reinforcement learning for robust and efficient STE. The proposed method promotes active exploration of novel belief state transitions that have not been sufficiently explored during training. We further introduce an uncertainty-adaptive active perception reward to guide efficient source search under uncertainty. Simulations under high-noise conditions and real-world experiments demonstrate the robustness and feasibility of the proposed framework, highlighting its potential for practical STE problems.

源项估计强化学习主动感知气体检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。