用强化学习与粒子滤波结合,实现不确定环境下钻井路径的智能决策。
Decision-Driven Geosteering Under Uncertainty: A Unified Framework for Sequential Decision Optimization

- 融合粒子滤波与价值型强化学习,实时更新地质不确定性并指导钻井决策
- 三种算法在相同条件下测试,双深度强化学习表现最优且路径更平滑
- 适用于油气勘探中复杂地质条件下的自动化钻井决策,尤其适合工业仿真验证
地质导向需在未知地层中导航钻井轨迹,并根据钻探过程中获取的间接测量数据动态调整决策。本文提出一种考虑不确定性的地质导向统一框架,将粒子滤波(PF)用于概率性地下结构解释,结合基于价值的强化学习进行序列决策优化。钻头前方的地质不确定性通过粒子滤波显式建模,实现基于信念的控制而非确定性纠偏。该框架将粒子滤波信念更新与信念驱动的决策策略耦合,评估三种在相同不确定性表示下运行的决策方案:可解释的近似动态规划(ADP)、深度Q学习基线和采用目标网络训练稳定性的双深度强化学习(Dual DRL),后者使用双重(价值/优势)分解参数化Q值。除最终位置性能外,还引入以稳定性为导向的度量指标,量化决策策略随时间演化的平稳性,提供操作层面的额外洞察。框架集成API,在工业级地质导向模拟器中进行验证,涵盖真实测量噪声和钻探约束。所有方法使用相同的地质实现、操作限制和奖励定义,实现对不同决策策略在整个钻井过程中的行为进行高保真、受控评估,而非仅关注最终井位。
原文摘要 · Abstract (English)
Geosteering requires navigating a well trajectory through an unknown geological configuration, while sequentially updating decisions based on indirect measurements acquired during drilling. This work presents an uncertainty-aware geosteering framework that tightly integrates particle filtering for probabilistic subsurface interpretation with value-based reinforcement learning for sequential decision-making. Geological uncertainty ahead of the drill bit is represented explicitly through a particle filter (PF), enabling belief-informed control rather than deterministic trajectory correction. The framework couples PF belief updates with belief-informed decision policies and evaluates three decision-making options that operate under identical uncertainty representations: an interpretable Approximate Dynamic Programming (ADP) scheme, a Deep Q-learning baseline, and a Dual Deep Reinforcement Learning (Dual DRL) architecture trained with a target Q-network scheme for stability, using a dueling (value/advantage) decomposition for Q-value parameterization. Beyond final placement performance, we assess policy behavior using stability-oriented metrics that quantify steering smoothness over time, providing additional operational insight into how decision policies respond as uncertainty evolves. The framework is integrated with an API for validation within an industrial geosteering simulator under realistic measurement noise and drilling constraints. Using identical geological realizations, operational limits, and reward definitions across methods, the experiments provide a controlled and high-fidelity evaluation of how alternative decision policies behave throughout the drilling process, rather than evaluating performance solely from the final well trajectory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。