在障碍物相关性环境中,用贝叶斯方法高效规划路径。
Stochastic Path Planning in Correlated Obstacle Fields
- 用高斯随机场建模障碍物空间相关性,实时更新阻塞概率。
- 两阶段学习框架提升路径规划效率,支持完整代价分布学习。
- 适合应对传感器噪声大、障碍物成簇的复杂导航场景。
我们提出随机相关障碍物场景(SCOS)问题,即在具有空间相关性的不确定阻塞障碍物环境中,面对受限且带噪声的传感器及高成本的确认操作。通过高斯随机场(GRF)建模空间相关性,设计贝叶斯信念更新机制以优化阻塞概率,并据此缩小搜索空间提升效率。为求最优遍历策略,提出新颖的两阶段学习框架:离线阶段通过增强信息奖励的乐观策略迭代学习鲁棒基础策略;在线阶段通过周期性贝叶斯更新基础策略,实现信息适应。该框架支持蒙特卡洛点估计与分布强化学习(RL),可学习完整代价分布,增强不确定性量化能力。理论分析证明了相关性感知更新的优势及后验采样下的收敛性。在不同障碍密度与传感器能力下进行充分实验,性能持续优于基线。该框架适用于存在对抗性干扰或聚集自然危害的导航环境。
原文摘要 · Abstract (English)
We introduce the Stochastic Correlated Obstacle Scene (SCOS) problem, a navigation setting with spatially correlated obstacles of uncertain blockage status, realistically constrained sensors that provide noisy readings and costly disambiguation. Modeling the spatial correlation with Gaussian Random Field (GRF), we develop Bayesian belief updates that refine blockage probabilities, and use the posteriors to reduce search space for efficiency. To find the optimal traversal policy, we propose a novel two-stage learning framework. An offline phase learns a robust base policy via optimistic policy iteration augmented with information bonus to encourage exploration in informative regions, followed by an online rollout policy with periodic base updates via a Bayesian mechanism for information adaptation. This framework supports both Monte Carlo point estimation and distributional reinforcement learning (RL) to learn full cost distributions, leading to stronger uncertainty quantification. We establish theoretical benefits of correlation-aware updating and convergence property under posterior sampling. Comprehensive empirical evaluations across varying obstacle densities, sensor capabilities demonstrate consistent performance gains over baselines. This framework addresses navigation challenges in environments with adversarial interruptions or clustered natural hazards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。