arXiv:2409.16392cs.AIcs.LG2024-09ICRA被引 2

用更少粒子实现更准定位,提升复杂环境下的决策效率

Rao-Blackwellized POMDP Planning

  • 将瑞利-布莱克韦尔化技术用于信念更新和在线规划
  • 相同计算资源下,粒子数减少一半仍保持更高定位精度
  • 适合需要高效推理的机器人导航、自动驾驶等场景

部分可观测马尔可夫决策过程(POMDP)为不确定性下的决策提供了结构化框架,但其应用依赖高效的信念更新。序列重要性重采样粒子滤波器(SIRPF)常用于大型近似POMDP求解器中的信念更新,但面临粒子匮乏和状态维度增加时计算成本升高等问题。为此,本文提出瑞利-布莱克韦尔化POMDP(RB-POMDP)近似求解器,并给出了在信念更新与在线规划中应用瑞利-布莱克韦尔化的通用方法。我们在一个模拟定位任务中对比了SIRPF与瑞利-布莱克韦尔化粒子滤波器(RBPF)的表现,该任务中智能体在无GPS环境下向目标移动,使用POMCPOW与RB-POMCPOW规划器。结果表明,RBPF在更少粒子条件下仍能长期维持准确的信念近似;更意外的是,在相同计算限制下,结合积分求积法的RBPF显著提升了规划质量,优于基于SIRPF的规划。

原文摘要 · Abstract (English)

Partially Observable Markov Decision Processes (POMDPs) provide a structured framework for decision-making under uncertainty, but their application requires efficient belief updates. Sequential Importance Resampling Particle Filters (SIRPF), also known as Bootstrap Particle Filters, are commonly used as belief updaters in large approximate POMDP solvers, but they face challenges such as particle deprivation and high computational costs as the system's state dimension grows. To address these issues, this study introduces Rao-Blackwellized POMDP (RB-POMDP) approximate solvers and outlines generic methods to apply Rao-Blackwellization in both belief updates and online planning. We compare the performance of SIRPF and Rao-Blackwellized Particle Filters (RBPF) in a simulated localization problem where an agent navigates toward a target in a GPS-denied environment using POMCPOW and RB-POMCPOW planners. Our results not only confirm that RBPFs maintain accurate belief approximations over time with fewer particles, but, more surprisingly, RBPFs combined with quadrature-based integration improve planning quality significantly compared to SIRPF-based planning under the same computational limits.

强化学习贝叶斯推理机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。