提出快速采样方法,让机器人在复杂不确定环境中长时序规划更高效。
Scaling Long-Horizon Online POMDP Planning via Rapid State Space Sampling
- 用快速采样生成多样宏观动作,指导信念空间搜索
- 在超100步规划任务中性能超越现有方法数倍
- 适合需要长期推理的机器人自主决策场景
部分可观测马尔可夫决策过程(POMDP)是不确定性环境下运动规划的通用且严谨框架。尽管POMDP求解器的可扩展性已有显著提升,但长时程POMDP(如≥15步)仍难以求解。本文提出一种新型近似在线POMDP求解器——基于参考的快速状态空间采样在线规划(ROP-RaS3)。该方法利用全新的高速采样式运动规划技术,快速采样状态空间并在线生成多样化宏观动作,进而用于引导信念空间采样,推断高质量策略,无需穷举动作空间——这是现代在线POMDP求解器的根本限制。ROP-RaS3在多个长时程POMDP问题上进行评估,包括规划时长达100步以上、状态维度达15维且需超过20步前瞻的问题。在所有测试中,其性能均显著优于其他最先进方法,最高提升达数倍。
原文摘要 · Abstract (English)
Partially Observable Markov Decision Processes (POMDPs) are a general and principled framework for motion planning under uncertainty. Despite tremendous improvement in the scalability of POMDP solvers, long-horizon POMDPs (e.g., $\geq15$ steps) remain difficult to solve. This paper proposes a new approximate online POMDP solver, called Reference-Based Online POMDP Planning via Rapid State Space Sampling (ROP-RaS3). ROP-RaS3 uses novel extremely fast sampling-based motion planning techniques to sample the state space and generate a diverse set of macro actions online which are then used to bias belief-space sampling and infer high-quality policies without requiring exhaustive enumeration of the action space -- a fundamental constraint for modern online POMDP solvers. ROP-RaS3 is evaluated on various long-horizon POMDPs, including on a problem with a planning horizon of more than 100 steps and a problem with a 15-dimensional state space that requires more than 20 look ahead steps. In all of these problems, ROP-RaS3 substantially outperforms other state-of-the-art methods by up to multiple folds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。