动态优化不确定环境下的决策,实时提升规划精度。
Anytime Incremental $ρ$POMDP Planning in Continuous Spaces
- 在线动态细化信念表示,支持持续优化
- 计算成本降低数量级,实现在连续空间高效求解
- 适合需持续收集信息的机器人任务,如探索与导航
部分可观测马尔可夫决策过程(POMDP)为自动驾驶和机器人探索等不确定性决策任务提供了稳健框架。其扩展形式ρPOMDP引入依赖信念的奖励,可显式建模不确定性。现有连续空间的在线ρPOMDP求解器依赖固定信念表示,限制了适应性与精细化能力,不利于信息采集类任务。本文提出ρPOMCPOW,一种任意时间求解器,可动态精炼信念表示,并提供随时间改善的理论保证。为缓解信念相关奖励更新的高计算开销,我们提出一种新颖的增量计算方法,适用于常见熵估计器,实现计算成本数量级降低。实验表明,ρPOMCPOW在效率与解质量上均优于当前最先进方法。
原文摘要 · Abstract (English)
Partially Observable Markov Decision Processes (POMDPs) provide a robust framework for decision-making under uncertainty in applications such as autonomous driving and robotic exploration. Their extension, $ρ$POMDPs, introduces belief-dependent rewards, enabling explicit reasoning about uncertainty. Existing online $ρ$POMDP solvers for continuous spaces rely on fixed belief representations, limiting adaptability and refinement - critical for tasks such as information-gathering. We present $ρ$POMCPOW, an anytime solver that dynamically refines belief representations, with formal guarantees of improvement over time. To mitigate the high computational cost of updating belief-dependent rewards, we propose a novel incremental computation approach. We demonstrate its effectiveness for common entropy estimators, reducing computational cost by orders of magnitude. Experimental results show that $ρ$POMCPOW outperforms state-of-the-art solvers in both efficiency and solution quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。