将主动SLAM建模为部分可观测马尔可夫决策过程,求解近优策略
SLAM as a Stochastic Control Problem with Partial Information: Optimal Solutions and Rigorous Approximations

- 从随机控制视角重构主动SLAM,引入几何感知的探索代价函数
- 在一般假设下建立正则性条件,推导出近最优的逼近解
- 通过数值实验验证学习算法可获得接近最优的导航策略
同时定位与地图构建(SLAM)是机器人领域基础的状态估计问题,要求机器人在构建环境地图的同时准确定位自身。本文从最优随机控制角度研究主动SLAM,将其重构为部分信息下的决策问题。在回顾常见模型基础上,提出一种通用的随机控制框架,严格处理运动、感知与地图表示。引入新的探索阶段代价函数,基于状态几何结构评估信息获取行为。该框架被形式化为非标准的部分可观测马尔可夫决策过程(POMDP),并在此基础上推导出严格合理的近似解。为支持分析,研究了在广泛机器人应用场景下成立的一般性正则性条件。针对特定情形,通过标准学习算法进行大规模数值实验,验证了所学策略接近最优。
原文摘要 · Abstract (English)
Simultaneous localization and mapping (SLAM) is a foundational state estimation problem in robotics in which a robot accurately constructs a map of its environment while also localizing itself within this construction. We study the active SLAM problem through the lens of optimal stochastic control, thereby recasting it as a decision-making problem under partial information. After reviewing several commonly studied models, we present a general stochastic control formulation of active SLAM together with a rigorous treatment of motion, sensing, and map representation. We introduce a new exploration stage cost that encodes the geometry of the state when evaluating information-gathering actions. This formulation, constructed as a nonstandard partially observable Markov decision process (POMDP), is then analyzed to derive rigorously justified approximate solutions that are near-optimal. To enable this analysis, the associated regularity conditions are studied under general assumptions that apply to a wide range of robotics applications. For a particular case, we conduct an extensive numerical study in which standard learning algorithms are used to learn near-optimal policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。