arXiv:2505.00596cs.ROcs.AI2025-05IJCAI被引 2

提出基于有限状态控制器的离线求解器,高效解决确定性部分可观测决策问题。

A Finite-State Controller Based Offline Solver for Deterministic POMDPs

  • 将蒙特卡洛价值迭代扩展为有限状态控制器形式,构建策略
  • 在大规模问题上成功率高,优于现有DetPOMDP基线方法
  • 已在真实移动机器人森林测绘场景验证,适合实际应用

确定性部分可观测马尔可夫决策过程(DetPOMDPs)常出现在智能体对环境状态不确定但可确定性执行动作与观测的规划问题中。本文提出DetMCVI,是蒙特卡洛价值迭代(MCVI)算法在DetPOMDP上的适配版本,能够以有限状态控制器(FSCs)的形式构建策略。该方法在大规模问题上表现出高成功率,显著优于现有的DetPOMDP基线算法。我们在一个真实的移动机器人森林地图构建场景中验证了该算法的有效性。

原文摘要 · Abstract (English)

Deterministic partially observable Markov decision processes (DetPOMDPs) often arise in planning problems where the agent is uncertain about its environmental state but can act and observe deterministically. In this paper, we propose DetMCVI, an adaptation of the Monte Carlo Value Iteration (MCVI) algorithm for DetPOMDPs, which builds policies in the form of finite-state controllers (FSCs). DetMCVI solves large problems with a high success rate, outperforming existing baselines for DetPOMDPs. We also verify the performance of the algorithm in a real-world mobile robot forest mapping scenario.

强化学习规划有限状态控制器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。