arXiv:2505.16732cs.LGcs.AI2025-05NeurIPS被引 3

用粒子滤波优化连续部分可观测环境下的决策,兼顾探索与利用。

Sequential Monte Carlo for Policy Optimization in Continuous POMDPs

  • 将策略优化转化为非马尔可夫的费曼-卡茨模型推断
  • 在标准基准上显著优于现有方法,尤其在不确定性下表现更优
  • 适合需要长期规划与信息收集的强化学习任务

在部分可观测环境下,智能体需权衡降低不确定性(探索)与追求即时目标(利用)。本文提出一种针对连续部分可观测马尔可夫决策过程(POMDP)的新策略优化框架,将策略学习建模为非马尔可夫费曼-卡茨模型中的概率推断,能自然捕捉信息获取的价值,无需次优近似或人工启发式。为此,我们设计了一种嵌套的序贯蒙特卡洛(SMC)算法,在最优轨迹分布样本上高效估计依赖历史的策略梯度。在多个标准连续POMDP基准测试中,该方法展现出优越性,尤其在处理不确定性时优于现有方法。

原文摘要 · Abstract (English)

Optimal decision-making under partial observability requires agents to balance reducing uncertainty (exploration) against pursuing immediate objectives (exploitation). In this paper, we introduce a novel policy optimization framework for continuous partially observable Markov decision processes (POMDPs) that explicitly addresses this challenge. Our method casts policy learning as probabilistic inference in a non-Markovian Feynman--Kac model that inherently captures the value of information gathering by anticipating future observations, without requiring suboptimal approximations or handcrafted heuristics. To optimize policies under this model, we develop a nested sequential Monte Carlo (SMC) algorithm that efficiently estimates a history-dependent policy gradient under samples from the optimal trajectory distribution induced by the POMDP. We demonstrate the effectiveness of our algorithm across standard continuous POMDP benchmarks, where existing methods struggle to act under uncertainty.

强化学习部分可观测蒙特卡洛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。