arXiv:2502.06523cs.AI2025-02被引 2

提出更紧的值函数上界,加速求解部分可观测马尔可夫决策过程。

Tighter Value-Function Approximations for POMDPs

  • 设计新上界方法,比常用快速有信息上界更紧
  • 在多个基准测试中显著提升求解速度
  • 适合需要高效求解复杂决策问题的研究者

求解部分可观测马尔可夫决策过程(POMDPs)通常需要对指数级多的状态信念值进行推理。现有先进求解器通过值界来引导这一过程。然而,可靠的上界计算往往代价高昂,且其紧度与计算成本之间存在权衡。本文提出了新的、可证明比常用快速有信息上界更紧的上界。实验表明,尽管增加了额外计算开销,新上界在广泛基准测试中仍能加速现有最优求解器的性能。

原文摘要 · Abstract (English)

Solving partially observable Markov decision processes (POMDPs) typically requires reasoning about the values of exponentially many state beliefs. Towards practical performance, state-of-the-art solvers use value bounds to guide this reasoning. However, sound upper value bounds are often computationally expensive to compute, and there is a tradeoff between the tightness of such bounds and their computational cost. This paper introduces new and provably tighter upper value bounds than the commonly used fast informed bound. Our empirical evaluation shows that, despite their additional computational overhead, the new upper bounds accelerate state-of-the-art POMDP solvers on a wide range of benchmarks.

强化学习决策优化值函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。