arXiv:2608.10936math.OCcs.LG2026-08中稿 · IEEE CDC 2026

发现重启部分可观马尔可夫决策中策略的阈值结构,可优化控制决策。

Threshold Structure of Optimal Policies in Restart POMDPs

  • 用最后观测状态和重启后时间构造充分统计量,将问题转化为可观测马尔可夫决策过程。
  • 在折扣与总成本准则下,最优策略对重启时间呈阈值结构,且阈值不随状态增加而上升。
  • 适用于具有单调性结构的系统,适合研究动态重启策略的智能体设计者。

我们研究定义在一般Borel状态空间上的重启部分可观马尔可夫决策过程(Restart POMDP),其中控制器可选择让隐状态持续演化或重启系统并观察新状态。通过利用包含最后一次观测状态和自重启以来经过时间的充分统计量表示,我们将问题简化为一个完全可观测的马尔可夫决策过程(MDP)。在自然的一步代价劣化条件下,我们证明了在折扣成本和总无折扣成本标准下,最优策略在经过时间上具有阈值结构。当状态空间具有偏序关系且转移核是随机单调时,进一步证明最优阈值关于状态是非增的。对于平均成本准则,在额外假设几何遍历性和瞬态收益主导的前提下,通过消失折扣法建立了类似的阈值结果,并证明了最优阈值和相对值函数的统一有界性。

原文摘要 · Abstract (English)

We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve unobserved or restarts the system and observes the new state. Exploiting a sufficient-statistic representation consisting of the last observed state and the elapsed time since restart, we reduce the problem to a fully observed MDP. Under a natural one-step cost deterioration condition, we prove that optimal policies have a threshold structure in the elapsed time for both the discounted and total undiscounted cost criteria. When the state space is partially ordered and the kernel is stochastically monotone, we further show that the optimal threshold is nonincreasing in the state. For the average cost criterion, under additional assumptions of geometric ergodicity and domination of the transient gain, we establish analogous threshold results via the vanishing discount approach, after showing the uniform boundedness of the optimal thresholds and relative value functions.

强化学习决策优化马尔可夫决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。