提出新方法分离马尔可夫链中的持久与瞬时行为,提升策略评估精度。
Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients
- 通过最小边界商构造精确分解,消除非衰减模式干扰
- 重构有限时域回报,恢复状态平均奖励,稳定估计器收敛
- 适合研究复杂马尔可夫链的理论分析与强化学习建模
我们研究可能不可约且周期性的有限马尔可夫链的固定策略评估问题。经典方法中基于增益与偏移的分解并非总是具有诊断性:增益仅记录不变的Cesàro平均,而持久的相位依赖行为与真正的瞬时效应一同被吸收进偏移项。我们识别出转移矩阵 $P$ 的真实边界不变子空间 $/mathcal{K}(P)$ 是造成模糊性的根源。对 $/mathcal{K}(P)$ 取商是去除所有非衰减模式并使剩余动力学严格稳定的最小精确商。在选择以 $/mathcal{K}(P)$ 为核的规范投影 $Π$ 后,奖励可唯一分解为 $r = g_Π^ ext{⋆} + (I-P)v_Π^ ext{⋆}$,其中 $g_Π^ ext{⋆}$ 为持久态特征,$v_Π^ ext{⋆}$ 为规范固定的瞬时分量。与经典归一化增益和偏移相比,新配对重新分配了相同信息,使所有持久模式均体现在 $g_Π^ ext{⋆}$,而 $v_Π^ ext{⋆}$ 纯为瞬时。该分解能重构有限时域回报,恢复状态平均奖励,支持瞬时成本解释,并在生成模型下给出稳定估计器。
原文摘要 · Abstract (English)
We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnostic: the gain records only invariant Cesàro averages, while persistent phase-dependent behavior is absorbed into the bias together with genuinely transient effects. We identify the real peripheral invariant subspace $\mathcal{K}(P)$ of the transition matrix $P$ as the source of this ambiguity. Quotienting by $\mathcal{K}(P)$ is the minimal exact quotient that removes all non-decaying modes and makes the remaining dynamics strictly stable. After choosing a gauge projection $Π$ with kernel $\mathcal{K}(P)$, the reward admits a unique decomposition $r = g_Π^\star + (I-P)v_Π^\star$, where $g_Π^\star$ is a persistent regime profile and $v_Π^\star$ is a gauge-fixed transient component. An exact comparison with classical normalized gain and bias shows that the new pair reallocates the same information so that all persistent modes are represented in $g_Π^\star$ and $v_Π^\star$ is transient. This decomposition reconstructs finite-horizon returns, recovers statewise average reward, admits a transient-cost interpretation, and yields a stable estimator under a generative model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。