用有限记忆逼近信念状态,量化控制性能损失。
Finite Memory Belief Approximation for Optimal Control in Partially Observable Markov Decision Processes
- 以截断输入输出历史构建信念近似,建立信息损失与控制性能的关联
- 基于Wasserstein度量,给出闭环轨迹上的性能退化上界
- 在LQG系统中验证信念误差指数衰减,性能损失随之变化
我们研究部分可观测随机最优控制(SOC)问题中的有限记忆信念逼近。尽管信念状态在部分可观测马尔可夫决策过程(POMDPs)中是充分的,但其通常为无穷维,难以实际应用。本文将截断的输入-输出(IO)历史解释为诱导信念逼近,并发展了一种基于度量的理论,直接关联信息损失与控制性能。利用Wasserstein度量,我们推导出政策条件下的性能边界,量化了在典型闭环轨迹上由有限记忆引起的值函数退化。分析通过固定策略比较进行:在相同闭环执行下评估两个代价函数,隔离真实信念被有限记忆近似替换时在信念级代价中的影响。对于线性二次高斯(LQG)系统,我们提供了闭式信念错配评估,并实证验证了预测机制,表明信念错配随记忆长度近似指数衰减,且由此引发的性能错配也相应变化。这些结果共同提供了对有限记忆信念逼近在部分可观测设置中能力与局限性的度量感知刻画。
原文摘要 · Abstract (English)
We study finite memory belief approximation for partially observable (PO) stochastic optimal control (SOC) problems. While belief states are sufficient for SOC in partially observable Markov decision processes (POMDPs), they are generally infinite-dimensional and impractical. We interpret truncated input-output (IO) histories as inducing a belief approximation and develop a metric-based theory that directly relates information loss to control performance. Using the Wasserstein metric, we derive policy-conditional performance bounds that quantify value degradation induced by finite memory along typical closed-loop trajectories. Our analysis proceeds via a fixed-policy comparison: we evaluate two cost functionals under the same closed-loop execution and isolate the effect of replacing the true belief by its finite memory approximation inside the belief-level cost. For linear quadratic Gaussian (LQG) systems, we provide closed-form belief mismatch evaluation and empirically validate the predicted mechanism, demonstrating that belief mismatch decays approximately exponentially with memory length and that the induced performance mismatch scales accordingly. Together, these results provide a metric-aware characterization of what finite memory belief approximation can and cannot achieve in PO settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。