arXiv:2607.27132cs.LG2026-07

提出最小记忆状态构造方法,解决部分可观测决策中的最优记忆压缩问题。

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

论文配图:Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes
图 1 · 摘自论文原文
  • 基于稳定商构造最小马尔可夫状态,通过观测与隐藏模式的置换结构实现
  • 实验验证可将原始状态压缩至商状态,三记忆状态即达最优配对精度
  • 适合研究部分可观测强化学习中记忆效率的学者与工程师

在部分可观测决策过程中,智能体需维护历史信息的递归更新统计量以恢复马尔可夫性,但最小此类统计量通常未知。本文针对一类结构化部分可观测马尔可夫决策过程(holonomy-cover POMDP),其中可见动态为马尔可夫,且每个可见转移均对隐藏模式施加固定置换,给出了最小马尔可夫充分统计量的刻画。具体地,构建了稳定商——一种保持一步奖励与商后继关系的最粗粒度观测抽象,并证明当前观测与稳定类的组合构成精确有限马尔可夫状态。当初始类正确时,精确类追踪所需记忆符号数恰好最小:在可达性与最大化观测下任意两决策分离条件下,无有限记忆控制器能使用更少符号。在可重置诊断下,最近原型类推断误差呈指数衰减;通过校准-重启归约,可将有限MDP保证转移至恢复状态。结果支持‘同伦记忆强化学习’,以当前稳定类表示记忆,通过有序边传输更新,诊断可用时识别局部类坐标,并在同步后应用标准有限MDP RL主干。实验显示,可从原始状态精确压缩至商状态,三决策时间记忆状态下实现完美配对顺序准确率,匹配商预言机,优于非预言机基线。

原文摘要 · Abstract (English)

An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized visible transition applies a fixed permutation to a hidden mode. In particular, we construct the stable quotient, the coarsest observation-wise abstraction preserving one-step rewards and quotient successors, and prove that the pair of the current observation and stable class forms an exact finite Markov state. When the current class is correctly initialized, exact class tracking requires exactly the minimal memory symbols, in the sense that under reachability and pairwise decision separation at a maximizing observation, no arbitrary finite-memory controller can use fewer. Under resettable diagnostics, nearest-prototype class inference has exponentially decaying error, and a calibrate-then-restart reduction transfers finite-MDP guarantees to the recovered state. The results enable \emph{Holonomy Memory Reinforcement Learning}. It represents memory by the current stable class, updates it through ordered edge transports, identifies local class coordinates when diagnostics are available, and applies a standard finite-MDP RL backbone after synchronization. Experiments recover an exact compression from raw states to quotient states and achieve perfect paired-order accuracy with three decision-time memory states, matching the quotient oracle and outperforming the non-oracle baselines.

强化学习部分可观测记忆压缩马尔可夫性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。