arXiv:2510.02149cs.LGmath.OC2025-10被引 1

动作触发观测的强化学习框架,解决部分可观测问题

Reinforcement Learning with Action-Triggered Observations

  • 动作决定是否获得完整状态观测,建模稀疏观测场景
  • 在线性MDP假设下,动作序列值函数可线性表示,支持高效算法设计
  • 算法达到与全观测相同的最优率,适合稀疏反馈环境

我们提出动作触发稀疏可追溯马尔可夫决策过程(ATST-MDP),一种针对部分可观测性的强化学习框架,其中每步以由所选动作决定概率的概率随机获得完整状态观测。我们推导了适用于该设定的贝尔曼方程,并证明了最优策略的存在性。利用稀疏观测能揭示完整状态的特性,我们给出一个等价表述:代理在连续观测之间承诺行动序列。在线性MDP假设下,我们证明该行动序列上的价值函数可在有限维特征映射中线性表示,从而支持标准回归方法。作为应用,我们提出了ATST-LSVI-UCB,一种乐观算法,在几何分布时长的回合制学习中实现$\widetilde{O}(\sqrt{Kd^3(1-γ)^{-3}})$的后悔率,其中$K$为回合数,$d$为特征维度,$γ$为折扣因子(回合延续概率),与全观测下线性MDP的已知最优率一致。

原文摘要 · Abstract (English)

We introduce Action-Triggered Sporadically Traceable Markov Decision Processes (ATST-MDPs), a reinforcement learning framework for partial observability in which full state observations occur stochastically at each step, with probability determined by the chosen action. We derive Bellman equations tailored to this setting and establish the existence of an optimal policy. Exploiting the fact that sporadic observations reveal the full state, we provide an equivalent formulation in which agents commit to action-sequences between consecutive observations. Under the linear MDP assumption, we show that the value function over such action-sequences admits a linear representation in a finite-dimensional feature map, enabling standard regression-based methods. As an application, we derive ATST-LSVI-UCB, an optimistic algorithm achieving regret $\widetilde{O}(\sqrt{Kd^3(1-γ)^{-3}})$ for episodic learning with geometrically distributed horizons, where $K$ is the number of episodes, $d$ the feature dimension, and $γ$ the discount factor (episode continuation probability), matching the known rate for linear MDPs with full observability.

强化学习部分可观测线性MDP稀疏观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。