arXiv:2604.08620cs.LGcs.AI2026-04

从分布强化学习中挖掘动态规划结构,提升采样效率。

StructRL: Recovering Dynamic Programming Structure from Learning Dynamics in Distributional Reinforcement Learning

论文配图:StructRL: Recovering Dynamic Programming Structure from Learning Dynamics in Distributional Reinforcement Learning
图 1 · 摘自论文原文
  • 通过分析回报分布演化,提取状态学习时机信号。
  • 该信号可诱导出类似动态规划的传播顺序。
  • 适合关注高效采样与结构化学习的算法研究者。

强化学习通常被视为统一的数据驱动优化过程,更新依赖奖励和时序差分误差,未显式利用全局结构。相反,动态规划方法依赖结构化信息传播,实现高效稳定的学习。本文提供证据表明,分布强化学习的学习动态中可恢复此类结构。通过分析回报分布的时序演化,我们识别出捕捉状态空间中学习发生时机与位置的信号。特别地,引入时间学习指标 t*(s),反映状态在训练过程中最强学习更新的时刻。实证显示,该信号对状态形成与动态规划传播一致的排序。基于此,我们提出 StructRL 框架,利用这些信号引导采样以契合浮现的传播结构。初步结果表明,分布学习动态为恢复并利用类似动态规划的结构提供了机制,无需显式模型。这为强化学习提供了新视角:学习可被理解为结构化传播过程,而非纯粹均匀优化。

原文摘要 · Abstract (English)

Reinforcement learning is typically treated as a uniform, data-driven optimization process, where updates are guided by rewards and temporal-difference errors without explicitly exploiting global structure. In contrast, dynamic programming methods rely on structured information propagation, enabling efficient and stable learning. In this paper, we provide evidence that such structure can be recovered from the learning dynamics of distributional reinforcement learning. By analyzing the temporal evolution of return distributions, we identify signals that capture when and where learning occurs in the state space. In particular, we introduce a temporal learning indicator t*(s) that reflects when a state undergoes its strongest learning update during training. Empirically, this signal induces an ordering over states that is consistent with a dynamic programming-style propagation of information. Building on this observation, we propose StructRL, a framework that exploits these signals to guide sampling in alignment with the emerging propagation structure. Our preliminary results suggest that distributional learning dynamics provide a mechanism to recover and exploit dynamic programming-like structure without requiring an explicit model. This offers a new perspective on reinforcement learning, where learning can be interpreted as a structured propagation process rather than a purely uniform optimization procedure.

强化学习动态规划分布学习采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。