arXiv:2602.02978cs.AI2026-02被引 9

用序理论重构价值函数,提升强化学习的稳定性和样本效率。

Structuring Value Representations via Geometric Coherence in Markov Decision Processes

  • 将价值函数学习转化为有序集的逐步精化过程
  • 在多个任务中实现更优的样本效率与稳定表现
  • 适合关注强化学习收敛性与结构化表示的研究者

几何特性可用来稳定并加速强化学习。现有方法包括编码对称结构、几何感知数据增强和施加结构约束。本文从序理论视角重新审视强化学习,将价值函数估计重构为学习期望的偏序集(poset)。我们提出GCR-RL(几何一致性正则化强化学习),通过逐步精化前序偏序集并从时序差分信号中学习新的序关系,确保支撑价值函数的偏序集序列具有几何一致性。基于Q-learning和演员-评论家框架开发了两种新算法,以高效实现超偏序集的精化。理论分析了其性质与收敛速率。在多个任务上的实验表明,GCR-RL相比强基线显著提升样本效率和性能稳定性。

原文摘要 · Abstract (English)

Geometric properties can be leveraged to stabilize and speed reinforcement learning. Existing examples include encoding symmetry structure, geometry-aware data augmentation, and enforcing structural restrictions. In this paper, we take a novel view of RL through the lens of order theory and recast value function estimates into learning a desired poset (partially ordered set). We propose \emph{GCR-RL} (Geometric Coherence Regularized Reinforcement Learning) that computes a sequence of super-poset refinements -- by refining posets in previous steps and learning additional order relationships from temporal difference signals -- thus ensuring geometric coherence across the sequence of posets underpinning the learned value functions. Two novel algorithms by Q-learning and by actor--critic are developed to efficiently realize these super-poset refinements. Their theoretical properties and convergence rates are analyzed. We empirically evaluate GCR-RL in a range of tasks and demonstrate significant improvements in sample efficiency and stable performance over strong baselines.

强化学习序理论价值函数几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。