arXiv:2511.03473cs.LG2025-11

利用环境对称性提升强化学习效率,显著减少样本需求。

Reinforcement Learning Using known Invariances

  • 基于核方法设计对称性感知的值迭代算法,编码奖励与动态的不变性。
  • 理论证明对称性使信息增益和覆盖数降低,样本效率提升明显。
  • 在冰湖与布局设计任务中表现优于标准核方法,适合有结构先验场景。

在许多现实世界的强化学习问题中,环境具有可利用的内在对称性以提高学习效率。本文构建了将已知群对称性融入核方法强化学习的理论与算法框架。提出一种对称性感知的乐观最小二乘值迭代(LSVI)变体,利用不变核编码奖励与转移动态中的对称性。分析建立了不变再生核希尔伯特空间(invariant RKHS)的最大信息增益与覆盖数的新界,明确量化了对称性带来的样本效率增益。在定制化的Frozen Lake环境与二维布局设计问题上的实验结果验证了理论优势,表明对称性感知的RL显著优于标准核方法。这些发现凸显了结构先验在设计更高效强化学习算法中的价值。

原文摘要 · Abstract (English)

In many real-world reinforcement learning (RL) problems, the environment exhibits inherent symmetries that can be exploited to improve learning efficiency. This paper develops a theoretical and algorithmic framework for incorporating known group symmetries into kernel-based RL. We propose a symmetry-aware variant of optimistic least-squares value iteration (LSVI), which leverages invariant kernels to encode invariance in both rewards and transition dynamics. Our analysis establishes new bounds on the maximum information gain and covering numbers for invariant RKHSs, explicitly quantifying the sample efficiency gains from symmetry. Empirical results on a customized Frozen Lake environment and a 2D placement design problem confirm the theoretical improvements, demonstrating that symmetry-aware RL achieves significantly better performance than their standard kernel counterparts. These findings highlight the value of structural priors in designing more sample-efficient reinforcement learning algorithms.

强化学习对称性核方法样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。