arXiv:2512.00915cs.LGcs.RO2025-12被引 3

让强化学习在对称性被破坏时仍能高效泛化

Partially Equivariant Reinforcement Learning in Symmetry-Breaking Environments

  • 根据对称性是否成立,选择性使用不变或标准贝尔曼更新
  • 在网格、运动和操作任务中显著提升样本效率
  • 适合对称性不完全成立的真实环境,提升鲁棒性

群对称性为强化学习提供强大归纳偏置,可通过群不变马尔可夫决策过程(MDP)实现跨对称状态与动作的高效泛化。然而真实环境几乎从不满足完全群不变的MDP;动态、执行限制和奖励设计常导致对称性局部破坏。在此类情况下,若仍采用群不变贝尔曼更新,局部对称性破坏会引发误差传播,造成全局价值估计偏差。为此,本文提出部分群不变MDP(PI-MDP),根据对称性是否成立,选择性应用群不变或标准贝尔曼更新。该框架有效抑制局部对称性破坏带来的误差传播,同时保留等变性优势,提升样本效率与泛化能力。基于此,我们提出了适用于离散控制的PE-DQN和连续控制的PE-SAC算法,融合等变性与对称性破坏鲁棒性。在网格世界、运动与操作基准测试中,两类算法均显著优于基线方法,验证了选择性利用对称性的必要性。

原文摘要 · Abstract (English)

Group symmetries provide a powerful inductive bias for reinforcement learning (RL), enabling efficient generalization across symmetric states and actions via group-invariant Markov Decision Processes (MDPs). However, real-world environments almost never realize fully group-invariant MDPs; dynamics, actuation limits, and reward design usually break symmetries, often only locally. Under group-invariant Bellman backups for such cases, local symmetry-breaking introduces errors that propagate across the entire state-action space, resulting in global value estimation errors. To address this, we introduce Partially group-Invariant MDP (PI-MDP), which selectively applies group-invariant or standard Bellman backups depending on where symmetry holds. This framework mitigates error propagation from locally broken symmetries while maintaining the benefits of equivariance, thereby enhancing sample efficiency and generalizability. Building on this framework, we present practical RL algorithms -- Partially Equivariant (PE)-DQN for discrete control and PE-SAC for continuous control -- that combine the benefits of equivariance with robustness to symmetry-breaking. Experiments across Grid-World, locomotion, and manipulation benchmarks demonstrate that PE-DQN and PE-SAC significantly outperform baselines, highlighting the importance of selective symmetry exploitation for robust and sample-efficient RL. Project page: https://pranaboy72.github.io/perl_page/

强化学习对称性等变性样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。