让扩散策略利用对称性,提升强化学习的样本效率与稳定性。
Symmetry-Aware Steering of Equivariant Diffusion Policies: Benefits and Limits
- 基于对称性设计专用强化学习框架,避免破坏模型几何特性。
- 在对称性受限时仍显著提升采样效率与策略性能。
- 适合对样本敏感、需泛化能力的机器人控制任务。
等变扩散策略(EDPs)结合了扩散模型的生成表达能力与几何对称性带来的强泛化和高样本效率。尽管通过强化学习(RL)微调可超越演示数据,但直接使用非等变的RL方法常因忽略对称性而出现样本效率低、训练不稳定的问题。本文理论证明:EDP的扩散过程具有等变性,从而诱导出一个群不变的潜在噪声马尔可夫决策过程(MDP),非常适合等变扩散策略的微调。基于此理论,我们提出一种原则性的对称性感知微调框架,并在多种对称程度不同的任务上对比标准、等变及近似等变的RL策略。虽然在对称性被破坏时严格等变存在实际边界,但利用对称性进行微调仍能显著提升样本效率,防止价值发散,并在极有限演示数据下实现强策略改进。
原文摘要 · Abstract (English)
Equivariant diffusion policies (EDPs) combine the generative expressivity of diffusion models with the strong generalization and sample efficiency afforded by geometric symmetries. While steering these policies with reinforcement learning (RL) offers a promising mechanism for fine-tuning beyond demonstration data, directly applying standard (non-equivariant) RL can be sample-inefficient and unstable, as it ignores the symmetries that EDPs are designed to exploit. In this paper, we theoretically establish that the diffusion process of an EDP is equivariant, which in turn induces a group-invariant latent-noise MDP that is well-suited for equivariant diffusion steering. Building on this theory, we introduce a principled symmetry-aware steering framework and compare standard, equivariant, and approximately equivariant RL strategies through comprehensive experiments across tasks with varying degrees of symmetry. While we identify the practical boundaries of strict equivariance under symmetry breaking, we show that exploiting symmetry during the steering process yields substantial benefits-enhancing sample efficiency, preventing value divergence, and achieving strong policy improvements even when EDPs are trained from extremely limited demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。