利用任务对称性实现跨任务高效泛化,突破传统方法局限。
Hereditary Geometric Meta-RL: Nonlocal Generalization via Task Symmetries
- 通过李群变换复用训练策略,将元强化学习转为对称性发现
- 在二维导航任务中实现全任务空间泛化,而基线仅限于局部
- 提升样本效率与数值稳定性,适合复杂结构化任务场景
元强化学习通常依赖任务编码的平滑性实现局部泛化,但需密集覆盖任务空间,难以挖掘深层结构。本文提出几何视角,使任务空间继承底层系统的固有对称性,形成“遗传几何”。具体地,智能体通过李群作用变换状态与动作,复用训练时策略,将元强化学习转化为对称性发现,从而实现更广范围的泛化。当任务空间源自系统对称性时,其嵌入到可线性化、连通且紧致的对称子群中,支持高效测试时学习与推理。为此,我们提出微分对称性发现方法,消解函数不变性约束,相比传统函数方法显著提升数值稳定性和样本效率。实验显示,在二维导航任务中,本方法高效恢复真实对称性并实现全任务空间泛化,而基线仅能近似训练任务区域。
原文摘要 · Abstract (English)
Meta-Reinforcement Learning (Meta-RL) commonly generalizes via smoothness in the task encoding. While this enables local generalization around each training task, it requires dense coverage of the task space and leaves richer task space structure untapped. In response, we develop a geometric perspective that endows the task space with a "hereditary geometry" induced by the inherent symmetries of the underlying system. Concretely, the agent reuses a policy learned at the train time by transforming states and actions through actions of a Lie group. This converts Meta-RL into symmetry discovery rather than smooth extrapolation, enabling the agent to generalize to wider regions of the task space. We show that when the task space is inherited from the symmetries of the underlying system, the task space embeds into a subgroup of those symmetries whose actions are linearizable, connected, and compact--properties that enable efficient learning and inference at the test time. To learn these structures, we develop a differential symmetry discovery method. This collapses functional invariance constraints and thereby improves numerical stability and sample efficiency over functional approaches. Empirically, on a two-dimensional navigation task, our method efficiently recovers the ground-truth symmetry and generalizes across the entire task space, while a common baseline generalizes only near training tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。