arXiv:2606.02280cs.RO2026-06

让智能体自己学动态变化,无需预设参数即可零样本适应新环境。

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

论文配图:Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation
图 1 · 摘自论文原文
  • 从结果出发学习动态几何结构,不依赖预设物理参数。
  • 在MuJoCo上大幅优于传统方法,支持未建模和时变动态变化。
  • 适合需要鲁棒适应性的机器人控制场景,提升泛化能力。

现实世界中的动态变化对机器人强化学习构成重大挑战,因紧密耦合于标准环境的策略在物理条件改变时常会失效。现有方法多依赖显式编码已知物理参数至潜在上下文,这种以参数为中心的范式需预先指定变化轴,面对未建模或复合动态变化时变得脆弱。本文从结果导向视角重新审视动态适应:不告诉策略动态是什么,而是让其学会动态如何影响交互结果。理论上,目标域遗憾与轨迹动态编码器的Lipschitz常数呈单调关系;实践中,该常数可通过对比学习上界,获得平滑且任务相关的潜在拓扑,无需特权动态信息。在MuJoCo基准测试中,本方法在严重动态漂移下(包括未建模与时变参数)持续优于参数中心基线,同时提升分布内稳定性和潜在空间可解释性。结果验证了控制潜在几何结构是实现鲁棒适应的合理机制。

原文摘要 · Abstract (English)

Real-world dynamics shifts pose a critical challenge for reinforcement learning in robotics, as policies tightly coupled to nominal environments often fail catastrophically when physical conditions change. Most existing methods rely on encoding explicitly identified physical parameters into a latent context, a parameter-centric paradigm that depends on pre-specified axes of variation and becomes brittle under unmodeled or compound dynamics changes. We revisit dynamics adaptation from an outcome-centric perspective: rather than telling policies what the dynamics are, we enable them to learn how dynamics affect interaction outcomes. Theoretically, this is grounded in a monotonic relationship between target-domain regret and the Lipschitz constant of a trajectory dynamics encoder. Practically, this constant can be upper-bounded through contrastive learning, yielding a smooth, task-relevant latent topology without privileged dynamics information. On MuJoCo benchmarks, our method consistently outperforms parameter-centric baselines under severe dynamics shifts, including unmodeled and time-varying parameters, while also improving in-distribution stability and latent interpretability. Overall, these results validate that controlling latent geometry is a principled mechanism for robust adaptation.

强化学习动态适应潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。