arXiv:2607.11624cs.ROcs.LG2026-07中稿 · publication at the…

用对称柯尔莫哥洛夫模型提升四足机器人强化学习效率

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning

论文配图:SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning
图 1 · 摘自论文原文
  • 结合系统对称性与自编码器学习动力学模型
  • 收敛速度加快,奖励提升30%以上,跨环境可迁移
  • 适合复杂非线性机器人的高效强化学习研究

强化学习算法通常样本效率低。近期研究通过在学习中融入物理先验来缓解此问题,但多数方法仅在低维基准系统上验证,未涉及高维、复杂非线性动力学的机器人。本文提出SKooP(对称柯尔莫哥洛夫预测),将形态对称性与自编码器学习的柯尔莫哥洛夫模型结合,同时学习策略与系统动力学模型。利用柯尔莫哥洛夫预测作为批评者的优势观测,使智能体基于更平滑、信息更丰富的特征进行学习。同时在演员、批评者、编码器和解码器网络中引入群对称性,生成高度等变策略。通过深入分析学习到的柯尔莫哥洛夫模型和对称策略,验证了其对性能的影响。结果表明,该方法在多个挑战性的双足行走任务中显著缩短收敛时间,并提高最终奖励,且策略可在不同仿真环境中迁移。项目页:https://evelyd.github.io/SymmetricKoopmanPredictions

原文摘要 · Abstract (English)

Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency. In robotics, a recent line of work has emerged addressing this problem by encoding physics priors in the learning process. However, most of these approaches are validated on well-defined, low-dimensional benchmark systems rather than high-dimensional robots with complex nonlinear dynamics. In this paper, we introduce \textit{SKooP (Symmetric Koopman Predictions)}, an approach combining the advantages of morphological symmetries with those of a Koopman model learned via autoencoder to enhance policy learning. SKooP learns a Koopman model of the system dynamics alongside the policy. The resulting Koopman predictions are used as privileged observations for the critic, allowing the agent to learn based on smoother, more informative features. We also incorporate group symmetries into the actor, critic, encoder and decoder networks to produce a highly equivariant policy. The SKooP approach is validated via in-depth analysis of the learned Koopman models and symmetric policies to showcase how each of these influences the agent's performance. We also show that the learned policies are transferable to different simulation environments. Our results show that SKooP consistently reduces convergence time and increases the learned reward for multiple challenging bipedal locomotion tasks on a quadruped robot. Project page: https://evelyd.github.io/SymmetricKoopmanPredictions

强化学习四足机器人柯尔莫哥洛夫模型对称性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。