通过分因子学习,让机器人自动发现可解释、安全且可部署的多样化行为。
Divide, Discover, Deploy: Factorized Skill Learning with Symmetry and Style Priors
- 按状态空间分因子,为每类因子匹配不同技能发现算法。
- 引入对称性先验与风格因子,提升行为结构化与安全性。
- 技能可零样本迁移到真实机器人,性能媲美人工奖励训练策略。
无监督技能发现(USD)使智能体能在无任务特定奖励的情况下自主学习多样行为。尽管近期方法展现出潜力,其在实际机器人中的应用仍有限。本文提出一种模块化USD框架,解决学习技能的安全性、可解释性与可部署性挑战。方法基于用户定义的状态空间因子化,学习解耦的技能表征,并根据不同因子分配相应的内在奖励函数与发现算法。为促进形态感知的结构化技能,我们针对各因子设计了对称性归纳偏置。同时引入风格因子与正则化惩罚,以增强行为的安全性与多样性。我们在四足机器人仿真中评估该框架,实现所学技能的零样本迁移至真实硬件。结果表明,因子化与对称性促使发现结构化且人类可理解的行为;风格因子与惩罚项显著提升安全性和多样性。此外,所学技能可用于下游任务,性能达到采用手工设计奖励训练的最优策略水平。
原文摘要 · Abstract (English)
Unsupervised Skill Discovery (USD) allows agents to autonomously learn diverse behaviors without task-specific rewards. While recent USD methods have shown promise, their application to real-world robotics remains underexplored. In this paper, we propose a modular USD framework to address the challenges in the safety, interpretability, and deployability of the learned skills. Our approach employs user-defined factorization of the state space to learn disentangled skill representations. It assigns different skill discovery algorithms to each factor based on the desired intrinsic reward function. To encourage structured morphology-aware skills, we introduce symmetry-based inductive biases tailored to individual factors. We also incorporate a style factor and regularization penalties to promote safe and robust behaviors. We evaluate our framework in simulation using a quadrupedal robot and demonstrate zero-shot transfer of the learned skills to real hardware. Our results show that factorization and symmetry lead to the discovery of structured human-interpretable behaviors, while the style factor and penalties enhance safety and diversity. Additionally, we show that the learned skills can be used for downstream tasks and perform on par with oracle policies trained with hand-crafted rewards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。