用参考动作引导高维机器人技能发现,实现模仿与创新双重突破
Reference Grounded Skill Discovery
- 以参考数据在语义空间中锚定动作方向,指导高效探索
- 在359维观测、69维动作下成功模仿走跑踢等动作并发现变体
- 适合需要风格化控制的机器人运动生成任务
将无监督技能发现算法扩展到高自由度代理仍具挑战性。随着维度增加,探索空间呈指数级增长,而有意义技能的流形却相对有限,因此语义意义成为有效引导高维空间探索的关键。本文提出参考锚定技能发现(RGSD),通过参考数据在语义有意义的隐空间中锚定技能发现。RGSD首先进行对比预训练,将动作嵌入单位超球面,使每个参考轨迹聚类为独立方向。该锚定机制使技能发现能同时实现对参考行为的模仿和语义相关多样化行为的发现。在具有359维观测和69维动作的模拟SMPL人形机器人上,RGSD成功模仿行走、跑步、出拳和侧步等技能,并发现这些动作的变体。在下游运动任务中,RGSD利用发现的技能忠实响应用户指定的风格指令,性能优于依赖模仿学习的基线方法,后者常无法保持指定风格。
原文摘要 · Abstract (English)
Scaling unsupervised skill discovery algorithms to high-DoF agents remains challenging. As dimensionality increases, the exploration space grows exponentially, while the manifold of meaningful skills remains limited. Therefore, semantic meaningfulness becomes essential to effectively guide exploration in high-dimensional spaces. In this work, we present Reference-Grounded Skill Discovery (RGSD), a novel algorithm that grounds skill discovery in a semantically meaningful latent space using reference data. RGSD first performs contrastive pretraining to embed motions on a unit hypersphere, clustering each reference trajectory into a distinct direction. This grounding enables skill discovery to simultaneously involve both imitation of reference behaviors and the discovery of semantically related diverse behaviors. On a simulated SMPL humanoid with $359$-D observations and $69$-D actions, RGSD successfully imitates skills such as walking, running, punching, and sidestepping, while also discover variations of these behaviors. In downstream locomotion tasks, RGSD leverages the discovered skills to faithfully satisfy user-specified style commands and outperforms imitation-learning baselines, which often fail to maintain the commanded style.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。