无监督学习机器人操作模式的离散表示,提升抓取泛化能力
Discovering Robotic Interaction Modes with Discrete Representation Learning
- 通过自监督学习生成操作模式的离散编码,无需专家标注
- 在真实任务中实现92%的成功率,显著优于基线方法
- 适合需要自主理解交互模式的机器人研究者
人类操控可动物体(如开关抽屉)的行为可划分为多种交互模式。传统机器人学习方法缺乏对这些模式的离散表示,难以实现有效采样与语义对齐。本文提出ActAIM2,一种完全无监督的方法,可学习机器人操作中的离散交互模式表示,不依赖专家标签或仿真器特权信息。该方法结合模拟器滚动采集的新数据,包含一个交互模式选择器和一个低层动作预测器:选择器通过自监督生成潜在交互模式的离散表征,预测器则输出对应的动作轨迹。实验验证了其在操控可动物体上的成功率和从离散表示中采样有意义动作的鲁棒性。大量对比实验表明,与基线及消融实验相比,ActAIM2在可操作性和泛化性上均有显著提升。更多视频与结果详见:https://actaim2.github.io/
原文摘要 · Abstract (English)
Human actions manipulating articulated objects, such as opening and closing a drawer, can be categorized into multiple modalities we define as interaction modes. Traditional robot learning approaches lack discrete representations of these modes, which are crucial for empirical sampling and grounding. In this paper, we present ActAIM2, which learns a discrete representation of robot manipulation interaction modes in a purely unsupervised fashion, without the use of expert labels or simulator-based privileged information. Utilizing novel data collection methods involving simulator rollouts, ActAIM2 consists of an interaction mode selector and a low-level action predictor. The selector generates discrete representations of potential interaction modes with self-supervision, while the predictor outputs corresponding action trajectories. Our method is validated through its success rate in manipulating articulated objects and its robustness in sampling meaningful actions from the discrete representation. Extensive experiments demonstrate ActAIM2's effectiveness in enhancing manipulability and generalizability over baselines and ablation studies. For videos and additional results, see our website: https://actaim2.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。