通过结构化运动信息提升机器人操控关节物体的样本效率
KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation

- 引入KAI接口,用几何与运动先验增强策略学习
- 仅用一半演示数据即达82.9%成功率,低数据下表现突出
- 可融合人类操作视频,适应复杂现实场景
关节物体操控需要理解其运动结构,仅靠机器人示范难以高效学习。本文提出运动感知的关节物体接口(KAI),一种捕捉关节物体运动结构的结构化中间表示。通过在策略学习中嵌入可解释的几何与运动先验,KAI提供了与关节运动本质一致的强归纳偏置,显著提升样本效率。在六项仿真任务中,本方法平均成功率达82.9%,在仅使用一半示范数据的情况下,性能达到或超越基线。该方法对未见背景和视觉干扰也表现出鲁棒泛化能力,可从单一清洁训练环境迁移到杂乱真实场景。KAI的无动作依赖设计支持与人类交互视频联合训练,进一步提升现实鲁棒性:在多种视觉干扰下,联合训练后平均成功率超70%。
原文摘要 · Abstract (English)
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。