用视觉预测相对位姿,让机器人零样本插入复杂环境中的新物体
EasyInsert: A Data-Efficient and Generalizable Insertion Policy
- 将插入任务转为位姿差回归,仅需1小时人工操作数据
- 在15个未见物体上实现超90%成功率,含电缆等挑战物
- 支持自动数据收集与快速微调,适合真实场景部署
机器人插入任务在杂乱环境中极具挑战性,现有方法泛化能力差,常依赖结构化环境或CAD模型。为此,我们提出EasyInsert,借鉴人类直觉,将插入建模为位姿差回归问题,实现高效、可扩展的数据采集,仅需少量人工标注即可训练端到端视觉策略。执行时,视觉策略预测插头与插座的相对位姿,驱动多阶段粗到精的插入流程。在真实实验中,利用仅1小时人类遥操作数据启动大规模自动化数据采集,对15个未见的新物体中13个实现超过90%的零样本成功率,包括Type-C、HDMI和以太网线等复杂对象。此外,仅需一次手动重置,通过自动化数据收集与微调,所有15个物体均达超90%成功率。
原文摘要 · Abstract (English)
Robotic insertion is a highly challenging task that requires exceptional precision in cluttered environments. Existing methods often have poor generalization capabilities. They typically function in restricted and structured environments, and frequently fail when the plug and socket are far apart, when the scene is densely cluttered, or when handling novel objects. They also rely on strong assumptions such as access to CAD models or a digital twin in simulation. To address these limitations, we propose EasyInsert. Inspired by human intuition, it formulates insertion as a delta-pose regression problem, which unlocks an efficient, highly scalable data collection pipeline with minimal human labor to train an end-to-end visual policy. During execution, the visual policy predicts the relative pose between plug and socket to drive a multi-phase, coarse-to-fine insertion process. EasyInsert demonstrates strong zero-shot generalization capability for unseen objects in cluttered environments, robustly handling cases with significant initial pose deviations. In real-world experiments, by leveraging just 1 hour of human teleoperation data to bootstrap a large-scale automated data collection process, EasyInsert achieves an over 90% success rate in zero-shot insertion for 13 out of 15 unseen novel objects, including challenging objects like Type-C cables, HDMI cables, and Ethernet cables. Furthermore, requiring only a single manual reset, EasyInsert allows for fast adaptation to novel test objects through automated data collection and fine-tuning, achieving an over 90% success rate across all 15 objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。