构建可动态更新的原子技能库,降低具身操作的数据需求
An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation
- 用视觉-语言-规划分解任务,抽象出原子技能
- 通过三轮更新机制动态扩展技能库,支持新任务快速适配
- 实测证明数据效率高,适合真实场景的机器人操作
具身操作是具身人工智能的基础能力。现有模型在特定环境下虽具泛化性,但面对新环境与任务时表现不佳,因现实场景复杂多样。传统端到端数据收集与训练方式数据消耗大。将端到端任务分解为原子技能可降低数据需求并提升成功率。然而,现有方法受限于预设技能集,无法动态更新。为此,我们提出一种三轮驱动的数据驱动方法构建原子技能库:首先利用视觉-语言-规划(VLP)将任务分解为子任务;其次通过抽象子任务形成原子技能定义;最后通过数据收集与视觉-语言-动作(VLA)微调构建技能库。随着三轮更新策略推进,技能库持续动态扩展,任务覆盖范围自然增长。该方法从端到端任务转向原子技能,显著降低数据成本,保持高性能,并实现对新任务的高效适应。大量真实场景实验验证了本方法的有效性与高效性。
原文摘要 · Abstract (English)
Embodied manipulation is a fundamental ability in the realm of embodied artificial intelligence. Although current embodied manipulation models show certain generalizations in specific settings, they struggle in new environments and tasks due to the complexity and diversity of real-world scenarios. The traditional end-to-end data collection and training manner leads to significant data demands. Decomposing end-to-end tasks into atomic skills helps reduce data requirements and improves the task success rate. However, existing methods are limited by predefined skill sets that cannot be dynamically updated. To address the issue, we introduce a three-wheeled data-driven method to build an atomic skill library. We divide tasks into subtasks using the Vision-Language-Planning (VLP). Then, atomic skill definitions are formed by abstracting the subtasks. Finally, an atomic skill library is constructed via data collection and Vision-Language-Action (VLA) fine-tuning. As the atomic skill library expands dynamically with the three-wheel update strategy, the range of tasks it can cover grows naturally. In this way, our method shifts focus from end-to-end tasks to atomic skills, significantly reducing data costs while maintaining high performance and enabling efficient adaptation to new tasks. Extensive experiments in real-world settings demonstrate the effectiveness and efficiency of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。