arXiv:2505.15098cs.ROcs.AI2025-05

仅用10次示范,实现高效灵活的机器人灵巧操作

Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation

  • 通过聚焦物体末端轨迹,构建分层控制流程提升泛化能力
  • 7项真实任务测试中,仅10次示范即超越基线方法表现
  • 适合数据稀缺场景下的灵巧操作训练,尤其适用于复杂环境

从人类示范中学习机器人操作可快速获取技能,但往往在不同场景和物体位置下泛化能力不足,限制了实际应用,尤其在需要灵巧操作的复杂任务中。视觉-语言-动作(VLA)范式虽依赖大规模数据提升泛化性,但受限于数据稀缺,性能仍有限。本文提出一种名为对象聚焦代理(Object-Focus Actor, OFA)的新方法,实现数据高效的通用灵巧操作。OFA利用灵巧操作任务中一致的末端轨迹特性,实现高效策略训练。其采用分层流程:物体感知与位姿估计、预操作位姿到达及OFA策略执行,确保在不同背景和布局下操作聚焦且高效。七项真实世界任务的综合实验表明,OFA在位置与背景泛化测试中显著优于基线方法,且仅需10次示范即可实现鲁棒性能,凸显其数据效率。

原文摘要 · Abstract (English)

Robot manipulation learning from human demonstrations offers a rapid means to acquire skills but often lacks generalization across diverse scenes and object placements. This limitation hinders real-world applications, particularly in complex tasks requiring dexterous manipulation. Vision-Language-Action (VLA) paradigm leverages large-scale data to enhance generalization. However, due to data scarcity, VLA's performance remains limited. In this work, we introduce Object-Focus Actor (OFA), a novel, data-efficient approach for generalized dexterous manipulation. OFA exploits the consistent end trajectories observed in dexterous manipulation tasks, allowing for efficient policy training. Our method employs a hierarchical pipeline: object perception and pose estimation, pre-manipulation pose arrival and OFA policy execution. This process ensures that the manipulation is focused and efficient, even in varied backgrounds and positional layout. Comprehensive real-world experiments across seven tasks demonstrate that OFA significantly outperforms baseline methods in both positional and background generalization tests. Notably, OFA achieves robust performance with only 10 demonstrations, highlighting its data efficiency.

灵巧操作数据效率机器人学习视觉-语言-动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。