arXiv:2505.21652cs.ROcs.AI2025-05被引 9

首个支持零件级指令的机器人精细操作基准,提升语言指令与物体部件的对齐能力。

PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

  • 构建包含513个物体、1302项任务的零件级操作数据集,支持高阶任务分解。
  • 在10000+仿真示范中验证模型对零件概念和3D动作的掌握能力,发现现有模型表现不足。
  • 适合研究视觉-语言导航、机器人操作泛化、细粒度任务规划的研究者使用。

精细机器人操作(如旋转瓶子展示瓶盖标签)需要对物体部件及其与任务关系的鲁棒推理。尽管已有基于语言指令训练通用机器人操作策略的进展,但缺乏大规模带零件级标注的细粒度操作数据集。本文提出PartInstruct,首个基于零件级指令的细粒度机器人操作基准。该数据集包含14类共513个物体实例,每例均标注零件级信息,并涵盖1302项细粒度操作任务,分为16类任务。训练集由超过10,000个在3D模拟器中合成的专家示范构成,每个示范配有高层任务指令、一系列基础部件技能指令及物体与部件的真值3D信息。此外,设计了全面测试套件,评估模型在新状态、新物体和新任务下的泛化能力。我们评估了多种先进机器人操作方法,包括端到端视觉-语言策略学习与双层规划模型。实验结果表明,当前模型在3D空间中对零件概念的语义接地和动作预测上仍存在显著困难,尤其在长时序任务中表现不佳。

原文摘要 · Abstract (English)

Fine-grained robot manipulation, such as lifting and rotating a bottle to display the label on the cap, requires robust reasoning about object parts and their relationships with intended tasks. Despite recent advances in training general-purpose robot manipulation policies guided by language instructions, there is a notable lack of large-scale datasets for fine-grained manipulation tasks with part-level instructions and diverse 3D object instances annotated with part-level labels. In this work, we introduce PartInstruct, the first large-scale benchmark for training and evaluating fine-grained robot manipulation models using part-level instructions. PartInstruct comprises 513 object instances across 14 categories, each annotated with part-level information, and 1302 fine-grained manipulation tasks organized into 16 task classes. Our training set consists of over 10,000 expert demonstrations synthesized in a 3D simulator, where each demonstration is paired with a high-level task instruction, a chain of base part-based skill instructions, and ground-truth 3D information about the object and its parts. Additionally, we designed a comprehensive test suite to evaluate the generalizability of learned policies across new states, objects, and tasks. We evaluated several state-of-the-art robot manipulation approaches, including end-to-end vision-language policy learning and bi-level planning models for robot manipulation on our benchmark. The experimental results reveal that current models struggle to robustly ground part concepts and predict actions in 3D space, and face challenges when manipulating object parts in long-horizon tasks.

机器人操作零件级指令细粒度控制3D仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。