arXiv:2409.20551cs.RO2024-09ICRA被引 14

统一建模工具与可动物体的使用属性,提升机器人泛化操作能力

UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models

  • 用视觉语言模型构建物体中心的操纵表征,统一处理工具与可动物
  • 在900类可动物体和600种工具上验证,显著提升真实与仿真环境泛化性能
  • 适合机器人操纵、具身智能研究者,可作为通用基准参考

现有机器人操纵研究受限于对3D运动约束和操作属性的浅层理解。为此,我们提出统一范式UniAff,将物体中心的操纵与任务理解整合为统一框架。构建了包含19类共900个可动物体和12类共600种工具的标注数据集。利用多模态大模型(MLLM)推断操纵任务中的物体中心表征,实现操作属性识别与3D运动约束推理。仿真与真实世界实验表明,UniAff显著提升了机器人对工具和可动物体的泛化操纵能力。项目网站已公开图像、视频、数据集及代码。

原文摘要 · Abstract (English)

Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset, and code are published on the project website at:https://sites.google.com/view/uni-aff/home

机器人操纵视觉语言模型操作属性统一表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。