自动为3D物体添加机器人操作所需语义与物理标签
AnnotateAnything: Automatic Annotation of 3D Assets for Robot Manipulation

- 用视觉语言模型推断物体交互区域和操作约束
- 自动生成抓取、插入、悬挂等多样化可执行动作
- 支持多机器人、多任务并行仿真,提升数据效率
仿真可实现机器人数据的规模化收集,但原始3D资产仅提供几何信息,缺乏语义、交互和物理知识以明确机器人的操作位置与方式。本文提出AnnotateAnything,一个通用的自动标注框架,将静态3D资产转化为具备结构化、多样性和可执行性操作标签的可操作资产。该框架包含两条互补流水线:第一,统一的视觉-语言标注流程,利用视觉语言推理推断物体语义、交互约束和3D定位线索,为人类先验提供交互区域识别指导;第二,全自动、大规模并行的物理标注流程,通过候选生成、几何优化和轨迹生成,将上述先验在每项资产的几何与物理约束中具象化。该流程生成多样化且可执行的动作标注,包括抓取位姿、灵巧接触点、关节运动路径、插入方向、悬挂可达性及导航目标。基于生成的标注,我们构建了跨多样化物体、任务与机器人形态的异步并行仿真数据采集系统。实验表明,AnnotateAnything在标注效率、数据采集效率和任务成功率方面均优于现有方法,同时支持下游任务如可达性检测、机器人视觉问答与视觉指令微调。项目材料已公开,计划发布完整代码、标注数据与基准测试集。视频、代码、演示资产与标注见补充材料。项目页面:https://tourmaline-caramel-169490.netlify.app。
原文摘要 · Abstract (English)
Simulation enables scalable robot data collection, but raw 3D assets provide only geometry, lacking the semantic, interactive, and physical knowledge needed to specify where and how robots should act. In this work, we present AnnotateAnything, a general automatic annotation framework that converts passive 3D assets into manipulation-ready assets with structured, diverse, and executable manipulation labels. AnnotateAnything is built around two complementary pipelines. First, a unified visual-language annotation pipeline using vision-language reasoning to infer object semantics, interaction constraints, and 3D-grounded cues, providing human-prior guidance for identifying meaningful interaction regions. Second, a fully automatic and massively parallel physics annotation pipeline grounds these priors in each asset's geometry and physical constraints through candidate generation, geometry optimization and trajectory generation. This pipeline produces diverse and executable action annotations, including grasp poses, dexterous contacts, articulation waypoints, insertion directions, hanging affordances, and navigation targets. Using the generated annotations, we further build an asynchronous parallel simulation data-collection system across diverse objects, tasks, and robot embodiments. Experiments demonstrate that AnnotateAnything achieves superior annotation efficiency, data-collection efficiency, and task success rates over existing annotation and data-generation pipelines, while also supporting downstream tasks such as affordance detection, robotic VQA, and visual instruction finetuning. We provide project materials on the project page and plan to release the full code, annotations, and benchmark to facilitate future research. Videos, code, demo assets, and annotations are provided in supplementary materials Project page: https://tourmaline-caramel-169490.netlify.app.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。