arXiv:2511.23230cs.CV2025-11中稿 · CVPR被引 1

用动作描述自动生成3D功能分割数据,解决标注难问题。

Action-guided generation of 3D functionality segmentation data

  • 根据动作描述构建场景,自动合成带精确掩码的3D数据
  • 合成数据+真实数据训练,性能提升2.2 mAP、6.3 mAR、5.7 mIoU
  • 适合做3D交互理解与智能机器人场景解析的研究者

3D功能分割旨在识别3D场景中执行自由语言描述动作(如“打开床边柜子的第二个抽屉”)所需的交互元素。该任务受限于真实世界标注数据稀缺,因精细3D掩码的采集与标注成本过高。为此,我们提出SynthFun3D,首个直接从动作描述生成3D功能分割数据的方法。给定动作描述,SynthFun3D从大规模资产库中检索带部件级标注的物体,在空间与语义约束下构建合理3D场景,并渲染多视角图像,自动识别目标功能元素,生成无需人工标注的精确真值掩码。我们通过训练基于视觉语言模型的3D功能分割模型验证了生成数据的有效性:用合成数据增强真实数据后,性能相较纯真实数据训练提升+2.2 mAP、+6.3 mAR、+5.7 mIoU。结果表明,动作引导的合成数据生成为3D功能理解提供了可扩展且高效的注释补充方案。

原文摘要 · Abstract (English)

3D functionality segmentation aims to identify the interactive element in a 3D scene required to perform an action described in free-form language (e.g., the handle to ``Open the second drawer of the cabinet near the bed''). Progress has been constrained by the scarcity of annotated real-world data, as collecting and labeling fine-grained 3D masks is prohibitively expensive. To address this limitation, we introduce SynthFun3D, the first method for generating 3D functionality segmentation data directly from action descriptions. Given an action description, SynthFun3D constructs a plausible 3D scene by retrieving objects with part-level annotations from a large-scale asset repository and arranging them under spatial and semantic constraints. SynthFun3D renders multi-view images and automatically identifies the target functional element, producing precise ground-truth masks without manual annotation. We demonstrate the effectiveness of the generated data by training a VLM-based 3D functionality segmentation model. Augmenting real-world data with our synthetic data consistently improves performance, with gains of +2.2 mAP, +6.3 mAR, and +5.7 mIoU over real-only training. This shows that action-guided synthetic data generation provides a scalable and effective complement to manual annotation for 3D functionality understanding. Project page: tev-fbk.github.io/synthfun3d.

3D分割动作理解合成数据视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。