提出细粒度3D指代表达分割新任务,实现短语到3D实例的精准映射。
3D-DRES: Detailed 3D Referring Expression Segmentation
- 设计短语-实例标注范式,精确匹配自然语言短语与3D元素
- 构建含54,432条描述的DetailRefer数据集,覆盖11,054个物体
- 提出Dual-mode基线模型,同时支持句级与短语级分割
当前3D视觉定位任务仅处理句子级别的检测或分割,难以利用自然语言中丰富的组合上下文推理。为此,我们提出详细3D指代表达分割(3D-DRES)新任务,旨在实现短语到3D实例的映射,提升细粒度3D视觉语言理解能力。为支持该任务,我们构建了DetailRefer数据集,包含54,432条描述,覆盖11,054个不同物体。不同于以往数据集,DetailRefer首次采用短语-实例标注范式,将每个被引用的名词短语明确映射到对应的3D元素。此外,我们提出DetailBase,一个专为双模式分割设计的轻量高效基线架构。实验表明,基于DetailRefer训练的模型不仅在短语级分割上表现优异,还在传统3D-RES基准上取得意外提升。
原文摘要 · Abstract (English)
Current 3D visual grounding tasks only process sentence level detection or segmentation, which critically fails to leverage the rich compositional contextual reasonings within natural language expressions. To address this challenge, we introduce Detailed 3D Referring Expression Segmentation (3D-DRES), a new task that provides a phrase to 3D instance mapping, aiming at enhancing fine-grained 3D vision language understanding. To support 3D-DRES, we present DetailRefer, a new dataset comprising 54,432 descriptions spanning 11,054 distinct objects. Unlike previous datasets, DetailRefer implements a pioneering phrase-instance annotation paradigm where each referenced noun phrase is explicitly mapped to its corresponding 3D elements. Additionally, we introduce DetailBase, a purposefully streamlined yet effective baseline architecture that supports dual-mode segmentation at both sentence and phrase levels. Our experimental results demonstrate that models trained on DetailRefer not only excel at phrase-level segmentation but also show surprising improvements on traditional 3D-RES benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。