通过跨模态增强和空间关系建模,提升3D视觉定位的准确性与泛化能力。
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
- 利用基础模型生成多样化的文本-3D配对数据,扩充训练集
- 设计语言-空间自适应解码器,融合语义与空间关系信息
- 适用于多种3D视觉定位任务,尤其适合数据稀缺场景
3D视觉定位旨在将自然语言描述与3D场景中的目标物体对应起来,是一项重要但具有挑战性的任务。尽管该领域已有进展,现有方法普遍存在文本-3D样本数量少、多样性不足的问题,且未能有效利用3D空间中的丰富空间关系等上下文线索。为此,我们提出AugRefer,一种推进3D视觉定位的新方法。AugRefer引入跨模态增强策略,通过将物体放置于3D场景中,并利用基础模型生成准确且语义丰富的描述,大规模生成多样化的文本-3D配对。这些生成的数据可被任意现有3DVG方法用于扩充训练集。此外,AugRefer设计了语言-空间自适应解码器,能根据语言描述和多种3D空间关系动态调整候选目标。在三个基准数据集上的大量实验充分验证了其有效性。
原文摘要 · Abstract (English)
3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly encounter a shortage: a limited amount and diversity of text3D pairs available for training. Moreover, they fall short in effectively leveraging different contextual clues (e.g., rich spatial relations within the 3D visual space) for grounding. To address these limitations, we propose AugRefer, a novel approach for advancing 3D visual grounding. AugRefer introduces cross-modal augmentation designed to extensively generate diverse text-3D pairs by placing objects into 3D scenes and creating accurate and semantically rich descriptions using foundation models. Notably, the resulting pairs can be utilized by any existing 3DVG methods for enriching their training data. Additionally, AugRefer presents a language-spatial adaptive decoder that effectively adapts the potential referring objects based on the language description and various 3D spatial relations. Extensive experiments on three benchmark datasets clearly validate the effectiveness of AugRefer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。