构建多样化的3D视觉定位数据,提升模型对复杂语言的泛化能力。
Scaling Diverse Language Generation for 3D Visual Grounding

- 结合场景图约束采样与大模型生成,实现可扩展的多样化查询合成。
- 在多个3DVG基准上性能提升,且生成文本多样性显著优于现有数据集。
- 适合研究3D视觉语言模型泛化能力的学者,尤其关注语言多样性的方向。
3D视觉定位(3DVG)旨在将自然语言描述中的实体与三维场景中的物体对应起来,是使智能体理解空间语言的关键。然而,缺乏大规模且多样化的描述限制了模型对复杂语言模式的泛化能力。现有方法在约束类型和语言表达上多样性不足,图像字幕生成方式难以精确区分物体,影响定位精度。为此,我们提出ViGiL3D++,一种可扩展、场景无关的方法,通过结合场景图中的约束采样与大语言模型(LLM)的语言生成能力,生成多样化的视觉定位查询。实验表明,该方法生成的数据在多样性上优于现有规模化数据集,并在多个3DVG基准上提升了模型性能,同时揭示了视觉语言模型(VLMs)仍存在的关键局限性。
原文摘要 · Abstract (English)
Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for enabling agents to correspond spatial language with objects in the physical world. However, the lack of diverse descriptions at scale prevents models from generalizing beyond simple linguistic patterns. Recent such attempts lack diversity in the constraint types and language used to ground objects. Captioning methods cannot precisely contrast objects, which is important for visual grounding. We therefore propose ViGiL3D++, a scalable, scene-agnostic method that generates diverse visual grounding queries by combining constraint sampling in scene graphs with the language generation of LLMs. We show that it has greater diversity over existing scaled datasets and improves model performance over several 3DVG benchmarks but also illuminates outstanding limitations of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。