arXiv:2501.01366cs.CVcs.AI2025-01ACL被引 9

构建多样语言3D视觉定位数据集,提升模型对复杂描述的理解能力

ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding

  • 基于语言多样性分析框架构建新数据集
  • 现有模型在复杂提示下准确率显著下降
  • 适合评估真实场景中视觉定位模型的泛化能力

3D视觉定位(3DVG)旨在根据自然语言描述定位3D场景中的目标实体,对具身智能和场景检索具有重要意义。现有研究虽通过大语言模型扩展了3DVG数据规模,但未涵盖英语中可能的完整提示类型。为此,我们提出一种语言学分析框架,构建了面向3D视觉定位的多样化语言数据集ViGiL3D,用于诊断和评估视觉定位方法在不同语言模式下的表现。通过对现有开放词汇3DVG方法的评测发现,这些模型在更具挑战性的分布外提示下表现不佳,难以胜任真实应用场景。

原文摘要 · Abstract (English)

3D visual grounding (3DVG) involves localizing entities in a 3D scene referred to by natural language text. Such models are useful for embodied AI and scene retrieval applications, which involve searching for objects or patterns using natural language descriptions. While recent works have focused on LLM-based scaling of 3DVG datasets, these datasets do not capture the full range of potential prompts which could be specified in the English language. To ensure that we are scaling up and testing against a useful and representative set of prompts, we propose a framework for linguistically analyzing 3DVG prompts and introduce Visual Grounding with Diverse Language in 3D (ViGiL3D), a diagnostic dataset for evaluating visual grounding methods against a diverse set of language patterns. We evaluate existing open-vocabulary 3DVG methods to demonstrate that these methods are not yet proficient in understanding and identifying the targets of more challenging, out-of-distribution prompts, toward real-world applications.

3D视觉定位语言多样性具身智能数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。