arXiv:2409.03757cs.CVcs.AI2024-09NeurIPS被引 54

对比七种视觉模型,揭示3D场景理解中各模型的优劣与适用场景。

Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

  • 测试7种视觉基础模型在4类任务中的表现,涵盖图像、视频与3D模型。
  • DINOv2表现最优,视频模型擅长物体级任务,扩散模型利于几何理解。
  • 语言预训练模型在语言任务中意外表现差,提示需灵活选择编码器。

复杂3D场景理解受到越来越多关注,场景编码策略在此过程中起关键作用。然而,不同场景下最优编码策略仍不明确,尤其相较于基于图像的方法。为此,我们对多种视觉编码模型在3D场景理解中的表现进行了全面探究,识别出各类模型在不同场景中的优势与局限。评估涵盖七种视觉基础编码器,包括基于图像、视频和3D的基础模型。在四个任务中进行测试:视觉-语言场景推理、视觉定位、分割与配准,分别聚焦于场景理解的不同方面。关键发现包括:DINOv2表现最佳,视频模型在物体级任务中占优,扩散模型在几何任务中受益,而语言预训练模型在语言相关任务中表现出意料之外的局限性。这些结果挑战了部分传统认知,为视觉-语言与场景理解任务中模型选择提供了新视角,并强调未来需要更灵活的编码器适配策略。

原文摘要 · Abstract (English)

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-based counterparts. To address this issue, we present a comprehensive study that probes various visual encoding models for 3D scene understanding, identifying the strengths and limitations of each model across different scenarios. Our evaluation spans seven vision foundation encoders, including image-based, video-based, and 3D foundation models. We evaluate these models in four tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration, each focusing on different aspects of scene understanding. Our evaluations yield key findings: DINOv2 demonstrates superior performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, and language-pretrained models show unexpected limitations in language-related tasks. These insights challenge some conventional understandings, provide novel perspectives on leveraging visual foundation models, and highlight the need for more flexible encoder selection in future vision-language and scene-understanding tasks. Code: https://github.com/YunzeMan/Lexicon3D

3D理解视觉模型多模态编码器评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。