arXiv:2410.10799cs.CV2024-10中稿 · 3DV 2025被引 14

构建3D视觉基准测试,发现现有模型距人类理解仍有差距

Towards Foundation Models for 3D Vision: How Close Are We?

  • 设计新基准UniQA-3D,评估3D视觉理解能力
  • 视觉语言模型表现差,专用模型虽准但不鲁棒
  • 神经网络比传统方法更接近人类3D认知机制

构建3D视觉基础模型仍是未解难题。为推进该目标,我们需理解当前模型的3D推理能力,并识别其与人类的差距。为此,我们构建了一个名为UniQA-3D的新3D视觉理解基准,涵盖视觉问答(VQA)格式下的基础3D任务。我们在该基准上评估了前沿视觉语言模型(VLMs)、专用模型及人类被试的表现。结果表明,VLMs整体表现较差,专用模型虽准确但对几何扰动敏感,而人类视觉仍是最可靠的3D视觉系统。进一步分析显示,神经网络在3D视觉机制上比传统计算机视觉方法更接近人类,且基于Transformer的ViT比CNN更贴近人类认知。代码已开源于https://github.com/princeton-vl/UniQA-3D。

原文摘要 · Abstract (English)

Building a foundation model for 3D vision is a complex challenge that remains unsolved. Towards that goal, it is important to understand the 3D reasoning capabilities of current models as well as identify the gaps between these models and humans. Therefore, we construct a new 3D visual understanding benchmark named UniQA-3D. UniQA-3D covers fundamental 3D vision tasks in the Visual Question Answering (VQA) format. We evaluate state-of-the-art Vision-Language Models (VLMs), specialized models, and human subjects on it. Our results show that VLMs generally perform poorly, while the specialized models are accurate but not robust, failing under geometric perturbations. In contrast, human vision continues to be the most reliable 3D visual system. We further demonstrate that neural networks align more closely with human 3D vision mechanisms compared to classical computer vision methods, and Transformer-based networks such as ViT align more closely with human 3D vision mechanisms than CNNs. We hope our study will benefit the future development of foundation models for 3D vision. Code is available at https://github.com/princeton-vl/UniQA-3D .

3D视觉基础模型视觉问答认知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。