arXiv:2506.16805cs.CV2025-06ICCV被引 2

提出新基准评估视觉模型在稀疏图像中识别共同可见区域的能力

Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes

  • 构建1000+个室内场景的稀疏视图数据集,评测跨视角共可见性推理能力
  • 现有模型性能远低于人类,即使领先视觉语言模型也差距显著
  • 设计新型多视角模型Covis,融合空间与语义信息提升推理表现

人类具备在复杂场景中从稀疏分布的多幅图像里识别共同可见3D区域的卓越能力,这构成了3D视觉与机器人感知的基础,不仅依赖低层特征匹配,更需高层空间推理与认知整合。然而,当前视觉模型能否复现这种人类级能力尚不明确。本文提出Co-VisiON基准,用于评估超过1,000个稀疏视图室内场景中的人类启发式共可见性推理能力。结果表明,尽管共可见性常被视为低层特征匹配任务,但在稀疏条件下仍对现有视觉模型构成挑战。值得注意的是,一个专有视觉语言模型优于所有纯视觉基线,但所有模型与人类表现仍有显著差距。这一差距凸显了现有架构的局限性,并推动发展能以类人方式整合空间与语义信息的模型。受人类视觉认知启发,我们提出一种新型多视角基线Covis,成为纯视觉模型中的最佳表现者,缩小了与专有VLM的差距。我们希望该基准与发现能推动更具鲁棒性、认知启发式推理能力的视觉模型发展。数据集与源代码详见https://ai4ce.github.io/CoVISION。

原文摘要 · Abstract (English)

Humans exhibit a remarkable ability to recognize co-visibility-the 3D regions simultaneously visible in multiple images-even when these images are sparsely distributed across a complex scene. This ability is foundational to 3D vision, robotic perception, and relies not only on low-level feature matching but also on high-level spatial reasoning and cognitive integration. Yet, it remains unclear whether current vision models can replicate this human-level proficiency. In this work, we introduce the Co-VisiON benchmark, designed to evaluate human-inspired co-visibility reasoning across more than 1,000 sparse-view indoor scenarios. Our results show that while co-visibility is often approached as a low-level feature-matching task, it remains challenging for existing vision models under sparse conditions. Notably, a proprietary vision-language model surpasses all vision-only baselines, but all models fall significantly short of human performance. This gap underscores the limitations of current architectures and motivates the need for models that integrate spatial and semantic information in a human-like manner. Inspired by human visual cognition, we propose a novel multi-view baseline, Covis, which achieves top performance among pure vision models and narrows the gap to the proprietary VLM. We hope our benchmark and findings will spur further advancements in developing vision models capable of robust, cognitively inspired reasoning in challenging, sparse environments. Our dataset and source code can be found at https://ai4ce.github.io/CoVISION.

3D视觉多视图推理认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。