arXiv:2505.21955cs.CVcs.AI2025-05NeurIPS被引 6

融合第一人称与第三人称视角,提升视觉语言模型的场景理解能力。

Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs

  • 用多视角融合构建统一场景表征,增强跨视图推理能力。
  • 在4000对高质量数据上测试,性能提升达4.84%至5.94%。
  • 适合研究多视角感知、虚拟现实交互的开发者和研究人员。

大型视觉语言模型(LVLM)正被广泛应用于虚拟现实、增强现实等交互式场景中,其中头戴设备捕捉的第一人称(即自身体验)视角是关键输入。尽管该视角能提供用户注意力与手物交互的细粒度线索,但其视野狭窄且缺乏全局上下文,常导致空间或情境复杂问题的失败。为此,我们提出一种框架,通过引入第三人称(即外部观察)视角来补充信息,如全局场景布局与物体可见性。我们构建了E3VQA,首个基于同步第一人称-第三人称图像对的多视角问答基准,包含4000个高质量问答对。同时提出M3CoT,一种无需训练的提示技术,通过整合来自三个互补视角的场景图,构建统一的场景表示。M3CoT使LVLM在跨视图推理中表现更优,在GPT-4o上提升4.84%,在Gemini 2.0 Flash上提升5.94%,优于近期的CoT基线。大量评估揭示了LVLM在多视角推理中的强项与局限,并验证了融合双视角输入的价值。数据集与源码已公开于https://github.com/Leeinsu1/Towards-Comprehensive-Scene-Understanding。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object interactions, its narrow field of view and lack of global context often lead to failures on spatially or contextually demanding queries. To address this, we introduce a framework that augments egocentric inputs with third-person (exocentric) views, providing complementary information such as global scene layout and object visibility to LVLMs. We present E3VQA, the first benchmark for multi-view question answering with 4K high-quality question-answer pairs grounded in synchronized ego-exo image pairs. Additionally, we propose M3CoT, a training-free prompting technique that constructs a unified scene representation by integrating scene graphs from three complementary perspectives. M3CoT enables LVLMs to reason more effectively across views, yielding consistent performance gains (4.84% for GPT-4o and 5.94% for Gemini 2.0 Flash) over a recent CoT baseline. Our extensive evaluation reveals key strengths and limitations of LVLMs in multi-view reasoning and highlights the value of leveraging both egocentric and exocentric inputs. The dataset and source code are available at https://github.com/Leeinsu1/Towards-Comprehensive-Scene-Understanding.

多视角理解视觉语言模型场景推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。