评测大模型在多视角场景下的理解能力,发现其远未达到人类水平。
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- 构建包含2100个问答对的多视角基准测试集
- 模型在遮挡视图对应和相机位姿估计上表现较差
- 适合研究多模态模型视觉推理与具身智能的学者参考
多视角理解是多模态大语言模型(MLLMs)作为具身智能体实现有效导航、操作与三维场景认知的基础挑战。尽管近期MLLM在高层推理与规划方面取得显著进展,但在多视角几何一致性与跨视角对应关系上仍表现不佳。为此,我们提出All-Angles Bench,一个涵盖90个真实场景、超过2100个人工精心标注的多视角问答对的基准测试。该基准包含六项任务:计数、属性识别、相对距离、相对方向、物体操作与相机位姿估计,专门检验模型的几何对应能力与跨视角信息一致性。我们在27个代表性MLLM(包括Gemini-2.0-Flash、Claude-3.7-Sonnet、GPT-4o)上进行广泛实验,并与人类评估者对比,揭示了显著性能差距,表明当前MLLM尚未达到人类水平。深入分析显示,模型在部分遮挡视图的跨视角对应以及粗粒度相机位姿建立方面尤为薄弱。这些发现强调了引入领域特定优化或嵌入更强多视角感知模块的必要性。我们相信All-Angles Bench为缩小MLLM与人类级多视角理解之间的差距提供了重要洞见。项目及基准已公开于https://danielchyeh.github.io/All-Angles-Bench/。
原文摘要 · Abstract (English)
Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown impressive advances in high-level reasoning and planning, they frequently fall short when confronted with multi-view geometric consistency and cross-view correspondence. To comprehensively evaluate the challenges of MLLMs in multi-view scene reasoning, we propose All-Angles Bench, a benchmark of over 2,100 human carefully annotated multi-view question-answer pairs across 90 diverse real-world scenes. Our six tasks (counting, attribute identification, relative distance, relative direction, object manipulation, and camera pose estimation) specifically test model's geometric correspondence and the capacity to align information consistently across views. Our extensive experiments, benchmark on 27 representative MLLMs including Gemini-2.0-Flash, Claude-3.7-Sonnet, and GPT-4o against human evaluators reveals a substantial performance gap, indicating that current MLLMs remain far from human-level proficiency. Through in-depth analysis, we show that MLLMs are particularly underperforming under two aspects: (1) cross-view correspondence for partially occluded views and (2) establishing the coarse camera poses. These findings highlight the necessity of domain-specific refinements or modules that embed stronger multi-view awareness. We believe that our All-Angles Bench offers valuable insights and contribute to bridging the gap between MLLMs and human-level multi-view understanding. The project and benchmark are publicly available at https://danielchyeh.github.io/All-Angles-Bench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。