arXiv:2410.04659cs.CV2024-10ACL被引 12

评测多模态大模型主动感知能力,发现其普遍不足。

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

  • 通过限制视觉感知范围,让模型主动调整视野来回答问题。
  • 30个模型测试显示,主动感知能力普遍较弱,存在明显差距。
  • 适合关注模型推理与环境交互的AI研究者使用。

主动感知是人类关键能力之一,指基于对环境的理解设定目标并采取行动以达成目标。尽管对多模态大语言模型(MLLMs)的评估已有诸多努力,但主动感知仍被严重忽视。为此,我们提出一个名为ActiView的新基准,用于评估MLLMs的主动感知能力。聚焦一种特殊形式的视觉问答(VQA),该形式简化了评估流程并可量化结果,但对现有MLLMs仍具挑战性。给定一张图像,我们限制模型的感知范围,要求其根据推理主动缩放或移动感知区域以成功回答问题。我们在30个模型上进行广泛评估,涵盖专有和开源模型,发现受限的感知范围在促进主动感知中起关键作用。结果揭示了MLLMs在主动感知能力上的显著差距,表明该领域亟需更多关注。我们希望ActiView能推动MLLMs以更自然、整体的方式理解多模态输入。

原文摘要 · Abstract (English)

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked. To address this gap, we propose a novel benchmark named ActiView to evaluate active perception in MLLMs. We focus on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. Meanwhile, intermediate reasoning behaviors of models are also discussed. Given an image, we restrict the perceptual field of a model, requiring it to actively zoom or shift its perceptual field based on reasoning to answer the question successfully. We conduct extensive evaluation over 30 models, including proprietary and open-source models, and observe that restricted perceptual fields play a significant role in enabling active perception. Results reveal a significant gap in the active perception capability of MLLMs, indicating that this area deserves more attention. We hope that ActiView could help develop methods for MLLMs to understand multimodal inputs in more natural and holistic ways.

多模态主动感知视觉问答大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。