arXiv:2511.11438cs.CV2025-11中稿 · AAAI被引 4

首个系统评估多模态大模型视觉提示理解能力的基准测试

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

  • 设计双阶段评测框架,覆盖30k个视觉提示样本
  • 发现模型对提示形状和属性敏感度差异显著
  • 适合研究视觉提示、多模态推理的学者与开发者

多模态大语言模型(MLLMs)已广泛应用于细粒度目标识别和上下文理解等任务。当用户需查询图像中特定区域或对象时,通常会使用边界框等视觉提示(VP)作为参考。然而,现有基准尚未系统评估MLLMs对这类视觉提示的理解能力。为此,本文提出VP-Bench,一个用于评估MLLM在视觉提示感知与利用方面能力的基准。该基准采用两阶段评估:第一阶段考察模型在自然场景中感知视觉提示的能力,包含30,000个视觉提示,覆盖八种形状与355种属性组合;第二阶段研究视觉提示对下游任务的影响,衡量其在真实问题解决场景中的有效性。我们使用该基准评估了28个MLLM,包括GPT-4o等专有模型与InternVL3、Qwen2.5-VL等开源模型,并分析了提示属性变化、问题排列方式及模型规模等因素对视觉提示理解的影响。VP-Bench为研究MLLM如何理解并解答具身指代问题提供了新范式。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "visual prompts" (VPs), such as bounding boxes, to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap leaves it unclear whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and use them to solve problems. To address this limitation, we introduce VP-Bench, a benchmark for assessing MLLMs' capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models' ability to perceive VPs in natural scenes, using 30k visualized prompts spanning eight shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 28 MLLMs, including proprietary systems (e.g., GPT-4o) and open-source models (e.g., InternVL3 and Qwen2.5-VL), and provide a comprehensive analysis of factors that affect VP understanding, such as variations in VP attributes, question arrangement, and model scale. VP-Bench establishes a new reference framework for studying how MLLMs comprehend and resolve grounded referring questions.

视觉提示多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。