arXiv:2410.12332cs.CV2024-10ICCV被引 59

评测大模型跨图像定位目标的能力,发现性能与人类差距明显。

MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs

  • 设计多上下文视觉定位任务,需在多图中根据自然语言描述找目标。
  • 构建2000个高质量标注样本,覆盖20种实用技能和三类开放文本风格。
  • 对比20+主流模型,揭示当前大模型在多图任务中的显著不足。

尽管多模态大语言模型(MLLMs)展现出卓越的视觉-语言理解能力,但其在单张图像之外解决实例级视觉-语言问题的能力仍待深入探索。为此,本文提出一种新的视觉定位任务——多上下文视觉定位,旨在基于开放式文本提示,在多张图像中定位感兴趣的目标实例。为推动该研究,我们构建了名为MC-Bench的新数据集,包含2000个高质量、人工标注的样本。每个样本由一对实例级标注图像及对应文本提示组成,文本提示高度开放,涵盖三种不同风格,覆盖20种实际应用场景。我们对超过20个最先进的MLLMs及基础模型进行了基准测试,包括自研的简单有效代理基线和通过多上下文指令微调的基线模型。评估结果揭示现有MLLMs与人类之间存在显著性能差距,并提供了若干具有启发性的观察,指明未来研究方向。我们希望MC-Bench及其发现能激励学术界进一步挖掘MLLMs在实例级任务、特别是多图像场景下的未充分利用潜力。项目页面:https://xuyunqiu.github.io/MC-Bench。

原文摘要 · Abstract (English)

While multimodal large language models (MLLMs) have demonstrated extraordinary vision-language understanding capabilities, their abilities to solve instance-level visual-language problems beyond a single image warrant further exploration. To assess these unproven abilities of MLLMs, this paper proposes a new visual grounding task called multi-context visual grounding, which aims to localize instances of interest across multiple images based on open-ended text prompts. In order to facilitate this research, we construct a new dataset MC-Bench that features 2K high-quality and manually annotated samples. Each sample consists of an instance-level labeled image pair and a corresponding text prompt that indicates the target instances in the images. These text prompts are highly open-ended and follow three distinct styles, covering 20 practical skills. We benchmark over 20 state-of-the-art MLLMs and foundation models with potential multi-context visual grounding capabilities, along with our developed simple yet effective agentic baseline and a finetuned baseline by multi-context instruction tuning. Our evaluation reveals a non-trivial performance gap between existing MLLMs and humans, along with some insightful observations that suggest potential future directions. We hope that MC-Bench and our empirical findings encourage the research community to further advance the untapped potentials of MLLMs in instance-level tasks, particularly in multi-image contexts. Project page: https://xuyunqiu.github.io/MC-Bench.

视觉定位多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。