arXiv:2503.12329cs.CVcs.CL2025-03ACL被引 45

评测大模型图像描述能力,发现顶级模型已超人类水平。

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

  • 构建6000+对比评测对,用人类偏好判断模型表现
  • GPT-4o表现达或超越人类,多数开源模型落后
  • 提出自动化评测工具,与人工评分相关性94.3%

图像描述是视觉语言研究中的长期挑战。随着大语言模型(LLM)兴起,现代视觉语言模型(VLM)能生成详尽全面的图像描述。然而,如何评估此类描述质量仍不明确。本文回答两个关键问题:(1) 当前VLM在图像描述上的实际表现如何,尤其与人类相比?我们构建了CapArena平台,包含超过6000个成对的描述对比和高质量的人类偏好投票。该竞技场式评估标志着重要进展:领先模型如GPT-4o的表现达到甚至超越人类水平,而大多数开源模型仍落后。(2) 自动化评价指标能否可靠评估详细描述质量?基于CapArena的人类标注,我们评估了传统及最新描述评价指标,以及以VLM作为裁判的方法。分析显示,尽管部分指标(如METEOR)在句子层面与人类评价有一定一致性,但其系统性偏差导致模型排名不一致;相比之下,VLM-as-a-Judge在句子和模型层面均展现出稳健判别力。基于这些发现,我们推出CapArena-Auto,一种准确高效的自动化评测工具,在仅需4美元/测试的情况下,实现与人工评分94.3%的相关性。数据与资源将开源至https://caparena.github.io。

原文摘要 · Abstract (English)

Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive image descriptions. However, benchmarking the quality of such captions remains unresolved. This paper addresses two key questions: (1) How well do current VLMs actually perform on image captioning, particularly compared to humans? We built CapArena, a platform with over 6000 pairwise caption battles and high-quality human preference votes. Our arena-style evaluation marks a milestone, showing that leading models like GPT-4o achieve or even surpass human performance, while most open-source models lag behind. (2) Can automated metrics reliably assess detailed caption quality? Using human annotations from CapArena, we evaluate traditional and recent captioning metrics, as well as VLM-as-a-Judge. Our analysis reveals that while some metrics (e.g., METEOR) show decent caption-level agreement with humans, their systematic biases lead to inconsistencies in model ranking. In contrast, VLM-as-a-Judge demonstrates robust discernment at both the caption and model levels. Building on these insights, we release CapArena-Auto, an accurate and efficient automated benchmark for detailed captioning, achieving 94.3% correlation with human rankings at just $4 per test. Data and resources will be open-sourced at https://caparena.github.io.

图像描述大模型评测自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。