arXiv:2509.25818cs.CVcs.AI2025-09EMNLP被引 7

用混合大模型评估长图文描述,速度更快更准。

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

  • 采用非自回归的混合框架,避免视觉信息过早融合
  • 在7805张图上超越现有方法,表现超人类水平
  • 专为长图文设计,适合评测多模态生成质量

本研究聚焦于多模态大语言模型(MLLMs)生成的长篇详细图文描述的自动评估。现有图像描述评估指标大多针对短文本设计,不适用于长描述。此外,近期基于大模型评判(LLM-as-a-Judge)的方法因依赖自回归推理和早期视觉信息融合,存在推理缓慢的问题。为此,我们提出VELA,一种基于新型LLM-Hybrid-as-a-Judge框架的长图文描述自动评估指标。同时构建了LongCap-Arena基准,包含7,805张图像、对应的人工撰写的长参考描述与候选描述,以及来自三个视角(描述性、相关性、流畅性)的32,246条人工判断。实验表明,VELA在该基准上表现优于现有指标,并达到超人类性能。

原文摘要 · Abstract (English)

In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions. Moreover, recent LLM-as-a-Judge approaches suffer from slow inference due to their reliance on autoregressive inference and early fusion of visual information. To address these limitations, we propose VELA, an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework. Furthermore, we propose LongCap-Arena, a benchmark specifically designed for evaluating metrics for long captions. This benchmark comprises 7,805 images, the corresponding human-provided long reference captions and long candidate captions, and 32,246 human judgments from three distinct perspectives: Descriptiveness, Relevance, and Fluency. We demonstrated that VELA outperformed existing metrics and achieved superhuman performance on LongCap-Arena.

图文生成评估方法大模型评测长文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。