arXiv:2512.21582cs.CV2025-12AAAI被引 4

提出无需大模型的图像描述评估方法,兼顾参考与无参考场景

LLM-Free Image Captioning Evaluation in Reference-Flexible Settings

  • 设计新机制学习图文与文间相似性表示
  • 在多个数据集上超越现有无大模型评估指标
  • 基于33万条人工标注,适合评估系统研发者使用

我们关注在有参考和无参考两种场景下图像描述的自动评估问题。基于大语言模型(LLM)的现有评估指标会倾向自身生成结果,其公平性存疑;而多数无大模型指标虽避免此问题,但性能并不稳定。为此,我们提出一种名为Pearl的无大模型监督评估指标,适用于两种设置。该方法引入新机制,学习图像-描述与描述-描述间的相似性表征。此外,我们构建了一个人工标注的数据集,包含来自2,360名标注者对超过7.5万张图像的约33.3万条人类判断。Pearl在Composite、Flickr8K-Expert、Flickr8K-CF、Nebula和FOIL数据集上,无论在有参考还是无参考设置中,均优于其他现有无大模型评估指标。

原文摘要 · Abstract (English)

We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics, that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings. Our project page is available at https://pearl.kinsta.page/.

图像描述评估指标无大模型多场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。