提出新图像描述评估指标,更贴近人类判断标准。
VCRScore: Image captioning metric based on V\&L Transformers, CLIP, and precision-recall
- 基于视觉语言模型与CLIP构建新型评估框架
- 在人工标注数据上验证,显著优于传统指标
- 适合评估先进图像生成模型的性能
图像描述已成为视觉与语言研究的重要任务,旨在根据给定图像或视频生成最准确的描述。尽管近年来模型和方法不断进步,性能评估指标却长期停滞不前。当前仍广泛使用BLEU、METEOR、CIDEr、ROUGE等经典指标,甚至包括BertScore和ClipScore等较新指标。因此,亟需改进对新模型进展、局限性和适用范围的衡量方式。本文提出一种新的图像描述评估指标。首先构建了一个人工标注数据集,用以评估描述与图像内容的相关性;以人工评分作为真实标签,提出新指标并与其他主流指标进行对比。实验结果表明,新指标表现更优,并揭示了重要洞察。
原文摘要 · Abstract (English)
Image captioning has become an essential Vision & Language research task. It is about predicting the most accurate caption given a specific image or video. The research community has achieved impressive results by continuously proposing new models and approaches to improve the overall model's performance. Nevertheless, despite increasing proposals, the performance metrics used to measure their advances have remained practically untouched through the years. A probe of that, nowadays metrics like BLEU, METEOR, CIDEr, and ROUGE are still very used, aside from more sophisticated metrics such as BertScore and ClipScore. Hence, it is essential to adjust how are measure the advances, limitations, and scopes of the new image captioning proposals, as well as to adapt new metrics to these new advanced image captioning approaches. This work proposes a new evaluation metric for the image captioning problem. To do that, first, it was generated a human-labeled dataset to assess to which degree the captions correlate with the image's content. Taking these human scores as ground truth, we propose a new metric, and compare it with several well-known metrics, from classical to newer ones. Outperformed results were also found, and interesting insights were presented and discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。