arXiv:2501.03567cs.CV2025-01被引 5

用图文循环生成评估图片描述,更贴近人类判断。

Evaluating Image Caption via Cycle-consistent Text-to-Image Generation

  • 通过文本生成图像再对比原图,避开跨模态鸿沟。
  • 在多个数据集上与人工评价相关性超越现有方法。
  • 分像素、语义、目标三层评估,细节更全面。

图像描述的评估通常依赖参考描述,但获取成本高且存在显著差异与主观性。尽管已有无参考评估指标,多数仍聚焦于描述与图像间的跨模态匹配。近期研究发现,基于对比学习的多模态系统中普遍存在的模态差距,削弱了如CLIPScore等跨模态指标的可靠性。本文提出CAMScore,一种循环式无参考自动图像描述评估指标。为规避上述模态差距,CAMScore利用文本到图像模型从描述生成图像,并将生成图像与原始图像进行评估。此外,为提供更细致的评估信息,设计了包含像素级、语义级和目标级的三层评估框架。在多个基准数据集上的大量实验结果表明,CAMScore相较于现有参考型与无参考型指标,与人类判断具有更优相关性,验证了该框架的有效性。

原文摘要 · Abstract (English)

Evaluating image captions typically relies on reference captions, which are costly to obtain and exhibit significant diversity and subjectivity. While reference-free evaluation metrics have been proposed, most focus on cross-modal evaluation between captions and images. Recent research has revealed that the modality gap generally exists in the representation of contrastive learning-based multi-modal systems, undermining the reliability of cross-modality metrics like CLIPScore. In this paper, we propose CAMScore, a cyclic reference-free automatic evaluation metric for image captioning models. To circumvent the aforementioned modality gap, CAMScore utilizes a text-to-image model to generate images from captions and subsequently evaluates these generated images against the original images. Furthermore, to provide fine-grained information for a more comprehensive evaluation, we design a three-level evaluation framework for CAMScore that encompasses pixel-level, semantic-level, and objective-level perspectives. Extensive experiment results across multiple benchmark datasets show that CAMScore achieves a superior correlation with human judgments compared to existing reference-based and reference-free metrics, demonstrating the effectiveness of the framework.

图像描述评估指标无参考生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。