不依赖参考句子,用重建效果评估图像描述质量。
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

- 基于重建能力评估描述质量,强调语义等效而非像素匹配。
- 通过下游视觉语言任务验证重建结果的语义一致性。
- 提出无需人工参考的评测框架,适合模型自评与自动化评估。
图像描述是视觉-语言研究的核心任务,但如何在不依赖人工参考描述的情况下评估描述对图像语义的忠实度仍无定论。现有评估方法依赖人工标注的参考句,其内容受标注者意图和描述能力影响。本文研究基于重建的评估原则:描述的好坏取决于其能否恢复原始图像。然而,由于描述会压缩视觉信息,无法完全还原细节,因此像素级对比既不可行也不合理。我们深入分析描述的本质——传递图像的语义内容,提出新原则:描述好坏取决于其能否生成语义等价的重建图像。为此,我们通过一系列下游视觉-语言任务测试重建图像与原图的语义一致性,构建了一种无参考、任务相关的描述评分机制。同时,我们分析了组件依赖性局限,并引入低成本的《描述图灵测试数据集》(CTTD)作为替代评估工具。
原文摘要 · Abstract (English)
Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unsettled. Prevailing evaluations rely on human-annotated references, whose content reflects annotator intent and captioning proficiency. In this paper, we study a reconstruction-based principle for caption evaluation: a caption is as good as its capacity to enable reconstruction of the original image. However, because captioning inherently compresses visual information, it is impossible to recover all details, and pixel-wise comparison between reconstructed and source images is neither feasible nor meaningful. Through our in-depth analysis of the nature of captions, whose fundamental purpose is to transmit the semantic content of an image, we propose a revised principle: a caption is as good as its capacity to enable a reconstruction that is semantically equivalent to the original. To assess semantic equivalence, we test whether the reconstruction matches the original image across a suite of downstream vision--language tasks, yielding a reference-free, task-conditioned caption score. We characterize component-dependent limitations and introduce the lower-cost Captioning Turing Test Dataset (CTTD) surrogate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。