arXiv:2506.24016cs.CLcs.AI2025-06ACL被引 7

提出可解释的图像描述评估指标,生成结构化解释。

EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations

  • 基于流畅性、相关性、描述性三标准生成结构化解释。
  • 在基准数据集上达当前最优,人类评估显示解释质量显著提升。
  • 适合需要透明评估结果的研究者与模型开发者。

大型语言模型和视觉-语言模型的发展推动了图像描述可解释评估指标的研究。然而,现有指标生成的解释缺乏统一标准,整体质量难以验证。本文提出EXPERT,一种无需参考文本的评估指标,基于流畅性、相关性和描述性三个基本准则生成结构化解释。通过构建大规模高质量结构化解释数据集,设计两阶段评估模板,有效监督视觉-语言模型完成评分与解释生成。EXPERT在基准数据集上达到领先性能,经全面人工评估验证,其生成解释的质量显著优于现有方法。代码与数据集已公开于https://github.com/hjkim811/EXPERT。

原文摘要 · Abstract (English)

Recent advances in large language models and vision-language models have led to growing interest in explainable evaluation metrics for image captioning. However, these metrics generate explanations without standardized criteria, and the overall quality of the generated explanations remains unverified. In this paper, we propose EXPERT, a reference-free evaluation metric that provides structured explanations based on three fundamental criteria: fluency, relevance, and descriptiveness. By constructing large-scale datasets of high-quality structured explanations, we develop a two-stage evaluation template to effectively supervise a vision-language model for both scoring and explanation generation. EXPERT achieves state-of-the-art results on benchmark datasets while providing significantly higher-quality explanations than existing metrics, as validated through comprehensive human evaluation. Our code and datasets are available at https://github.com/hjkim811/EXPERT.

可解释性图像描述评估指标视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。