arXiv:2506.07631cs.CLcs.CV2025-06被引 2

为详细图像描述的精准评估提供新基准与自动化工具

Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline

  • 构建包含10216条细粒度标注的DOCCI-Critique数据集
  • 自动评析模型在多个基准上达到顶尖表现,相关性达0.98
  • 提出批判-修正流程,使描述准确率提升46%

大型视觉语言模型(VLM)现在能生成高度详细的段落级图像描述,但评估其事实准确性仍具挑战。现有方法常遗漏细粒度错误,因设计用于较短文本或缺乏经验证的错误数据集。我们引入DOCCI-Critique基准,包含1,400条由14个VLM生成的段落级描述(100张图像),涵盖超过10,216条句子级别的事实正确性人工标注及错误解释,全部置于段落上下文中。基于此,我们开发了VNLI-Critique模型,实现自动化句子级事实判断与批判生成。主要应用包括:(1) VNLI-Critique在M-HalDetect基准上表现优异,且在CHOCOLATE主张验证中取得强结果;(2) 基于VNLI-Critique的AutoRater在DOCCI-Critique上提供可靠排名,与人类判断高度一致(如0.98斯皮尔曼相关性);(3) 创新性地采用批判-修正流水线,利用VNLI-Critique的批判指导大模型修正,显著提升描述事实性(如DetailCaps-4870上提升46%)。本工作提供关键基准与实用工具,旨在大幅提升细粒度评估标准,推动VLM图像理解能力进步。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) now generate highly detailed, paragraphlength image captions, yet evaluating their factual accuracy remains challenging. Current methods often miss fine-grained errors, being designed for shorter texts or lacking datasets with verified inaccuracies. We introduce DOCCI-Critique, a benchmark with 1,400 VLM-generated paragraph captions (100 images, 14 VLMs) featuring over 10,216 sentence-level human annotations of factual correctness and explanatory rationales for errors, all within paragraph context. Building on this, we develop VNLI-Critique, a model for automated sentence-level factuality classification and critique generation. We highlight three key applications: (1) VNLI-Critique demonstrates robust generalization, validated by state-of-the-art performance on the M-HalDetect benchmark and strong results in CHOCOLATE claim verification. (2) The VNLI-Critique driven AutoRater for DOCCI-Critique provides reliable VLM rankings, showing excellent alignment with human factuality judgments (e.g., 0.98 Spearman). (3) An innovative Critic-and-Revise pipeline, where critiques from VNLI-Critique guide LLM-based corrections, achieves substantial improvements in caption factuality (e.g., a 46% gain on DetailCaps-4870). Our work offers a crucial benchmark alongside practical tools, designed to significantly elevate the standards for fine-grained evaluation and foster the improvement of VLM image understanding. Project page: https://google.github.io/unblocking-detail-caption

图像描述事实评估自动评析批判修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。