arXiv:2502.14914cs.CVcs.CL2025-02NeurIPS被引 18

新基准CAPability全面评估图文生成的准确与完整度。

CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

  • 构建12维多视角评测体系,覆盖11,000+张图视频数据
  • 引入精确率与命中率指标,稳定衡量生成质量
  • 发现模型懂但说不清,暴露语言表达短板

视觉描述基准已因多模态大模型兴起而过时,传统简短标准答案和评价指标难以有效评估详细描述。尽管近期基准尝试通过关键词提取或物体中心评价改进,仍局限于模糊视图或物体视图分析,且视觉元素覆盖不全。本文提出CAPability,一个涵盖12个维度、六种关键视角的综合性多视图评测基准。我们收集了近11,000张人工标注的图像与视频,并附有视觉元素标注,用于评估生成描述。CAPability通过精确率(precision)与命中率(hit)指标,稳定评估描述的正确性与完整性。通过将标注转化为问答对,进一步提出启发式指标“知但不能说”(K̄T),揭示了问答能力与描述能力间显著性能差距。本研究对多模态大模型的描述能力进行全方位分析,识别其在各维度的优势与不足,为未来提升特定能力提供方向。

原文摘要 · Abstract (English)

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions with \textit{precision} and \textit{hit} metrics. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides a holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of their capabilities.

视觉描述评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。