arXiv:2412.13989cs.CL2024-12中稿 · COLM被引 14

评测四大图文一致性评估方法,发现它们都有明显缺陷。

What makes a good metric? Evaluating automatic metrics for text-to-image consistency

  • 分析四种基于语言模型的图文一致性评估方法
  • 无一满足所有理想评估标准,且对语言视觉特性不敏感
  • 基于VQA的方法可能依赖常见文本捷径,可靠性存疑

语言模型越来越多地被用作大型AI系统中的组件,用于提示优化到自动评估等任务。本文分析了四种近期常用、用于衡量图文一致性的方法——CLIPScore、TIFA、VPEval和DSG——其构建有效性。我们定义了图文一致性度量应具备的理想特性集合,并发现所测试的任何一种度量均未满足全部条件。研究发现这些度量对语言与视觉属性的敏感性不足;此外,TIFA、VPEval和DSG虽提供超越CLIPScore的新信息,但彼此间相关性很高;通过消融实验发现并非所有模型组件都不可或缺,进一步表明对视觉信息的敏感性不足。最后,我们证明三种基于VQA的度量可能依赖于熟悉的文本捷径(如问答中的‘是’偏见),这使其作为模型性能量化评估工具的有效性受到质疑。

原文摘要 · Abstract (English)

Language models are increasingly being incorporated as components in larger AI systems for various purposes, from prompt optimization to automatic evaluation. In this work, we analyze the construct validity of four recent, commonly used methods for measuring text-to-image consistency - CLIPScore, TIFA, VPEval, and DSG - which rely on language models and/or VQA models as components. We define construct validity for text-image consistency metrics as a set of desiderata that text-image consistency metrics should have, and find that no tested metric satisfies all of them. We find that metrics lack sufficient sensitivity to language and visual properties. Next, we find that TIFA, VPEval and DSG contribute novel information above and beyond CLIPScore, but also that they correlate highly with each other. We also ablate different aspects of the text-image consistency metrics and find that not all model components are strictly necessary, also a symptom of insufficient sensitivity to visual information. Finally, we show that all three VQA-based metrics likely rely on familiar text shortcuts (such as yes-bias in QA) that call their aptitude as quantitative evaluations of model performance into question.

图文评估语言模型测评基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。