指出当前图文生成评估方法的缺陷并提出改进建议。
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
- 识别可信评估需满足的两个关键特性。
- 实证发现主流评估框架在多模型多指标下表现不一致。
- 提出改进图文对齐评估的实用建议,适合评估研究者参考。
文本到图像模型常难以生成与文本提示精确匹配的图像。先前研究广泛探讨了图文对齐的评估问题,但现有评估主要关注与人类判断的一致性,忽略了可信评估框架的其他关键属性。本文首先识别出可靠评估应满足的两个核心方面,随后通过实验证明当前主流评估框架在多种度量标准和模型上未能充分满足这些特性。最后,我们提出了改进图文对齐评估的若干建议。
原文摘要 · Abstract (English)
Text-to-image models often struggle to generate images that precisely match textual prompts. Prior research has extensively studied the evaluation of image-text alignment in text-to-image generation. However, existing evaluations primarily focus on agreement with human assessments, neglecting other critical properties of a trustworthy evaluation framework. In this work, we first identify two key aspects that a reliable evaluation should address. We then empirically demonstrate that current mainstream evaluation frameworks fail to fully satisfy these properties across a diverse range of metrics and models. Finally, we propose recommendations for improving image-text alignment evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。