arXiv:2509.21227cs.CVcs.CL2025-09中稿 · NeurIPS被引 3

评估文本生成图像的评测指标,发现没有哪种指标在所有情况下都可靠。

Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation

  • 对比多种评测指标在复杂组合任务中的表现差异
  • 发现不同指标在不同任务中优劣不一,无统一最优解
  • 提醒研究者谨慎选择指标,避免误导生成模型优化

文本到图像生成技术快速发展,但评估生成结果是否准确捕捉提示中的对象、属性和关系仍是核心挑战。当前评估高度依赖自动化指标,却常因惯例或流行而被采用,缺乏与人类判断的验证。由于评估结果直接影响领域进展,必须明确指标与人类偏好的一致性。为此,我们系统研究了广泛使用的组合性文本-图像评测指标。分析不仅考察相关性,还深入比较各类指标在多样组合挑战下的行为,以及与人类判断的一致性。结果表明:无单一指标在所有任务中表现稳定,性能随组合问题类型变化;尽管基于VQA的指标广受青睐,但并非始终更优;部分基于嵌入的指标在特定场景下表现更强;图像仅有的指标因侧重感知质量,对组合性评估贡献甚微。研究强调需审慎透明地选择指标,以确保评估可信,并作为生成过程中的奖励模型使用。

原文摘要 · Abstract (English)

Text-image generation has advanced rapidly, but assessing whether outputs truly capture the objects, attributes, and relations described in prompts remains a central challenge. Evaluation in this space relies heavily on automated metrics, yet these are often adopted by convention or popularity rather than validated against human judgment. Because evaluation and reported progress in the field depend directly on these metrics, it is critical to understand how well they reflect human preferences. To address this, we present a broad study of widely used metrics for compositional text-image evaluation. Our analysis goes beyond simple correlation, examining their behavior across diverse compositional challenges and comparing how different metric families align with human judgments. The results show that no single metric performs consistently across tasks: performance varies with the type of compositional problem. Notably, VQA-based metrics, though popular, are not uniformly superior, while certain embedding-based metrics prove stronger in specific cases. Image-only metrics, as expected, contribute little to compositional evaluation, as they are designed for perceptual quality rather than alignment. These findings underscore the importance of careful and transparent metric selection, both for trustworthy evaluation and for their use as reward models in generation. Project page is available at https://amirkasaei.com/eval-the-evals/ .

图像生成评测指标组合性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。