arXiv:2505.11314cs.CVcs.CL2025-05Transactions of th…被引 1

提出CROC框架,用百万级伪标签数据自动评估文本到图像生成指标鲁棒性。

CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks

  • 通过合成对比测试对图像属性进行系统性探测
  • 构建超百万对伪标签数据集,发现多数指标在否定句上失效
  • 适合关注生成质量评估的科研人员与模型开发者

文本到图像生成任务中,评估指标的可靠性(元评估)至关重要。人工元评估成本高、耗时长,自动化替代方案稀缺。本文提出可扩展的CROC框架,通过合成跨图像属性的对比测试用例,系统性地探测并量化指标鲁棒性。利用CROC构建了超过100万组伪标签对比提示-图像对(CROC$^{syn}$),实现对评估指标的细粒度比较。同时基于该数据集训练出新指标CROCScore,其在开源方法中表现领先。为补充数据,还引入人类标注基准(CROC$^{hum}$),聚焦高难度类别。结果揭示现有指标存在严重鲁棒性问题:许多指标在含否定词的提示下失效,所有测试的开源指标均至少在24%涉及身体部位识别的案例中失败。

原文摘要 · Abstract (English)

The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) generation tasks. Human-based meta-evaluation is costly and time-intensive, and automated alternatives are scarce. We address this gap and propose CROC: a scalable framework for automated Contrastive Robustness Checks that systematically probes and quantifies metric robustness by synthesizing contrastive test cases across a comprehensive taxonomy of image properties. With CROC, we generate a pseudo-labeled dataset (CROC$^{syn}$) of over 1 million contrastive prompt-image pairs to enable a fine-grained comparison of evaluation metrics. We also use this dataset to train CROCScore, a new metric that achieves state-of-the-art performance among open-source methods, demonstrating an additional key application of our framework. To complement this dataset, we introduce a human-supervised benchmark (CROC$^{hum}$) targeting especially challenging categories. Our results highlight robustness issues in existing metrics: for example, many fail on prompts involving negation, and all tested open-source metrics fail on at least 24% of cases involving correct identification of body parts.

文本生成图像评估指标鲁棒性自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。