arXiv:2412.18150cs.CVcs.AI2024-12被引 44

构建4万对带细粒度标注的图文数据集,用于更精准评估生成模型。

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

  • 构建40K图文对,含细粒度人类标注,提升评估可靠性。
  • 提出FGA-BLIP2与PN-VQA新方法,实现零样本细粒度评估。
  • 可用于排名当前AIGC模型,助力未来研究发展。

近期,文本到图像(T2I)生成模型取得了显著进展,相应地涌现出大量自动化评估指标来衡量生成模型的图文对齐能力。然而,现有小规模数据集限制了这些指标间的性能比较,且缺乏细粒度评估能力。本研究构建了EvalMuse-40K基准,包含40,000个图文对,并配有细粒度的人类标注,用于图文对齐任务评估。在构建过程中,采用平衡提示采样和数据重标注等策略,确保数据多样性与可靠性。基于此,可全面评估T2I模型的图文对齐指标有效性。同时,我们提出两种新方法:FGA-BLIP2通过端到端微调视觉语言模型生成细粒度对齐分数;PN-VQA则采用新颖的正负样本VQA方式,在零样本条件下实现细粒度评估。两种方法均表现优异。我们还利用这些方法对当前AIGC模型进行排序,结果可为后续研究提供参考,推动T2I生成技术发展。数据与代码将公开共享。

原文摘要 · Abstract (English)

Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment capabilities of generative models. However, the performance comparison among these automated metrics is limited by existing small datasets. Additionally, these datasets lack the capacity to assess the performance of automated metrics at a fine-grained level. In this study, we contribute an EvalMuse-40K benchmark, gathering 40K image-text pairs with fine-grained human annotations for image-text alignment-related tasks. In the construction process, we employ various strategies such as balanced prompt sampling and data re-annotation to ensure the diversity and reliability of our benchmark. This allows us to comprehensively evaluate the effectiveness of image-text alignment metrics for T2I models. Meanwhile, we introduce two new methods to evaluate the image-text alignment capabilities of T2I models: FGA-BLIP2 which involves end-to-end fine-tuning of a vision-language model to produce fine-grained image-text alignment scores and PN-VQA which adopts a novel positive-negative VQA manner in VQA models for zero-shot fine-grained evaluation. Both methods achieve impressive performance in image-text alignment evaluations. We also use our methods to rank current AIGC models, in which the results can serve as a reference source for future study and promote the development of T2I generation. The data and code will be made publicly available.

图文对齐评估基准AIGC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。