构建大规模图文对齐评测基准,揭示现模型在多实例场景下的生成缺陷
M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark
- 提出多类别、多实例、多关系的文本到图像评测基准
- 现有开源模型在该基准上表现不佳,平均得分低于60分
- 提出无需训练的后处理方法,显著提升多类扩散模型对齐效果
文本到图像模型常难以精准匹配文本提示。以往评估多聚焦于简单场景,忽视同一类别中多个实例的生成挑战,或使用与人类评价相关性差的指标。本文提出M$^3$T2IBench,一个大规模、多类别、多实例、多关系的文本到图像评测基准,并引入基于目标检测的评估指标AlignScore,其与人类评价高度一致。实验发现,当前开源文本到图像模型在此基准上表现较差。为此,我们提出Revise-Then-Enforce方法,一种无需训练的后处理策略,在多种扩散模型上均显著提升图像与文本的对齐能力。
原文摘要 · Abstract (English)
Text-to-image models are known to struggle with generating images that perfectly align with textual prompts. Several previous studies have focused on evaluating image-text alignment in text-to-image generation. However, these evaluations either address overly simple scenarios, especially overlooking the difficulty of prompts with multiple different instances belonging to the same category, or they introduce metrics that do not correlate well with human evaluation. In this study, we introduce M$^3$T2IBench, a large-scale, multi-category, multi-instance, multi-relation along with an object-detection-based evaluation metric, $AlignScore$, which aligns closely with human evaluation. Our findings reveal that current open-source text-to-image models perform poorly on this challenging benchmark. Additionally, we propose the Revise-Then-Enforce approach to enhance image-text alignment. This training-free post-editing method demonstrates improvements in image-text alignment across a broad range of diffusion models. \footnote{Our code and data has been released in supplementary material and will be made publicly available after the paper is accepted.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。