动态评估框架提升文本图像生成模型真实性和对齐度的评测能力
DynEval: Holistic Evaluations of T2I Generative Models in the Wild

- 构建动态评估框架,结合大规模生成数据与指令微调提升评测鲁棒性
- 在11个基准上与人工评分相关性更高,覆盖36个模型42个子类别的细粒度分析
- 适合需要全面评估生成模型性能的研究者和开发者使用
文本到图像(T2I)生成技术虽已能生成高度逼真的图像,但可靠评估仍具挑战性,尤其在大规模场景下。现有自动评估方法多依赖静态提示集,难以捕捉局部提示错位、组合错误或语义错误但视觉合理等细微缺陷。本文提出DynEval动态评估框架,用于联合评估T2I模型的文本-图像对齐与图像质量。为支持大规模训练,我们构建了两个大数据集:首先,基于DiffusionDB中人类撰写的提示,采用分层提示-模型生成策略构建包含50万组提示-图像对的GenDB;其次,在GenDB基础上,通过结构化评估流程提炼出25万条提示-图像-反馈三元组,形成DynEvalInstruct指令数据集。基于此数据,采用课程学习策略对小型评估器进行全量微调,以蒸馏大型教师模型的评估能力,最终得到DynEval-2B和DynEval-4B。在11个基准上的广泛对比显示,该评估器与人工评分的相关性更高。同时,可对36个T2I模型在42个子类别和9个语义维度上的能力与失败模式进行细粒度分析。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, especially at scale. Existing automatic evaluators, often relying on a static prompt set, struggle to capture subtle failure modes such as partial prompt misalignment, compositional errors, or visually plausible but semantically incorrect generations. In this work, we introduce DynEval, a Dynamic Evaluation framework designed to jointly assess text-to-image alignment and image quality of T2I models. To support scalable training beyond limited human-annotated data, we construct two large datasets. First, we build GenDB, a collection of 500K prompt-image pairs generated from human-written prompts drawn from DiffusionDB using a tiered prompt-model generation strategy. Second, building upon GenDB, we construct DynEvalInstruct, a 250K instruction dataset comprising prompt-image-response triplets distilled from a structured evaluation pipeline that decomposes evaluation into text-image alignment and visual quality reasoning. Using this dataset, we perform full fine-tuning of a compact evaluator through a curriculum learning strategy to effectively distill the superior evaluation capabilities of a larger teacher vision-language model, resulting in DynEval-2B and DynEval-4B. In extensive comparisons against existing evaluators across 11 benchmarks, our evaluator achieves a higher overall correlation with human judgments. Furthermore, it provides fine-grained analysis of the capabilities and failure modes of 36 T2I models across 42 subcategories and 9 semantic dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。