arXiv:2506.05108cs.CV2025-06ICCV被引 4

提出新评估框架,量化文本生成图像模型的多样性与泛化能力。

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models

  • 构建无参考的DIM-CIM框架,区分默认模式多样性和泛化能力。
  • 1.5B到8.1B参数模型提升泛化但牺牲默认多样性,相关性达0.85。
  • 可识别细粒度失败案例,适合模型优化与训练数据分析者使用。

近期文本到图像(T2I)模型在质量和一致性上取得显著进展,但以牺牲表示多样性为代价。现有自动评估方法要么依赖参考图像数据集,要么对多样性类型测量不具体,限制了可适应性和可解释性。为此,我们提出无参考的Does-it/Can-it框架——DIM-CIM,用于衡量默认模式多样性(模型是否生成预期属性)和泛化能力(能否为特定概念生成多样化属性)。我们构建了基于COCO概念和描述的COCO-DIMCIM基准,并通过大语言模型进行增强。在该基准上,我们发现从1.5B到8.1B参数规模的主流模型在提升泛化能力的同时,其默认模式多样性下降。此外,DIM-CIM能识别细粒度失败案例,如通用提示下可生成某些属性,但明确请求时却极少出现。最后,我们用DIM-CIM评估模型训练数据,发现训练图像多样性与默认模式多样性相关性为0.85。本工作提供了一个灵活且可解释的评估框架,推动对T2I模型性能的全面理解。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) models have achieved impressive quality and consistency. However, this has come at the cost of representation diversity. While automatic evaluation methods exist for benchmarking model diversity, they either require reference image datasets or lack specificity about the kind of diversity measured, limiting their adaptability and interpretability. To address this gap, we introduce the Does-it/Can-it framework, DIM-CIM, a reference-free measurement of default-mode diversity ("Does" the model generate images with expected attributes?) and generalization capacity ("Can" the model generate diverse attributes for a particular concept?). We construct the COCO-DIMCIM benchmark, which is seeded with COCO concepts and captions and augmented by a large language model. With COCO-DIMCIM, we find that widely-used models improve in generalization at the cost of default-mode diversity when scaling from 1.5B to 8.1B parameters. DIMCIM also identifies fine-grained failure cases, such as attributes that are generated with generic prompts but are rarely generated when explicitly requested. Finally, we use DIMCIM to evaluate the training data of a T2I model and observe a correlation of 0.85 between diversity in training images and default-mode diversity. Our work provides a flexible and interpretable framework for assessing T2I model diversity and generalization, enabling a more comprehensive understanding of model performance.

文本生成图像生成模型评估多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。