arXiv:2410.22592cs.CV2024-10被引 6

用语义驱动方法量化文生图模型的样本多样性,发现模型生成结果高度同质。

GRADE: Quantifying Sample Diversity in Text-to-Image Models

  • 利用大语言模型和视觉问答系统识别概念特异性差异轴(如饼干的形状)
  • 12个模型生成72万张图像中,98%饼干为圆形,显示显著默认行为
  • 指出训练数据描述不明确是导致多样性低的关键原因

我们提出GRADE,一种自动量化文生图模型样本多样性的方法。该方法利用大语言模型和视觉问答系统的世界知识,识别与特定概念相关的多样性维度(如‘饼干’的概念对应‘形状’)。通过估计概念及其属性的频率分布并以熵值衡量多样性,我们在总计72万张图像上评估了12个模型的表现,发现所有模型均存在明显多样性不足,且更强模型的多样性反而下降。进一步研究发现,模型常表现出‘默认行为’——即稳定生成相同属性的概念(如98%的饼干为圆形)。最后,我们揭示训练数据中描述不明确是导致多样性低的主要原因。本工作提出了一种自动、语义驱动的多样性度量方法,揭示了文生图模型输出的高度同质性。

原文摘要 · Abstract (English)

We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and visual question-answering systems to identify relevant concept-specific axes of diversity (e.g., ``shape'' for the concept ``cookie''). It then estimates frequency distributions of concepts and their attributes and quantifies diversity using entropy. We use GRADE to measure the diversity of 12 models over a total of 720K images, revealing that all models display limited variation, with clear deterioration in stronger models. Further, we find that models often exhibit default behaviors, a phenomenon where a model consistently generates concepts with the same attributes (e.g., 98% of the cookies are round). Lastly, we show that a key reason for low diversity is underspecified captions in training data. Our work proposes an automatic, semantically-driven approach to measure sample diversity and highlights the stunning homogeneity in text-to-image models.

文生图多样性评估生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。