arXiv:2511.10547cs.CVcs.LG2025-11被引 4

用人类评估法系统测试文生图模型的多样性表现

Benchmarking Diversity in Image Generation via Attribute-Conditional Human Evaluation

  • 设计可细化评估概念与变化因素的人类评测模板
  • 通过二项式检验对比模型在多样性的实际表现差异
  • 发现模型在颜色、姿态等特定因素上生成重复性强

尽管文生图(T2I)模型生成质量持续提升,但多样性不足问题依然存在,常生成同质化输出。本文提出一种系统性评估框架,用于衡量T2I模型的多样性。核心贡献包括:(1) 设计新颖的人类评估模板,实现对概念及其变化因素的精细评估;(2) 构建包含多种概念及对应变化因素的提示集(如:苹果,颜色);(3) 采用二项式检验方法,基于人工标注比较不同模型的多样性表现。此外,本文还严格对比多种图像嵌入方式在多样性度量中的有效性。所提方法可对模型按多样性进行排序,并揭示其在特定类别上的薄弱环节。该研究为提升文生图模型多样性与评价指标发展提供了可靠方法与洞见。

原文摘要 · Abstract (English)

Despite advances in generation quality, current text-to-image (T2I) models often lack diversity, generating homogeneous outputs. This work introduces a framework to address the need for robust diversity evaluation in T2I models. Our framework systematically assesses diversity by evaluating individual concepts and their relevant factors of variation. Key contributions include: (1) a novel human evaluation template for nuanced diversity assessment; (2) a curated prompt set covering diverse concepts with their identified factors of variation (e.g. prompt: An image of an apple, factor of variation: color); and (3) a methodology for comparing models in terms of human annotations via binomial tests. Furthermore, we rigorously compare various image embeddings for diversity measurement. Notably, our principled approach enables ranking of T2I models by diversity, identifying categories where they particularly struggle. This research offers a robust methodology and insights, paving the way for improvements in T2I model diversity and metric development.

文生图多样性评估人类评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。