评估文本生成图像模型的多样性,发现过度修正问题并提出更合理的解决方案。
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
- 构建了评估生成多样性的新基准DivBench,区分不足与过度多样化。
- 多数模型多样性不足,但现有方法常错误修改提示中明确指定的属性。
- 基于大模型引导的重写方法能更好平衡多样性与语义准确性,适合需要公平生成的场景。
当前文本到图像(T2I)模型的多样化策略常忽略上下文合理性,导致在提示中明确指定属性时仍出现过度多样化。本文提出DivBench,一个用于衡量T2I生成中多样性不足与过度多样化的评估基准与框架。对主流T2I模型的系统评估显示,多数模型存在多样性不足问题,而许多多样化方法则过矫正,不恰当地修改了上下文明确指定的属性。研究证明,上下文感知的方法,尤其是基于大语言模型(LLM)引导的FairDiffusion和提示重写技术,能够有效缓解多样性不足,同时避免过度多样化,在表现力与语义保真之间取得更好平衡。
原文摘要 · Abstract (English)
Current diversification strategies for text-to-image (T2I) models often ignore contextual appropriateness, leading to over-diversification where demographic attributes are modified even when explicitly specified in prompts. This paper introduces DIVBENCH, a benchmark and evaluation framework for measuring both under- and over-diversification in T2I generation. Through systematic evaluation of state-of-the-art T2I models, we find that while most models exhibit limited diversity, many diversification approaches overcorrect by inappropriately altering contextually-specified attributes. We demonstrate that context-aware methods, particularly LLM-guided FairDiffusion and prompt rewriting, can already effectively address under-diversity while avoiding over-diversification, achieving a better balance between representation and semantic fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。