arXiv:2512.20666cs.LGcs.AI2025-12被引 1

发现扩散模型生成多概念图像时存在主导与被压制现象

Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models

  • 通过控制微调实验揭示视觉同质数据导致概念主导
  • 早期去噪阶段主导词集中注意力,抑制其他概念表征
  • 注意力头分布性失效,适合研究可控多概念生成的团队

文本到图像的扩散模型在生成多样化高保真图像方面表现优异,但在多概念生成中常出现某一概念词主导输出而其他概念被压制的现象,我们称之为主导-被主导(DvD)失衡。为系统研究此故障模式,我们引入DominanceBench,并从数据和内部机制两个角度分析其成因。受控微调实验显示,基于视觉同质(低变异性)特定概念训练图像学习的概念,在与其他概念组合时表现出更强的主导性。交叉注意力分析表明,主导词在早期去噪步骤集中注意力,随后竞争概念的表征减弱。头消融分析进一步显示,这种主导性分布在多个注意力头中而非局部集中。总体而言,这些发现将DvD定义为系统性的概念级失败模式,为更可靠、可控的多概念生成提供基础。DominanceBench将在发表后公开。

原文摘要 · Abstract (English)

Text-to-image diffusion models have attracted significant attention for their ability to generate diverse, high-fidelity images. However, in multi-concept generation, one concept token often dominates the output while others are suppressed-a phenomenon we term the Dominant-vs-Dominated (DvD) imbalance. To systematically study this failure mode, we introduce DominanceBench and examine its underlying causes from both data and internal-mechanistic perspectives. Our controlled fine-tuning study, which mimics concept learning during diffusion-model training, shows that concepts learned from visually homogeneous (low-variation) concept-specific training images exhibit stronger dominance when composed with others. Cross-attention analysis indicates that dominant tokens concentrate attention in early denoising steps, followed by reduced representation of competing concepts. Head-ablation analysis further shows that this dominance is distributed across attention heads rather than localized. Overall, these findings characterize DvD as a systematic concept-level failure mode and provide a basis for more reliable and controllable multi-concept generation. DominanceBench will be released upon publication.

扩散模型多概念生成注意力机制生成失衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。