通过动态干扰文本-图像关联空间,实现生成多样性提升而不损失质量。
On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers
- 在Transformer前向传播中实时干预多模态注意力通道
- 显著提升生成多样性,同时保持视觉质量和语义一致性
- 计算开销小,适用于快速模型如Turbo和蒸馏版
当前文生图扩散模型虽具优异语义对齐能力,但生成结果常趋于单一,缺乏多样性。现有方法或需代价高昂的优化反馈,或在中间潜在空间干预时破坏结构导致伪影。本文提出在上下文空间中引入动态排斥机制,于Transformer前向过程中,在文本条件与图像结构融合后的块间注入干预,引导生成路径在结构形成后、组合定型前发生偏移。实验表明,该方法显著增强多样性,同时维持高视觉保真度与语义准确性;且计算开销极低,对现代轻量级(如Turbo)及蒸馏模型仍有效,而传统轨迹干预在此类模型中通常失效。
原文摘要 · Abstract (English)
Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt. This typicality bias presents a challenge for creative applications that require a wide range of generative outcomes. We identify a fundamental trade-off in current approaches to diversity: modifying model inputs requires costly optimization to incorporate feedback from the generative path. In contrast, acting on spatially-committed intermediate latents tends to disrupt the forming visual structure, leading to artifacts. In this work, we propose to apply repulsion in the Contextual Space as a novel framework for achieving rich diversity in Diffusion Transformers. By intervening in the multimodal attention channels, we apply on-the-fly repulsion during the transformer's forward pass, injecting the intervention between blocks where text conditioning is enriched with emergent image structure. This allows for redirecting the guidance trajectory after it is structurally informed but before the composition is fixed. Our results demonstrate that repulsion in the Contextual Space produces significantly richer diversity without sacrificing visual fidelity or semantic adherence. Furthermore, our method is uniquely efficient, imposing a small computational overhead while remaining effective even in modern "Turbo" and distilled models where traditional trajectory-based interventions typically fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。