arXiv:2606.06813cs.CVcs.AI2026-06中稿 · ICML

通过抑制早期特征中的直流分量,无需训练即可提升文本到图像生成的多样性。

Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation

论文配图:Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation
图 1 · 摘自论文原文
  • 在生成早期抑制特征的直流分量,打破采样轨迹锁定
  • 相比原模型,多样性提升显著且图像质量保持不变
  • 无需额外训练或采样,适合快速部署于现有生成系统

近期基于大规模Transformer和流模型的目标函数构建的文本到图像生成模型虽能实现良好的图文对齐与高质量视觉输出,但在固定提示下常产生高度相似的样本。现有增强多样性的方法通常需昂贵的采样或辅助优化,带来显著开销。我们分析中间Transformer特征发现,零频空间平均(DC)分量在生成初期即快速收敛,导致早期轨迹锁定,限制后续变化。为此,提出无需训练的表示层干预方法DAVE,通过选择性衰减早期阶段的该分量来提升多样性。DAVE不改变原有采样流程,开销极小,在保持竞争性图像质量的同时显著增强提示一致下的多样性。

原文摘要 · Abstract (English)

Recent text-to-image models built on large-scale Transformer backbones and flow-based objectives deliver strong text-image alignment and high visual quality, yet often produce overly similar samples under a fixed prompt. Existing diversity-enhancement methods alleviate this issue, but typically require expensive sampling or auxiliary optimization, incurring non-trivial overhead. To investigate the root cause of this homogeneity, we examine intermediate Transformer features and observe that the zero-frequency spatial average (DC) component rapidly converges across seeds early in generation, causing early trajectory lock-in that limits downstream variation. Building on this observation, we propose DC Attenuation for diVersity Enhancement (DAVE), a training-free representation-level intervention that selectively attenuates this component in the early regime. DAVE preserves the sampling pipeline with negligible overhead, improving prompt-consistent diversity while maintaining competitive image quality.

文本生成图像多样性提升表示调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。