arXiv:2602.21596cs.CV2026-02中稿 · ICLR被引 2

发现扩散模型条件嵌入存在语义瓶颈,可大幅压缩而不影响生成质量。

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

  • 系统分析扩散Transformer的条件嵌入结构,发现其角度高度相似。
  • 99%以上类条件嵌入方向重合,关键语义仅集中在少数维度。
  • 剪枝低幅值维度可删减三分之二空间,生成效果反而提升。

扩散Transformer在类别条件和多模态生成上达到顶尖性能,但其学习到的条件嵌入结构仍不清晰。本文首次系统研究这些嵌入,发现显著冗余:类条件嵌入在ImageNet-1K上角度相似度超过99%,而姿态引导图像生成、视频转音频等连续条件任务相似度超99.9%。我们进一步发现,语义信息集中于少数维度,头部维度承载主要信号,尾部维度贡献极小。通过剪枝低幅值维度——最多删除嵌入空间的三分之二——生成质量和保真度基本不受影响,某些情况下甚至提升。结果揭示了基于Transformer的扩散模型中存在语义瓶颈,为语义编码机制提供了新理解,并提示更高效的条件化路径。

原文摘要 · Abstract (English)

Diffusion Transformers have achieved state-of-the-art performance in class-conditional and multimodal generation, yet the structure of their learned conditional embeddings remains poorly understood. In this work, we present the first systematic study of these embeddings and uncover a notable redundancy: class-conditioned embeddings exhibit extreme angular similarity, exceeding 99\% on ImageNet-1K, while continuous-condition tasks such as pose-guided image generation and video-to-audio generation reach over 99.9\%. We further find that semantic information is concentrated in a small subset of dimensions, with head dimensions carrying the dominant signal and tail dimensions contributing minimally. By pruning low-magnitude dimensions--removing up to two-thirds of the embedding space--we show that generation quality and fidelity remain largely unaffected, and in some cases improve. These results reveal a semantic bottleneck in Transformer-based diffusion models, providing new insights into how semantics are encoded and suggesting opportunities for more efficient conditioning mechanisms.

扩散模型条件生成嵌入压缩语义瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。