arXiv:2605.13974cs.CVcs.AI2026-05被引 1

发现扩散模型中少数通道主导图像生成,可精准操控语义内容。

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

论文配图:Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
图 1 · 摘自论文原文
  • 识别出极少数高激活通道,它们是生成质量的关键
  • 仅保留这些通道就能还原图像主体与显著区域
  • 可跨提示迁移激活,实现无需训练的语义插值

扩散变换器(DiTs)等基于流的架构是当前最强的文生图模型,但提示如何塑造图像语义的内部机制仍不清晰。本文研究了‘大规模激活’:一个极小比例的隐藏状态通道,其响应远高于其他通道。尽管数量稀少,这些通道在三个层面有效构建了整幅图像。第一,功能关键:人为置零这些通道导致生成质量急剧下降,而同等数量的低激活通道被破坏则影响微乎其微。第二,空间有序:仅保留这些通道并聚类,能生成与主体和显著区域高度一致的结构化分割图,揭示出看似异常的子空间中隐藏的结构化空间编码。第三,可迁移:将一个提示下的大规模激活迁移到另一个提示中,能使最终图像向源提示偏移,同时保留目标内容,实现局部语义插值而非无序像素混合。我们将其用于文本和图像条件下的语义传输,实现无需额外训练的提示插值与主体驱动生成。这些结果表明,大规模激活并非异常,而是现代DiT模型中组织和控制语义信息的稀疏提示条件载体子空间。

原文摘要 · Abstract (English)

Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood. In this work, we study massive activations: a small subset of hidden-state channels whose responses are consistently much larger than the rest. We show that, despite their sparsity, these few channels effectively draw the whole picture, in three complementary senses. First, they are functionally critical: a controlled disruption probe that zeroes the massive channels causes a sharp collapse in generation quality, while disrupting an equally-sized set of low-statistic channels has marginal effect. Second, they are spatially organized: restricting image-stream tokens to massive channels and clustering them yields coherent partitions that closely align with the main subject and salient regions, exposing a structured spatial code hidden inside an apparently outlier-like subspace. Third, they are transferable: transporting massive activations from one prompt-conditioned trajectory into another, shifts the final image toward the source prompt while preserving substantial content from the target, producing localized semantic interpolation rather than unstructured pixel blending. We exploit this property in two use cases: text-conditioned and image-conditioned semantic transport, where massive activations transport enables prompt interpolation and subject-driven generation without any additional training. Together, these results recast massive activations not as activation anomalies, but as a sparse prompt-conditioned carrier subspace that organizes and controls semantic information in modern DiT models.

扩散模型激活分析语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。