arXiv:2602.01609cs.CV2026-02被引 2

提出无需训练的令牌剪枝方法,加速扩散模型上下文生成

Token Pruning for In-Context Generation in Diffusion Transformers

  • 基于离线校准的敏感性分析定位关键注意力层,评估上下文令牌冗余度
  • 设计新影响度量指标与动态更新策略,实现跨时空维度的精准剪枝
  • 在不损失图像结构与视觉一致性前提下,推理速度提升超30%

上下文生成通过参考样例显著增强扩散Transformer(DiTs)的可控图像到图像生成能力。然而,输入拼接导致序列长度激增,带来严重计算瓶颈。现有令牌压缩技术多针对文本到图像合成,采用统一剪枝策略,忽视了参考上下文与目标潜在表示在空间、时间及功能维度上的固有角色差异。为此,我们提出ToPi——一种专为DiTs中上下文生成设计的无训练令牌剪枝框架。具体而言,ToPi利用离线校准驱动的敏感性分析识别关键注意力层,作为冗余度估计的可靠代理;基于这些层,构建新颖的影响度量以量化各上下文令牌贡献,并结合随扩散轨迹演化的时序更新策略进行选择性剪枝。实证评估表明,ToPi可在复杂图像生成任务中实现超过30%的推理加速,同时保持结构保真度与视觉一致性。

原文摘要 · Abstract (English)

In-context generation significantly enhances Diffusion Transformers (DiTs) by enabling controllable image-to-image generation through reference examples. However, the resulting input concatenation drastically increases sequence length, creating a substantial computational bottleneck. Existing token reduction techniques, primarily tailored for text-to-image synthesis, fall short in this paradigm as they apply uniform reduction strategies, overlooking the inherent role asymmetry between reference contexts and target latents across spatial, temporal, and functional dimensions. To bridge this gap, we introduce ToPi, a training-free token pruning framework tailored for in-context generation in DiTs. Specifically, ToPi utilizes offline calibration-driven sensitivity analysis to identify pivotal attention layers, serving as a robust proxy for redundancy estimation. Leveraging these layers, we derive a novel influence metric to quantify the contribution of each context token for selective pruning, coupled with a temporal update strategy that adapts to the evolving diffusion trajectory. Empirical evaluations demonstrate that ToPi can achieve over 30\% speedup in inference while maintaining structural fidelity and visual consistency across complex image generation tasks.

扩散模型令牌剪枝推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。