arXiv:2603.25758cs.CVcs.AI2026-03

自动选最优时间步,让扩散Transformer更高效学表征。

A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation Learning

  • 动态从Transformer特征中选出信息最丰富的时刻
  • 在分类与分割任务上超越所有先前扩散模型方法
  • 适合做生成预训练的下游表征学习研究者

扩散模型已深刻改变生成式人工智能领域,现被越来越多用于判别性表征学习。扩散Transformer(DiT)作为传统U-Net型扩散模型的有前途替代方案,通过生成预训练展现了在下游判别任务中的潜力。然而,其当前训练效率和表征能力仍受限于不足的时间步搜索以及对DiT特有特征表示的利用不充分。为此,本文提出自动时间步选择(A-SelecT),可在单次运行中动态定位DiT最具信息量的时间步,无需耗时的穷举搜索和次优的判别特征选择。在分类与分割基准上的大量实验表明,经A-SelecT增强的DiT在效率与效果上均显著超越此前所有基于扩散模型的方法。

原文摘要 · Abstract (English)

Diffusion models have significantly reshaped the field of generative artificial intelligence and are now increasingly explored for their capacity in discriminative representation learning. Diffusion Transformer (DiT) has recently gained attention as a promising alternative to conventional U-Net-based diffusion models, demonstrating a promising avenue for downstream discriminative tasks via generative pre-training. However, its current training efficiency and representational capacity remain largely constrained due to the inadequate timestep searching and insufficient exploitation of DiT-specific feature representations. In light of this view, we introduce Automatically Selected Timestep (A-SelecT) that dynamically pinpoints DiT's most information-rich timestep from the selected transformer feature in a single run, eliminating the need for both computationally intensive exhaustive timestep searching and suboptimal discriminative feature selection. Extensive experiments on classification and segmentation benchmarks demonstrate that DiT, empowered by A-SelecT, surpasses all prior diffusion-based attempts efficiently and effectively.

扩散模型Transformer表征学习时间步选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。