arXiv:2411.17522stat.MLcs.AI2024-11被引 26

揭示条件扩散Transformer的统计最优性,为模型设计提供理论依据

On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality

  • 通过泰勒展开与分段常数逼近,精细分析条件扩散得分函数
  • 证明条件与潜在空间的DiT在特定条件下达到统计最优,误差率更优
  • 适合研究扩散模型理论、优化生成效率的研究者参考

我们研究了带无分类器引导的条件扩散Transformer(DiTs)的近似与估计速率。在四种常见数据假设下,对“上下文内”条件DiTs进行了全面分析。结果显示,在特定设定下,条件DiTs及其潜在变体均能达到无条件DiTs的极小极大最优性。具体而言,将输入域离散化为无穷小网格,并在霍尔德光滑数据假设下对条件扩散得分函数进行逐项泰勒展开,从而更精细地利用Transformer的通用逼近能力,实现更紧的误差界。此外,在线性潜在子空间假设下扩展至潜在空间设置,不仅表明潜在条件DiTs在近似与估计上均优于普通条件DiTs,还证明了潜在无条件DiTs的极小极大最优性。研究成果确立了条件与无条件DiTs的统计极限,为构建更高效、准确的DiT模型提供了实践指导。

原文摘要 · Abstract (English)

We investigate the approximation and estimation rates of conditional diffusion transformers (DiTs) with classifier-free guidance. We present a comprehensive analysis for ``in-context'' conditional DiTs under four common data assumptions. We show that both conditional DiTs and their latent variants lead to the minimax optimality of unconditional DiTs under identified settings. Specifically, we discretize the input domains into infinitesimal grids and then perform a term-by-term Taylor expansion on the conditional diffusion score function under Hölder smooth data assumption. This enables fine-grained use of transformers' universal approximation through a more detailed piecewise constant approximation and hence obtains tighter bounds. Additionally, we extend our analysis to the latent setting under the linear latent subspace assumption. We not only show that latent conditional DiTs achieve lower bounds than conditional DiTs both in approximation and estimation, but also show the minimax optimality of latent unconditional DiTs. Our findings establish statistical limits for conditional and unconditional DiTs, and offer practical guidance toward developing more efficient and accurate DiT models.

扩散模型理论分析最优性Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。