arXiv:2605.27102cs.CVcs.LG2026-05被引 1

在潜在空间中直接预测干净图像,比预测速度更有效。

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

论文配图:JLT: Clean-Latent Prediction in Latent Diffusion Transformers
图 1 · 摘自论文原文
  • 直接预测潜在空间中的干净图像,而非噪声速度。
  • 在ImageNet上达到FID 2.50,显著优于速度预测方法。
  • 适合研究扩散模型优化与潜在表示的从业者。

基于干净数据预测的流匹配方法表明,回归干净样本能更有效地利用低维结构,而非预测高维噪声量。我们探究这一原理在图像映射到学习的潜在空间后是否依然有效,此时压缩已消除大量原始像素变化。本文提出JLT,一个基于冻结FLUX.2 VAE编码的130M潜在扩散Transformer,与匹配的速率预测DiT在相同表示、主干和训练设置下对比。尽管在固定污染时间下,x、ε、v三者线性可转换,但局部高斯分析显示,速率回归继承各向同性的目标协方差下界,并放大低方差潜在方向,而干净预测则抑制它们。在ImageNet 256×256上,JLT-B/1使用无分类器引导获得FID-50K 2.50,与速率预测存在显著差距。结果表明,潜在扩散中的预测目标是依赖表示的几何选择,而非可互换的代数参数化。

原文摘要 · Abstract (English)

Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predicting an ambient noised quantity. We ask whether this principle remains useful after images are mapped into a learned latent space, where compression has already removed much of the raw pixel variability. We introduce JLT, a 130M latent diffusion Transformer over frozen FLUX.2 VAE codes, and compare clean-latent prediction with a matched velocity-prediction DiT under the same representation, backbone, and training settings. Although the three variables x, epsilon, and v are linearly convertible for a fixed corruption time, a local Gaussian analysis shows that velocity regression inherits an isotropic target-covariance floor and amplifies low-variance latent directions, while clean prediction damps them. On ImageNet 256 x 256, JLT-B/1 obtains FID-50K 2.50 with classifier-free guidance, with a large matched-target gap over velocity prediction. These results suggest that prediction targets in latent diffusion are representation-dependent geometric choices, rather than interchangeable algebraic parameterizations.

扩散模型潜在空间生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。