arXiv:2607.00647cs.CV2026-07中稿 · ECCV

对比不同预测目标对无训练引导稳定性的提升效果。

Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold

论文配图:Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
图 1 · 摘自论文原文
  • 提出新理论分析不同预测目标在高噪声下的估计精度。
  • 发现x-prediction能最可靠保持生成样本在数据流形上。
  • 适合研究扩散模型推理稳定性与生成质量优化的读者。

无训练引导(TFG)在推理阶段引导预训练扩散模型实现特定属性。为保证效果,引导必须从采样初期的高噪声阶段开始应用。由于其目标函数(分类器或能量模型)定义于干净图像,ε-和v-预测模型需在每一步从噪声状态估计干净图像〤〥,估计准确性决定了引导是否偏离数据流形。x-预测作为新方法,直接输出干净图像,消除了高噪声下的估计误差。本文提供理论分析,揭示各预测目标对估计精度的影响,并引入引导类FID(Child FID),可暴露标准评估忽略的流形损伤。在新的细粒度鸟类基准和风格迁移任务上的实验表明,x-预测能最可靠地保持引导样本在数据流形上,是无训练引导的最佳基础。代码已开源。

原文摘要 · Abstract (English)

Training-free guidance (TFG) steers a pretrained diffusion model toward a desired attribute at inference. To be effective, this guidance must be applied from the earliest, high-noise steps of sampling. Because its objective (a classifier or energy) is defined on clean images, $ε$- and $v$-prediction models must first estimate the clean image $\hat{x}$ from the noisy state at each step, and the accuracy of that estimate determines how easily guidance drifts off the data manifold. $x$-prediction, a recent alternative, outputs the clean image directly, removing this source of error even at high noise. This is our motivation. We provide a theoretical analysis of how each prediction target shapes this accuracy, and introduce guided-class FID (Child FID), a metric that exposes the manifold damage standard evaluation misses. Experiments on a new fine-grained bird benchmark and on style transfer confirm that $x$-prediction keeps guided samples on the manifold most reliably, making it the strongest foundation for training-free guidance. Code is available at https://github.com/ManLuML/on-manifold-tfg

扩散模型生成质量无训练引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。