arXiv:2412.03439cs.CV2024-12CVPR被引 49

让扩散模型直接输出无噪声语义特征,性能更优且成本更低

CleanDIFT: Diffusion Features without Noise

  • 提出轻量级无监督微调方法,使扩散模型可直接处理无噪声图像
  • 无噪声特征在多种任务中显著超越旧方法,且推理成本大幅降低
  • 适合需要高效特征提取的视觉应用,如图像检索与分类

大规模预训练扩散模型的内部特征最近被证明是多种下游任务的强大语义描述符。现有方法通常需在输入图像上添加噪声才能获得有效特征,因为模型在低噪声或无噪声输入下表现不佳。本文发现,噪声对特征质量有决定性影响,且无法通过集成不同随机噪声来弥补。为此,我们提出一种轻量级、无监督的微调方法,使扩散主干网络能直接生成高质量的无噪声语义特征。实验表明,这些特征在多种提取设置和下游任务中均显著优于先前方法,性能甚至超过基于集成的方案,且计算开销仅为后者的极小部分。

原文摘要 · Abstract (English)

Internal features from large-scale pre-trained diffusion models have recently been established as powerful semantic descriptors for a wide range of downstream tasks. Works that use these features generally need to add noise to images before passing them through the model to obtain the semantic features, as the models do not offer the most useful features when given images with little to no noise. We show that this noise has a critical impact on the usefulness of these features that cannot be remedied by ensembling with different random noises. We address this issue by introducing a lightweight, unsupervised fine-tuning method that enables diffusion backbones to provide high-quality, noise-free semantic features. We show that these features readily outperform previous diffusion features by a wide margin in a wide variety of extraction setups and downstream tasks, offering better performance than even ensemble-based methods at a fraction of the cost.

扩散模型特征提取无噪声微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。