arXiv:2505.09358cs.CVcs.LG2025-05TPAMI被引 96

用少量数据微调扩散模型,实现无需标注的高精度图像分析。

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

  • 基于预训练扩散模型,仅需小规模合成数据微调。
  • 零样本迁移下在深度估计等任务上达顶尖性能。
  • 适合资源有限但需高质量图像分析的研究者。

过去十年计算机视觉的发展依赖大规模标注数据集和强预训练模型。在数据稀缺场景中,预训练模型的质量对有效迁移学习至关重要。传统预训练方法以图像分类和自监督学习为主,而近年来基于文本到图像生成、采用潜空间去噪扩散的模型,凭借海量带字幕图像数据训练,展现出对视觉世界的深层理解。本文提出Marigold,一类条件生成模型及其微调协议,可从Stable Diffusion等预训练潜空间扩散模型中提取知识,并适配于密集图像分析任务,包括单目深度估计、表面法线预测和固有分解。Marigold对预训练模型架构改动极小,仅需在单张GPU上用小规模合成数据训练数日,即可实现当前最优的零样本泛化能力。

原文摘要 · Abstract (English)

The success of deep learning in computer vision over the past decade has hinged on large labeled datasets and strong pretrained models. In data-scarce settings, the quality of these pretrained models becomes crucial for effective transfer learning. Image classification and self-supervised learning have traditionally been the primary methods for pretraining CNNs and transformer-based architectures. Recently, the rise of text-to-image generative models, particularly those using denoising diffusion in a latent space, has introduced a new class of foundational models trained on massive, captioned image datasets. These models' ability to generate realistic images of unseen content suggests they possess a deep understanding of the visual world. In this work, we present Marigold, a family of conditional generative models and a fine-tuning protocol that extracts the knowledge from pretrained latent diffusion models like Stable Diffusion and adapts them for dense image analysis tasks, including monocular depth estimation, surface normals prediction, and intrinsic decomposition. Marigold requires minimal modification of the pre-trained latent diffusion model's architecture, trains with small synthetic datasets on a single GPU over a few days, and demonstrates state-of-the-art zero-shot generalization. Project page: https://marigoldcomputervision.github.io

扩散模型图像分析零样本微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。