用视频模型做图像修复,少样本也能达到顶尖效果
V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration
- 把图像修复看作逐步生成过程,借视频模型模拟修复细节
- 仅用1000个样本训练,就实现多任务修复,性能媲美专用模型
- 适合想用大模型做少样本视觉任务的研究者和开发者
大规模视频生成模型在海量多样的视觉数据上训练,内化了丰富的结构、语义和动态先验。尽管这些模型展现出强大的生成能力,其作为通用视觉学习者的潜力仍远未被挖掘。本文提出V-Bridge框架,将视频生成模型的潜在能力迁移至多样化的少样本图像修复任务中。我们重新定义图像修复为一个渐进的生成过程,利用视频模型模拟从退化输入到高保真输出的逐步优化。令人意外的是,仅需1000个多元任务训练样本(不足现有修复方法的2%),预训练视频模型即可被引导完成具有竞争力的图像修复,单模型实现多项任务,性能可与专为该任务设计的架构相媲美。研究发现,视频生成模型隐式学习到了强大且可迁移的修复先验,只需极少数据即可激活,挑战了生成建模与低层视觉之间的传统界限,为视觉任务中的基础模型设计开辟了新范式。
原文摘要 · Abstract (English)
Large-scale video generative models are trained on vast and diverse visual data, enabling them to internalize rich structural, semantic, and dynamic priors of the visual world. While these models have demonstrated impressive generative capability, their potential as general-purpose visual learners remains largely untapped. In this work, we introduce V-Bridge, a framework that bridges this latent capacity to versatile few-shot image restoration tasks. We reinterpret image restoration not as a static regression problem, but as a progressive generative process, and leverage video models to simulate the gradual refinement from degraded inputs to high-fidelity outputs. Surprisingly, with only 1,000 multi-task training samples (less than 2% of existing restoration methods), pretrained video models can be induced to perform competitive image restoration, achieving multiple tasks with a single model, rivaling specialized architectures designed explicitly for this purpose. Our findings reveal that video generative models implicitly learn powerful and transferable restoration priors that can be activated with only extremely limited data, challenging the traditional boundary between generative modeling and low-level vision, and opening a new design paradigm for foundation models in visual tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。