arXiv:2506.07280cs.CVcs.AI2025-06被引 6

视频扩散模型可零样本迁移,用少量样本学会新任务。

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

  • 将任务转为视觉变化序列,用少量样本微调模型
  • 仅靠几例训练,即在分割、姿态估计等任务上表现良好
  • 适合想快速部署视觉模型的研究者与开发者

视频扩散模型(VDMs)作为强大的生成工具,能合成高质量时空内容。但其潜力远超生成本身。我们发现,为建模连贯时序而进行的训练,自然促使模型内化结构化表征与对视觉世界的隐含理解。为此,我们提出一种少样本微调框架,仅用少量示例即可将冻结的VDM重用于新任务。该方法将每个任务转化为一个视觉过渡过程,通过短输入-输出序列训练LoRA权重,无需改变生成接口。尽管监督极小,模型在多种任务中展现出强大泛化能力,涵盖低层视觉(如分割、姿态估计)和高层推理(如ARC-AGI)。这些结果重新定义了VDMs:它们不仅是生成引擎,更是可适配的视觉学习者,有望成为未来视觉基础模型的核心。

原文摘要 · Abstract (English)

Video Diffusion Models (VDMs) have emerged as powerful generative tools, capable of synthesizing high-quality spatiotemporal content. Yet, their potential goes far beyond mere video generation. We argue that the training dynamics of VDMs, driven by the need to model coherent sequences, naturally pushes them to internalize structured representations and an implicit understanding of the visual world. To probe the extent of this internal knowledge, we introduce a few-shot fine-tuning framework that repurposes VDMs for new tasks using only a handful of examples. Our method transforms each task into a visual transition, enabling the training of LoRA weights on short input-output sequences without altering the generative interface of a frozen VDM. Despite minimal supervision, the model exhibits strong generalization across diverse tasks, from low-level vision (for example, segmentation and pose estimation) to high-level reasoning (for example, on ARC-AGI). These results reframe VDMs as more than generative engines. They are adaptable visual learners with the potential to serve as the backbone for future foundation models in vision.

视频生成扩散模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。