arXiv:2512.20000cs.CV2025-12

用少量样本训练动作模块,实现精准视频生成控制。

Few-Shot-Based Modular Image-to-Video Adapter for Diffusion Models

  • 设计可并行的轻量动作模块,每个模块学习一种运动模式。
  • 仅需约十张样本,单张消费级显卡即可训练完成。
  • 支持用户自由组合动作模块,无需提示词工程。

扩散模型(DMs)在图像和视频生成中实现了惊人的逼真效果,但在图像动画应用上仍受限,即使在大规模数据集上训练亦然。主要挑战在于视频信号维度高导致训练数据稀缺,使模型更倾向记忆而非遵循提示生成动作;同时,模型难以泛化到训练集中未出现的新运动模式,尤其在有限数据下微调仍不充分。为此,我们提出模块化图像到视频适配器(MIVA),一个可附加于预训练扩散模型的轻量子网络,每个模块专门捕获单一运动模式,可通过并行扩展。MIVA可在约十组样本上高效训练,使用单个消费级显卡即可完成。推理时,用户通过选择一个或多个MIVA指定动作,无需提示词工程。大量实验表明,MIVA在保持甚至超越大规模训练模型生成质量的同时,实现更精确的运动控制。

原文摘要 · Abstract (English)

Diffusion models (DMs) have recently achieved impressive photorealism in image and video generation. However, their application to image animation remains limited, even when trained on large-scale datasets. Two primary challenges contribute to this: the high dimensionality of video signals leads to a scarcity of training data, causing DMs to favor memorization over prompt compliance when generating motion; moreover, DMs struggle to generalize to novel motion patterns not present in the training set, and fine-tuning them to learn such patterns, especially using limited training data, is still under-explored. To address these limitations, we propose Modular Image-to-Video Adapter (MIVA), a lightweight sub-network attachable to a pre-trained DM, each designed to capture a single motion pattern and scalable via parallelization. MIVAs can be efficiently trained on approximately ten samples using a single consumer-grade GPU. At inference time, users can specify motion by selecting one or multiple MIVAs, eliminating the need for prompt engineering. Extensive experiments demonstrate that MIVA enables more precise motion control while maintaining, or even surpassing, the generation quality of models trained on significantly larger datasets.

视频生成扩散模型少样本学习动作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。