arXiv:2601.12761cs.CV2026-01被引 1

让视频扩散模型学会感知运动,实现零样本运动迁移。

Moaw: Unleashing Motion Awareness for Video Diffusion Models

  • 将扩散模型从图像生成转向视频到稠密追踪的运动感知训练。
  • 构建运动标注数据集,定位强运动信息特征并注入生成模型。
  • 无需额外适配器即可实现零样本运动迁移,适合可控视频生成研究者。

视频扩散模型在大规模数据上训练后,天然捕捉帧间共享特征的对应关系。近期工作利用此特性在零样本设置下实现光流预测与跟踪。受此启发,我们探究监督训练能否更充分挖掘视频扩散模型的追踪能力。为此,提出Moaw框架,释放视频扩散模型的运动感知能力,并用于运动迁移。具体而言,训练一个扩散模型进行运动感知,将其模态从图像到视频生成转变为视频到稠密追踪。随后构建运动标注数据集,识别编码最强运动信息的特征,并注入结构相同的视频生成模型。由于两网络同构,这些特征可零样本自然适配,实现无需额外适配器的运动迁移。本工作为生成建模与运动理解的融合提供新范式,推动更统一、可控的视频学习框架发展。

原文摘要 · Abstract (English)

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot setting. Motivated by these findings, we investigate whether supervised training can more fully harness the tracking capability of video diffusion models. To this end, we propose Moaw, a framework that unleashes motion awareness for video diffusion models and leverages it to facilitate motion transfer. Specifically, we train a diffusion model for motion perception, shifting its modality from image-to-video generation to video-to-dense-tracking. We then construct a motion-labeled dataset to identify features that encode the strongest motion information, and inject them into a structurally identical video generation model. Owing to the homogeneity between the two networks, these features can be naturally adapted in a zero-shot manner, enabling motion transfer without additional adapters. Our work provides a new paradigm for bridging generative modeling and motion understanding, paving the way for more unified and controllable video learning frameworks.

视频生成扩散模型运动感知零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。