arXiv:2507.16116cs.CV2025-07被引 4

用向量化时间步适配技术,实现视频生成的精细时序控制。

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

  • 提出向量化时间步适配(VTA),无需重训练即可精准控制视频时序。
  • 零样本实现图像转视频、起止帧指定与视频扩展,性能媲美全量微调。
  • 保持原始模型能力,适合快速部署与多任务应用的开发者使用。

视频扩散模型的快速发展受限于时间建模的固有缺陷,尤其是传统标量时间步变量带来的帧演化刚性同步。尽管特定任务适配和自回归模型尝试解决此问题,但仍受计算效率低、灾难性遗忘或适用范围窄的制约。本文提出Pusa V1.0,采用向量化时间步适配(VTA)技术,在统一框架内实现细粒度时序控制。VTA为非破坏性适配,完全保留基础模型能力。相较于需大量资源微调的Wan-I2V方法,Pusa在极简微调后即实现零样本图像到视频生成,效果相当。此外,该方法同时解锁起始-结束帧控制、视频扩展等多类零样本功能,且不牺牲原文本到视频(T2V)能力。机制分析表明,该方法保留基础模型生成先验,仅精准注入时序动态,避免向量化时间步带来的组合爆炸。本工作建立了可扩展、高效、通用的下一代视频生成范式,推动高质量视频生成在科研与产业中的普及。

原文摘要 · Abstract (English)

The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchronization of frame evolution imposed by conventional scalar timestep variables. While task-specific adaptations and autoregressive models have sought to address these challenges, they remain constrained by computational inefficiency, catastrophic forgetting, or narrow applicability. In this work, we present \textbf{Pusa} V1.0, a versatile model that leverages \textbf{vectorized timestep adaptation (VTA)} to enable fine-grained temporal control within a unified video diffusion framework. Note that VTA is a non-destructive adaptation, which means that it fully preserves the capabilities of the base model. Unlike conventional methods like Wan-I2V, which finetune a base text-to-video (T2V) model with abundant resources to do image-to-video (I2V), we achieve comparable results in a zero-shot manner after an ultra-efficient finetuning process based on VTA. Moreover, this method also unlocks many other zero-shot capabilities simultaneously, such as start-end frames and video extension -- all without task-specific training. Meanwhile, it keeps the T2V capability from the base model. Mechanistic analyses also reveal that our approach preserves the foundation model's generative priors while surgically injecting temporal dynamics, avoiding the combinatorial explosion inherent to the vectorized timestep. This work establishes a scalable, efficient, and versatile paradigm for next-generation video synthesis, democratizing high-fidelity video generation for research and industry alike.

视频生成扩散模型时序控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。