提出URSA框架,让离散视频生成更高效精准
Uniform Discrete Diffusion with Metric Path for Video Generation
- 通过全局迭代优化离散时空标记,实现高效视频生成
- 仅需少量推理步数即可生成高分辨率长时视频
- 支持图像到视频、插值等多任务统一建模
连续空间视频生成发展迅速,而离散方法因误差累积和长时序不一致问题进展滞后。本文重新审视离散生成建模,提出均匀离散扩散与度量路径框架(URSA),一种简单但强大的可扩展视频生成方法。核心是将视频生成视为离散时空标记的迭代全局优化过程,融合线性度量路径与分辨率相关的时间步偏移机制,使模型能高效拓展至高分辨率图像合成与长时视频生成,同时显著减少推理步数。此外,引入异步时间微调策略,统一多个任务(如插值、图像到视频生成)于单一模型。在多个挑战性视频与图像生成基准上实验表明,URSA持续优于现有离散方法,性能接近顶尖连续扩散模型。代码与模型已开源。
原文摘要 · Abstract (English)
Continuous-space video generation has advanced rapidly, while discrete approaches lag behind due to error accumulation and long-context inconsistency. In this work, we revisit discrete generative modeling and present Uniform discRete diffuSion with metric pAth (URSA), a simple yet powerful framework that bridges the gap with continuous approaches for the scalable video generation. At its core, URSA formulates the video generation task as an iterative global refinement of discrete spatiotemporal tokens. It integrates two key designs: a Linearized Metric Path and a Resolution-dependent Timestep Shifting mechanism. These designs enable URSA to scale efficiently to high-resolution image synthesis and long-duration video generation, while requiring significantly fewer inference steps. Additionally, we introduce an asynchronous temporal fine-tuning strategy that unifies versatile tasks within a single model, including interpolation and image-to-video generation. Extensive experiments on challenging video and image generation benchmarks demonstrate that URSA consistently outperforms existing discrete methods and achieves performance comparable to state-of-the-art continuous diffusion methods. Code and models are available at https://github.com/baaivision/URSA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。