0.6B参数模型,5秒内于手机生成5秒视频,媲美云端效果。
SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
- 轻量架构+时序层优化,适配移动端硬件效率。
- 仅需4步去噪,0.6B参数模型在iPhone上5秒生成5秒视频。
- 适合移动设备用户快速创作高质量短视频。
过去一年,基于扩散的视频生成取得了前所未有的成功。近期社区提出的模型可从任意提示词生成具有流畅运动和电影级画质的高分辨率视频。然而,作为图像生成的超复杂任务,视频生成模型需要更多计算资源,主要部署在云服务器上,限制了内容创作者的广泛使用。本文提出一个全面的加速框架,将大规模视频扩散模型的能力带到边缘用户手中。在网络架构层面,我们从紧凑的图像骨干网络出发,搜索最优的时序层设计与布局以最大化硬件效率。此外,提出专用对抗微调算法,并将去噪步骤减少至4步。所提模型仅含0.6B参数,可在iPhone 16 PM上于5秒内生成一段5秒视频。相比依赖强大GPU、单视频生成需数分钟的服务器模型,本方案实现数量级加速,同时保持相当的生成质量。
原文摘要 · Abstract (English)
We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。