用压缩+蒸馏+分层合成,20倍提速生成高质量2K视频。
Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
- 在压缩潜空间中运行,降低计算开销
- 5秒24帧2K视频生成速度提升20倍
- 适合需要高效高分辨率视频生成的场景
随着消费者对超清视觉体验需求上升,2K视频合成需求日益增长。尽管扩散变换器(DiTs)在高质量视频生成方面表现优异,但将其扩展至2K分辨率因内存和计算成本的二次增长而难以实现。本文提出Turbo2K框架,通过在高度压缩的潜空间中运行,显著降低计算复杂度与内存占用,使2K视频合成成为可能。为克服高压缩比与小模型带来的生成质量限制,引入知识蒸馏策略,使小型学生模型继承大型教师模型的生成能力。分析表明,不同架构的DiTs在内部表示上存在结构相似性,支持有效知识迁移。此外,设计分层两阶段合成框架:先在低分辨率生成多层级特征,再引导高分辨率视频生成,确保结构连贯性与细节精细度,同时消除冗余编码解码开销。Turbo2K实现业界领先的效率,可生成5秒、24帧每秒、2K分辨率视频,推理速度相比现有方法最高提升20倍,显著提升高分辨率视频生成的可扩展性与实用性。
原文摘要 · Abstract (English)
Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals. While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scaling them to 2K resolution remains computationally prohibitive due to quadratic growth in memory and processing costs. In this work, we propose Turbo2K, an efficient and practical framework for generating detail-rich 2K videos while significantly improving training and inference efficiency. First, Turbo2K operates in a highly compressed latent space, reducing computational complexity and memory footprint, making high-resolution video synthesis feasible. However, the high compression ratio of the VAE and limited model size impose constraints on generative quality. To mitigate this, we introduce a knowledge distillation strategy that enables a smaller student model to inherit the generative capacity of a larger, more powerful teacher model. Our analysis reveals that, despite differences in latent spaces and architectures, DiTs exhibit structural similarities in their internal representations, facilitating effective knowledge transfer. Second, we design a hierarchical two-stage synthesis framework that first generates multi-level feature at lower resolutions before guiding high-resolution video generation. This approach ensures structural coherence and fine-grained detail refinement while eliminating redundant encoding-decoding overhead, further enhancing computational efficiency.Turbo2K achieves state-of-the-art efficiency, generating 5-second, 24fps, 2K videos with significantly reduced computational cost. Compared to existing methods, Turbo2K is up to 20$\times$ faster for inference, making high-resolution video generation more scalable and practical for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。