arXiv:2508.17756cs.LGcs.SY2025-08被引 6

无需重训练,高效生成超高清视频。

SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling

  • 基于分块技术,无需额外训练即可支持多种分辨率。
  • 通过自适应缓存减少内存和计算开销,加速生成过程。
  • 适合需要高分辨率视频生成的创作者与工业级应用。

扩散模型在生成任务(如图像和视频生成)中取得了显著进展,各领域对高质量内容(如2K/4K视频)的需求迅速增长。然而,现有标准分辨率(如720p)平台在生成超高清视频时仍面临过度重训练需求及高昂的计算与内存成本。为此,我们提出SUPERGEN,一种高效的基于分块的超高清视频生成框架。SUPERGEN采用无需训练的算法创新,结合分块策略,可无缝支持多种分辨率,同时显著降低内存占用和计算复杂度。此外,其引入了面向分块的自适应、区域感知缓存机制,利用去噪步骤与空间区域间的冗余性加速生成。SUPERGEN还集成缓存引导、通信最小化的分块并行策略,提升吞吐量并降低延迟。评估表明,SUPERGEN在多个基准测试中实现高性能与高输出质量的平衡。

原文摘要 · Abstract (English)

Diffusion models have recently achieved remarkable success in generative tasks (e.g., image and video generation), and the demand for high-quality content (e.g., 2K/4K videos) is rapidly increasing across various domains. However, generating ultra-high-resolution videos on existing standard-resolution (e.g., 720p) platforms remains challenging due to the excessive re-training requirements and prohibitively high computational and memory costs. To this end, we introduce SUPERGEN, an efficient tile-based framework for ultra-high-resolution video generation. SUPERGEN features a novel training-free algorithmic innovation with tiling to successfully support a wide range of resolutions without additional training efforts while significantly reducing both memory footprint and computational complexity. Moreover, SUPERGEN incorporates a tile-tailored, adaptive, region-aware caching strategy that accelerates video generation by exploiting redundancy across denoising steps and spatial regions. SUPERGEN also integrates cache-guided, communication-minimized tile parallelism for enhanced throughput and minimized latency. Evaluations show that SUPERGEN maximizes performance gains while achieving high output quality across various benchmarks.

视频生成扩散模型超分辨率高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。