arXiv:2509.19222cs.LG2025-09中稿 · NeurIPS被引 4

分析开源文生视频模型的延迟与能耗,揭示其规模增长规律。

Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models

  • 建立计算瓶颈模型,预测分辨率、时长和去噪步数对性能的影响。
  • 实验显示延迟随分辨率和时长呈二次增长,随去噪步数线性上升。
  • 对比六种模型,为绿色生成视频系统设计提供实证参考。

近期文本到视频(T2V)生成技术已实现从自然语言提示生成高质量、时间连贯视频片段。然而,这些系统伴随显著的计算开销,其能源需求仍不明确。本文系统研究了前沿开源T2V模型的延迟与能耗。我们首先构建一个计算瓶颈分析模型,预测空间分辨率、时间长度和去噪步数的缩放规律。随后在WAN2.1-T2V上通过细粒度实验验证,发现延迟随空间与时间维度呈二次增长,随去噪步数呈线性增长。最后,我们将分析扩展至六种不同T2V模型,在默认设置下比较其运行时与能耗表现。结果既提供了基准参考,也为设计更可持续的生成式视频系统提供了实践洞见。

原文摘要 · Abstract (English)

Recent advances in text-to-video (T2V) generation have enabled the creation of high-fidelity, temporally coherent clips from natural language prompts. Yet these systems come with significant computational costs, and their energy demands remain poorly understood. In this paper, we present a systematic study of the latency and energy consumption of state-of-the-art open-source T2V models. We first develop a compute-bound analytical model that predicts scaling laws with respect to spatial resolution, temporal length, and denoising steps. We then validate these predictions through fine-grained experiments on WAN2.1-T2V, showing quadratic growth with spatial and temporal dimensions, and linear scaling with the number of denoising steps. Finally, we extend our analysis to six diverse T2V models, comparing their runtime and energy profiles under default settings. Our results provide both a benchmark reference and practical insights for designing and deploying more sustainable generative video systems.

文生视频能耗分析生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。