arXiv:2607.04553cs.MMcs.AI2026-07

用架构原理估算视频生成能耗,无需模型参数即可评估碳排放。

Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption

论文配图:Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption
图 1 · 摘自论文原文
  • 从分辨率、时长等参数反推能耗,基于视频扩散模型的计算特性。
  • 能耗符合理论缩放规律,误差低于3%(MAPE),跨模型验证有效。
  • 适合关注视频生成可持续性的研究者与开发者参考。

我们提出一种双向框架,仅通过可观察的生成参数(如分辨率、时长)和架构基本原理,估算文本到视频(T2V)及文本到视频音频(T2VA)模型的能耗,无需访问权重、模型规模或实现细节。正向可预测能耗,反向可从推理时间恢复架构缩放行为,并以精度作为架构有效性判断标准。基于视频扩散模型固有的计算密集特性,我们证明各模型能耗均服从理论推导的缩放律,其分解为二次项与线性项,系数直接反映架构复杂度。在覆盖8.3B-27B参数的六款开源模型及三种GPU配置下,该分解方法整体MAPE低于3%。该方法为T2V模型与架构提供了标准化、兼具实证与理论基础的可持续性评估框架。

原文摘要 · Abstract (English)

We present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models from architectural first principles and observable generation parameters such as resolution and duration, requiring no access to weights, model size, or implementation details. Forward, it predicts energy from generation parameters and architectural principles; backward, it recovers architectural scaling behavior from observed inference times, with accuracy serving as a criterion for architectural validity. Building on the established compute-bound nature of video diffusion models, we demonstrate that each model's energy profile obeys theoretically derived scaling laws, decomposing into quadratic and linear terms whose coefficients directly reflect the underlying architectural complexity. Validated across six open-source models spanning 8.3B-27B parameters and three GPU configurations, this decomposition achieves below 3% MAPE across all architectures. This approach offers a standardized, empirically and theoretically grounded framework for sustainability benchmarking across T2V models and architectures.

视频生成能耗估算可持续性缩放律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。