提出视频扩散模型的精准缩放定律,降低训练成本40.1%
Towards Precise Scaling Laws for Video Diffusion Transformers
- 基于计算预算和模型规模,建立可预测最优超参数的新缩放法则
- 在1e10 TFlops预算下,推理成本降低40.1%且性能相当
- 适用于实际部署中算力受限场景,支持非最优模型的性能预估
由于视频扩散变换器训练成本高昂,如何在给定数据与算力预算下实现最优性能至关重要。尽管语言模型中已有缩放定律用于性能预测,但视觉生成模型中的缩放规律仍缺乏系统研究。本文首次系统分析视频扩散变换器的缩放规律并证实其存在。发现与语言模型不同,视频扩散模型对学习率和批量大小更敏感,而这两项超参数常被忽略。为此,我们提出一种新缩放定律,可预测任意模型规模与算力预算下的最优超参数。在1e10 TFlops算力预算下,该方法实现与传统方法相当的性能,同时推理成本降低40.1%。此外,我们建立了验证损失、模型规模与算力预算间的更通用精确关系,支持对非最优模型规模的性能预测,有助于在实际推理成本约束下实现更好权衡。
原文摘要 · Abstract (English)
Achieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training hyperparameters before large-scale training. While scaling laws are employed in language models to predict performance, their existence and accurate derivation in visual generation models remain underexplored. In this paper, we systematically analyze scaling laws for video diffusion transformers and confirm their presence. Moreover, we discover that, unlike language models, video diffusion models are more sensitive to learning rate and batch size, two hyperparameters often not precisely modeled. To address this, we propose a new scaling law that predicts optimal hyperparameters for any model size and compute budget. Under these optimal settings, we achieve comparable performance and reduce inference costs by 40.1% compared to conventional scaling methods, within a compute budget of 1e10 TFlops. Furthermore, we establish a more generalized and precise relationship among validation loss, any model size, and compute budget. This enables performance prediction for non-optimal model sizes, which may also be appealed under practical inference cost constraints, achieving a better trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。