首次发现扩散Transformer的缩放规律,可预测生成质量与资源需求。
Scaling Laws For Diffusion Transformers
- 通过大规模实验验证DiT损失与计算量呈幂律关系。
- 在1e21 FLOPs预算下,可准确预测10亿参数模型的图文生成损失。
- 损失趋势与生成质量(如FID)一致,适合快速评估模型性能。
扩散Transformer(DiT)已在内容复现任务(如图像和视频生成)中展现出优异的合成与扩展能力。然而,其缩放规律尚未被深入探索,而缩放规律通常能基于特定算力预算精确预测最优模型规模与数据需求。为此,本文首次在广泛算力范围内(1e17至6e18 FLOPs)进行实验,证实了DiT存在缩放规律。具体而言,预训练损失与计算量也遵循幂律关系。基于该规律,我们不仅能确定最优模型尺寸与所需数据量,还可准确预测在10亿参数模型、1e21 FLOPs算力预算下的文本到图像生成损失。此外,我们还证明,预训练损失趋势与生成性能(如FID)高度一致,即使跨不同数据集也是如此,从而补充了从算力到合成质量的映射关系,提供了一种低成本、可预测的模型与数据质量评估基准。
原文摘要 · Abstract (English)
Diffusion transformers (DiT) have already achieved appealing synthesis and scaling properties in content recreation, e.g., image and video generation. However, scaling laws of DiT are less explored, which usually offer precise predictions regarding optimal model size and data requirements given a specific compute budget. Therefore, experiments across a broad range of compute budgets, from 1e17 to 6e18 FLOPs are conducted to confirm the existence of scaling laws in DiT for the first time. Concretely, the loss of pretraining DiT also follows a power-law relationship with the involved compute. Based on the scaling law, we can not only determine the optimal model size and required data but also accurately predict the text-to-image generation loss given a model with 1B parameters and a compute budget of 1e21 FLOPs. Additionally, we also demonstrate that the trend of pre-training loss matches the generation performances (e.g., FID), even across various datasets, which complements the mapping from compute to synthesis quality and thus provides a predictable benchmark that assesses model performance and data quality at a reduced cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。