arXiv:2505.07615eess.AScs.AI2025-05中稿 · WASPAA 2025被引 3

分析7种生成音频模型的能耗,找质量与省电的平衡点

Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models

  • 测试7个文本转音频扩散模型的推理能耗
  • 发现生成参数变化显著影响能耗,存在性能-能耗权衡
  • 为高效音频生成提供节能优化方向,适合关注可持续AI的研究者

文本到音频模型近年来成为从文本描述生成声音的强大技术,但其高计算需求引发对能耗和环境影响的担忧。本文分析了7个最先进的基于扩散的文本到音频生成模型在推理阶段的能源消耗,研究不同生成参数对能耗的影响程度。同时,通过考察所有选定模型的帕累托最优解,旨在找到音频质量与能耗之间的最佳平衡。研究结果揭示了性能与环境影响间的权衡关系,为开发更高效的生成式音频模型提供参考。

原文摘要 · Abstract (English)

Text-to-audio models have recently emerged as a powerful technology for generating sound from textual descriptions. However, their high computational demands raise concerns about energy consumption and environmental impact. In this paper, we conduct an analysis of the energy usage of 7 state-of-the-art text-to-audio diffusion-based generative models, evaluating to what extent variations in generation parameters affect energy consumption at inference time. We also aim to identify an optimal balance between audio quality and energy consumption by considering Pareto-optimal solutions across all selected models. Our findings provide insights into the trade-offs between performance and environmental impact, contributing to the development of more efficient generative audio models.

音频生成扩散模型能耗分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。