实测判别与生成式AI能耗,给出绿色部署实用指南
Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations
- 通过软件测功法分析模型、超参与硬件对能耗影响
- 优化架构与配置可降耗不降效,大模型低负载时未必更耗能
- 为绿色MLOps/GenOps提供可复现的能耗评估基准
本研究针对真实世界MLOps流水线中判别型与生成型AI模型的能耗开展实证分析。对判别模型,考察训练与推理阶段不同架构与超参对能耗的影响,识别出节能实践;对生成式AI,重点评估大语言模型(LLMs)在不同模型规模与服务请求下的能耗表现。采用软件级功耗测量方法,确保在多种配置、模型与数据集间可复现。分析多模型与硬件组合,揭示各项指标间的关联,识别能耗关键驱动因素。结果表明,优化模型架构、超参数与硬件可显著降低判别模型能耗而不牺牲性能;对于LLMs,能耗效率取决于模型规模、推理复杂度与请求处理能力的平衡,模型越大并不必然导致更高能耗,当利用率低时尤为明显。研究为设计绿色可持续的ML运营提供实用指导,强调在保持性能前提下降低能耗与碳足迹,并可作为估算各类AI模型总能耗的基准参考。
原文摘要 · Abstract (English)
This study presents an empirical investigation into the energy consumption of Discriminative and Generative AI models within real-world MLOps pipelines. For Discriminative models, we examine various architectures and hyperparameters during training and inference and identify energy-efficient practices. For Generative AI, Large Language Models (LLMs) are assessed, focusing primarily on energy consumption across different model sizes and varying service requests. Our study employs software-based power measurements, ensuring ease of replication across diverse configurations, models, and datasets. We analyse multiple models and hardware setups to uncover correlations among various metrics, identifying key contributors to energy consumption. The results indicate that for Discriminative models, optimising architectures, hyperparameters, and hardware can significantly reduce energy consumption without sacrificing performance. For LLMs, energy efficiency depends on balancing model size, reasoning complexity, and request-handling capacity, as larger models do not necessarily consume more energy when utilisation remains low. This analysis provides practical guidelines for designing green and sustainable ML operations, emphasising energy consumption and carbon footprint reductions while maintaining performance. This paper can serve as a benchmark for accurately estimating total energy use across different types of AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。