arXiv:2501.15398cs.CL2025-01被引 5

对比三模型微调能耗,发现大模型碳足迹显著更高

How Green are Neural Language Models? Analyzing Energy Consumption in Text Summarization Fine-tuning

  • 用多个评估指标测试三种模型在摘要任务中的表现
  • LLaMA-3-8B微调碳足迹最大,远高于T5-base和BART-base
  • 呼吁将能效纳入语言模型设计考量,适合关注AI环保的研究者

人工智能系统对环境影响显著,尤其在自然语言处理任务中。这些任务常需大量计算资源训练深层神经网络,包括含数十亿参数的大规模语言模型。本研究分析了三种神经语言模型在文本摘要微调中的能源消耗与性能权衡:两个预训练模型(T5-base 和 BART-base)及一个大语言模型(LLaMA-3-8B)。模型用于生成科研论文摘要,捕捉核心主题。通过测量各模型微调过程的碳足迹,全面评估其环境影响。结果显示,LLaMA-3-8B 的碳足迹在三者中最大。采用多种评价指标(包括 ROUGE、METEOR、MoverScore、BERTScore、SciBERTScore)评估模型性能。研究强调在模型设计与实现中纳入环境因素的重要性,并呼吁发展更节能的 AI 方法。

原文摘要 · Abstract (English)

Artificial intelligence systems significantly impact the environment, particularly in natural language processing (NLP) tasks. These tasks often require extensive computational resources to train deep neural networks, including large-scale language models containing billions of parameters. This study analyzes the trade-offs between energy consumption and performance across three neural language models: two pre-trained models (T5-base and BART-base), and one large language model (LLaMA-3-8B). These models were fine-tuned for the text summarization task, focusing on generating research paper highlights that encapsulate the core themes of each paper. The carbon footprint associated with fine-tuning each model was measured, offering a comprehensive assessment of their environmental impact. It is observed that LLaMA-3-8B produces the largest carbon footprint among the three models. A wide range of evaluation metrics, including ROUGE, METEOR, MoverScore, BERTScore, and SciBERTScore, were employed to assess the performance of the models on the given task. This research underscores the importance of incorporating environmental considerations into the design and implementation of neural language models and calls for the advancement of energy-efficient AI methodologies.

语言模型碳足迹能效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。