arXiv:2508.14170cs.CLcs.CY2025-08中稿 · Nature Sci Rep, 32…被引 2

对比大模型推理的能耗与准确率,发现二者不总相关。

Comparing energy consumption and accuracy in text classification inference

  • 在不同模型和硬件上实测推理能耗与准确率关系
  • 大模型能耗高但零样本下准确率未必更高
  • 运行时间可作为能耗的实用替代指标

大型语言模型(LLMs)在自然语言处理任务中的广泛应用引发了对能效与可持续性的关注。尽管以往研究多聚焦于训练阶段的能耗,推理阶段却较少被关注。本研究系统评估了多种模型架构与硬件配置下文本分类推理中准确率与能耗之间的权衡。实证分析显示,在某些场景下表现最佳的模型同时具备较高的能效。虽然LLMs通常比传统机器学习模型消耗更多能量,但在我们的零样本分类设置中,其准确率仅持平或更低。我们观察到推理能耗存在显著差异(<1mWh至>1kWh),受模型类型、规模及硬件配置影响。此外,推理能耗与运行时间呈强相关性,表明在无法直接测量时,执行时间可作为能耗的实用代理。结果表明,能效与准确率是独立的评估维度,不一定一致。我们认为,可持续人工智能发展需系统评估性能与资源效率。

原文摘要 · Abstract (English)

The increasing deployment of large language models (LLMs) in natural language processing (NLP) tasks raises concerns about energy efficiency and sustainability. While prior research has largely focused on energy consumption during model training, the inference phase has received comparatively less attention. This study systematically evaluates the trade-offs between model accuracy and energy consumption in text classification inference across various model architectures and hardware configurations. Our empirical analysis shows that in some contexts the best-performing model in terms of accuracy can also be energy-efficient. While LLMs tend to consume significantly more energy than traditional machine learning models, they show the same or even lower levels of accuracy in our zero-shot classification setting. We observe substantial variability in inference energy consumption ($<$mWh to $>$kWh), influenced by model type, model size, and hardware specifications. Additionally, we find a strong correlation between inference energy consumption and model runtime, indicating that execution time can serve as a practical proxy for energy usage in settings where direct measurement is not feasible. Our findings demonstrate that energy efficiency and accuracy represent distinct evaluation dimensions that do not necessarily align. We argue that sustainable AI development requires systematic evaluation of both performance and resource efficiency.

能耗评估大模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。