让人类评估大模型时看到能耗数据,发现更节能的模型更受欢迎。
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
- 在公开评测平台中加入模型能耗信息,引导用户决策
- 多数问题下用户倾向选择能耗更低的小模型
- 适合关注绿色AI与模型效率的研究者和开发者
大语言模型的评估是一项复杂任务。传统自动化基准测试虽高效,但与人类判断的相关性较差;人工评估则因模型数量激增而难以规模化。现有公共评测平台(如LM Arena)允许用户自由对比模型回答并排名,但未考虑能源消耗。本文提出生成式能量竞技场(GEA),在评测过程中引入模型能耗信息。初步结果显示,在多数问题中,当用户了解能耗后,更倾向于选择更小、更节能的模型。这表明,对于大多数应用场景,高性能大模型带来的额外能耗与计算成本,并未换来用户感知质量的显著提升,因此不具性价比。
原文摘要 · Abstract (English)
The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to answer multiple-choice questions of different topics. However, this method has certain limitations, being the most concerning, the poor correlation with the humans. An alternative approach, is to have humans evaluate the LLMs. This poses scalability issues as there is a large and growing number of models to evaluate making it impractical (and costly) to run traditional studies based on recruiting a number of evaluators and having them rank the responses of the models. An alternative approach is the use of public arenas, such as the popular LM arena, on which any user can freely evaluate models on any question and rank the responses of two models. The results are then elaborated into a model ranking. An increasingly important aspect of LLMs is their energy consumption and, therefore, evaluating how energy awareness influences the decisions of humans in selecting a model is of interest. In this paper, we present GEA, the Generative Energy Arena, an arena that incorporates information on the energy consumption of the model in the evaluation process. Preliminary results obtained with GEA are also presented, showing that for most questions, when users are aware of the energy consumption, they favor smaller and more energy efficient models. This suggests that for most user interactions, the extra cost and energy incurred by the more complex and top-performing models do not provide an increase in the perceived quality of the responses that justifies their use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。