对比11个大模型在10项日常任务中的表现,找性价比最高的实用模型。
Sustainability via LLM Right-sizing
- 用双模型框架自动评测任务完成质量,覆盖输出、准确率、伦理等10项标准。
- 小模型如Gemma-3和Phi-4在多数任务中表现可靠,成本与能耗显著更低。
- 提出按任务类型评估模型“够用性”,更适合企业可持续部署需求。
大型语言模型(LLMs)已广泛嵌入组织工作流,引发对其能耗、财务成本和数据主权的担忧。尽管性能基准常推崇顶尖模型,但实际部署需更全面考量:何时更小、可本地部署的模型已“足够好”?本研究通过评估11个专有及开源的LLM在10项日常职业任务(如文本摘要、日程生成、邮件与提案撰写)上的表现,给出实证答案。采用基于双模型的评估框架,自动化执行任务并标准化评估10项指标,涵盖输出质量、事实准确性与伦理责任。结果表明,GPT-4o表现始终领先,但成本与环境足迹显著更高。值得注意的是,较小模型如Gemma-3和Phi-4在多数任务中实现强而稳定的表现,显示其在成本效率、本地部署或隐私要求高的场景中具备可行性。聚类分析揭示三类模型群组——高端全能型、胜任通用型、局限但安全型,凸显质量、控制与可持续性间的权衡。任务类型显著影响模型效能:概念类任务挑战普遍,而聚合与转换类任务表现更优。我们主张从追求性能最大化的基准转向任务与情境感知的“够用性”评估,更契合组织优先级。本方法提供可扩展的可持续性评估范式,并为负责任的LLM部署提供实践指导。
原文摘要 · Abstract (English)
Large language models (LLMs) have become increasingly embedded in organizational workflows. This has raised concerns over their energy consumption, financial costs, and data sovereignty. While performance benchmarks often celebrate cutting-edge models, real-world deployment decisions require a broader perspective: when is a smaller, locally deployable model "good enough"? This study offers an empirical answer by evaluating eleven proprietary and open-weight LLMs across ten everyday occupational tasks, including summarizing texts, generating schedules, and drafting emails and proposals. Using a dual-LLM-based evaluation framework, we automated task execution and standardized evaluation across ten criteria related to output quality, factual accuracy, and ethical responsibility. Results show that GPT-4o delivers consistently superior performance but at a significantly higher cost and environmental footprint. Notably, smaller models like Gemma-3 and Phi-4 achieved strong and reliable results on most tasks, suggesting their viability in contexts requiring cost-efficiency, local deployment, or privacy. A cluster analysis revealed three model groups -- premium all-rounders, competent generalists, and limited but safe performers -- highlighting trade-offs between quality, control, and sustainability. Significantly, task type influenced model effectiveness: conceptual tasks challenged most models, while aggregation and transformation tasks yielded better performances. We argue for a shift from performance-maximizing benchmarks to task- and context-aware sufficiency assessments that better reflect organizational priorities. Our approach contributes a scalable method to evaluate AI models through a sustainability lens and offers actionable guidance for responsible LLM deployment in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。