arXiv:2508.15478cs.CLcs.CY2025-08EMNLP被引 10

首个评估小模型性能与环保影响的综合基准,助力绿色AI落地。

SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version

  • 构建多维度评测体系,覆盖准确率、计算效率与能耗
  • 在4种硬件上测试15个模型,横跨9项任务与23个数据集
  • 开源工具链支持可复现研究,适合模型选型与低碳设计

小型语言模型(SLMs)具备计算高效和易部署的优势,但其性能与环境影响的系统性评估仍不充分。本文提出SLM-Bench,首个专为评估SLMs在准确性、计算效率和可持续性方面表现而设计的综合性基准。该基准在4种硬件配置下,对15个SLMs在9项NLP任务上的表现进行评估,涵盖23个数据集及14个应用领域。不同于以往基准,SLM-Bench量化了11项指标,涵盖正确性、计算开销与资源消耗,实现对效率权衡的全面分析。所有测试在受控硬件条件下进行,确保模型间公平比较。我们开发了开源基准测试流水线,采用标准化协议,提升可复现性与研究可扩展性。结果表明,不同模型在准确率与能效之间存在显著权衡:部分模型精度领先,另一些则表现出更优的能量效率。SLM-Bench为小模型评估树立新标准,弥合资源效率与实际应用间的差距。

原文摘要 · Abstract (English)

Small Language Models (SLMs) offer computational efficiency and accessibility, yet a systematic evaluation of their performance and environmental impact remains lacking. We introduce SLM-Bench, the first benchmark specifically designed to assess SLMs across multiple dimensions, including accuracy, computational efficiency, and sustainability metrics. SLM-Bench evaluates 15 SLMs on 9 NLP tasks using 23 datasets spanning 14 domains. The evaluation is conducted on 4 hardware configurations, providing a rigorous comparison of their effectiveness. Unlike prior benchmarks, SLM-Bench quantifies 11 metrics across correctness, computation, and consumption, enabling a holistic assessment of efficiency trade-offs. Our evaluation considers controlled hardware conditions, ensuring fair comparisons across models. We develop an open-source benchmarking pipeline with standardized evaluation protocols to facilitate reproducibility and further research. Our findings highlight the diverse trade-offs among SLMs, where some models excel in accuracy while others achieve superior energy efficiency. SLM-Bench sets a new standard for SLM evaluation, bridging the gap between resource efficiency and real-world applicability.

小模型绿色AI基准测试能效评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。