为大模型在企业场景下的评估构建了25个领域数据集的系统性基准。
Enterprise Benchmarks for Large Language Model Evaluation
- 基于金融、法律等13个企业领域,整合25个公开数据集进行评估
- 13个模型在不同任务上表现差异显著,体现任务适配重要性
- 适合关注大模型实际落地的企业开发者与评测研究者
大语言模型(LLMs)的发展带来了对复杂任务评估的更高要求,尤其在企业应用中。因此,需要针对企业数据集对LLMs进行系统化评估。本文提出了一套面向大模型评估的基准策略,聚焦于领域特定数据集,涵盖多种自然语言处理任务。所提出的评估框架包含来自金融、法律、网络安全、气候与可持续发展等领域的25个公开数据集。13个模型在不同企业任务上的表现差异显著,凸显根据具体任务需求选择合适模型的重要性。代码和提示语已在GitHub开源。
原文摘要 · Abstract (English)
The advancement of large language models (LLMs) has led to a greater challenge of having a rigorous and systematic evaluation of complex tasks performed, especially in enterprise applications. Therefore, LLMs need to be able to benchmark enterprise datasets for various tasks. This work presents a systematic exploration of benchmarking strategies tailored to LLM evaluation, focusing on the utilization of domain-specific datasets and consisting of a variety of NLP tasks. The proposed evaluation framework encompasses 25 publicly available datasets from diverse enterprise domains like financial services, legal, cyber security, and climate and sustainability. The diverse performance of 13 models across different enterprise tasks highlights the importance of selecting the right model based on the specific requirements of each task. Code and prompts are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。