arXiv:2506.20274cs.AI2025-06被引 5

为企事业场景设计14项评估任务,更真实测试大模型能力。

Enterprise Large Language Model Evaluation Benchmark

  • 基于布鲁姆分类法构建14个企业任务评估框架。
  • 用9700条样本验证,开源模型在推理强但判断弱。
  • 适合想定制评测的大模型应用企业参考。

大型语言模型(LLMs)在提升人工智能工具效率方面展现出潜力,但现有基准如大规模多任务语言理解(MMLU)难以有效评估企业场景下的复杂任务。本文提出一个基于布鲁姆分类法的14项任务评估框架,全面衡量大模型在企业环境中的能力。为应对数据噪声和标注成本高的问题,开发了一套可扩展的流水线,结合大模型作为标注者、大模型作为评判者与校正检索增强生成(CRAG),构建了一个稳健的9,700样本基准数据集。对六款主流模型的评估显示,开源模型如DeepSeek R1在推理任务中可媲美闭源模型,但在依赖判断的任务中表现落后,可能因过度思考所致。该基准揭示了企业在实际应用中的关键性能差距,并为模型优化提供可行建议。本研究为企业量身定制评估方案提供了蓝图,推动大模型的实用部署。

原文摘要 · Abstract (English)

Large Language Models (LLMs) ) have demonstrated promise in boosting productivity across AI-powered tools, yet existing benchmarks like Massive Multitask Language Understanding (MMLU) inadequately assess enterprise-specific task complexities. We propose a 14-task framework grounded in Bloom's Taxonomy to holistically evaluate LLM capabilities in enterprise contexts. To address challenges of noisy data and costly annotation, we develop a scalable pipeline combining LLM-as-a-Labeler, LLM-as-a-Judge, and corrective retrieval-augmented generation (CRAG), curating a robust 9,700-sample benchmark. Evaluation of six leading models shows open-source contenders like DeepSeek R1 rival proprietary models in reasoning tasks but lag in judgment-based scenarios, likely due to overthinking. Our benchmark reveals critical enterprise performance gaps and offers actionable insights for model optimization. This work provides enterprises a blueprint for tailored evaluations and advances practical LLM deployment.

大模型评估企业应用评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。