测试大模型在企业环境中的工作能力,发现当前表现仅41.8%。
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- 构建模拟企业场景的基准测试集EnterpriseBench
- 500个任务中顶尖模型仅完成41.8%
- 适合研究企业AI agent或智能办公系统的人参考
企业系统对提升员工与客户生产力和决策能力至关重要。将基于大语言模型(LLM)的系统融入企业环境,可实现智能自动化、个性化体验和高效信息检索,推动运营效率与战略增长。然而,由于企业环境数据分散于多个来源且受复杂访问控制约束,开发与评估此类系统极具挑战性。我们提出EnterpriseBench,一个全面的基准测试,模拟真实企业场景,包含500个跨软件工程、人力资源、财务及行政领域的多样化任务。该基准独特地捕捉了数据源碎片化、访问控制层级和跨职能工作流等关键企业特征。此外,我们设计了一种新颖的数据生成流程,基于组织元数据生成内部一致的企业任务。对最先进LLM代理的实验表明,即使最强模型也仅能完成41.8%的任务,凸显企业在面向AI系统方面仍有巨大改进空间。
原文摘要 · Abstract (English)
Enterprise systems are crucial for enhancing productivity and decision-making among employees and customers. Integrating LLM based systems into enterprise systems enables intelligent automation, personalized experiences, and efficient information retrieval, driving operational efficiency and strategic growth. However, developing and evaluating such systems is challenging due to the inherent complexity of enterprise environments, where data is fragmented across multiple sources and governed by sophisticated access controls. We present EnterpriseBench, a comprehensive benchmark that simulates enterprise settings, featuring 500 diverse tasks across software engineering, HR, finance, and administrative domains. Our benchmark uniquely captures key enterprise characteristics including data source fragmentation, access control hierarchies, and cross-functional workflows. Additionally, we provide a novel data generation pipeline that creates internally consistent enterprise tasks from organizational metadata. Experiments with state-of-the-art LLM agents demonstrate that even the most capable models achieve only 41.8% task completion, highlighting significant opportunities for improvement in enterprise-focused AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。