arXiv:2411.02305cs.CLcs.AI2024-11NAACL被引 66

测试大模型在真实客户管理任务中的表现,发现现有能力不足一半。

CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments

  • 设计九类真实客服任务,覆盖三类角色和16种业务对象。
  • 顶尖模型用ReAct提示仅完成不到40%任务,有工具调用也难超55%。
  • 适合想评估AI代理实战能力的研究者与企业应用开发者。

客户关系管理(CRM)系统是现代企业的核心,支持客户交互与数据管理。将AI代理集成到CRM中可自动化流程并提升服务个性化,但缺乏反映真实工作复杂性的评估基准,导致部署与评测困难。为此,我们提出CRMArena,一个基于行业专家指导和最佳实践的新型基准,涵盖九项真实职场任务,分属服务专员、分析师和管理者三类角色。该基准包含16种常见工业对象(如账户、订单、知识文章、工单)及其高度关联性,并引入投诉习惯、政策违规等隐变量以模拟真实数据分布。实验表明,当前最先进大模型使用ReAct提示时,在任务完成率低于40%;即使具备函数调用能力,也难以超过55%。结果凸显了增强代理函数调用与规则遵循能力的迫切需求。CRMArena向社区开放挑战:能稳定完成任务的系统将直接体现商业价值,适用于主流工作场景。

原文摘要 · Abstract (English)

Customer Relationship Management (CRM) systems are vital for modern enterprises, providing a foundation for managing customer interactions and data. Integrating AI agents into CRM systems can automate routine processes and enhance personalized service. However, deploying and evaluating these agents is challenging due to the lack of realistic benchmarks that reflect the complexity of real-world CRM tasks. To address this issue, we introduce CRMArena, a novel benchmark designed to evaluate AI agents on realistic tasks grounded in professional work environments. Following guidance from CRM experts and industry best practices, we designed CRMArena with nine customer service tasks distributed across three personas: service agent, analyst, and manager. The benchmark includes 16 commonly used industrial objects (e.g., account, order, knowledge article, case) with high interconnectivity, along with latent variables (e.g., complaint habits, policy violations) to simulate realistic data distributions. Experimental results reveal that state-of-the-art LLM agents succeed in less than 40% of the tasks with ReAct prompting, and less than 55% even with function-calling abilities. Our findings highlight the need for enhanced agent capabilities in function-calling and rule-following to be deployed in real-world work environments. CRMArena is an open challenge to the community: systems that can reliably complete tasks showcase direct business value in a popular work environment.

大模型代理客户管理任务评估真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。