arXiv:2505.18878cs.CLcs.AI2025-05被引 38

评测大模型在真实商业场景中的综合表现,发现其多轮对话和保密意识严重不足。

CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions

  • 设计十九项专家验证任务,覆盖售前、客服等多元商业流程
  • 单轮成功率仅58%,多轮对话下降至35%,保密意识几乎为零
  • 适合关注企业级AI应用落地与模型短板的研究者

尽管人工智能代理在商业领域具有变革潜力,但有效性能评估受限于公开、真实的业务数据稀缺。现有基准普遍缺乏环境、数据及人机交互的真实性,且覆盖的商业场景和行业有限。为此,我们提出CRMArena-Pro,一个面向多样化专业场景的全链路大模型代理评估基准。该基准在原CRMArena基础上扩展了十九项由专家验证的任务,涵盖销售、服务及‘配置、定价、报价’流程,支持B2B与B2C场景。其特色在于引入多轮交互、多样人格角色模拟以及强保密意识评估。实验表明,顶级大模型在CRMArena-Pro上的单轮成功率为约58%,多轮场景下骤降至约35%。尽管工作流执行表现较好(单轮超83%成功),其他核心业务能力仍具挑战。此外,模型本身几乎无保密意识;虽可通过提示改善,但常以牺牲任务表现为代价。结果揭示当前大模型能力与企业实际需求之间存在显著差距,亟需提升多轮推理、保密合规与多技能融合能力。

原文摘要 · Abstract (English)

While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platforms. Existing benchmarks often lack fidelity in their environments, data, and agent-user interactions, with limited coverage of diverse business scenarios and industries. To address these gaps, we introduce CRMArena-Pro, a novel benchmark for holistic, realistic assessment of LLM agents in diverse professional settings. CRMArena-Pro expands on CRMArena with nineteen expert-validated tasks across sales, service, and 'configure, price, and quote' processes, for both Business-to-Business and Business-to-Customer scenarios. It distinctively incorporates multi-turn interactions guided by diverse personas and robust confidentiality awareness assessments. Experiments reveal leading LLM agents achieve only around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings. While Workflow Execution proves more tractable for top agents (over 83% single-turn success), other evaluated business skills present greater challenges. Furthermore, agents exhibit near-zero inherent confidentiality awareness; though targeted prompting can improve this, it often compromises task performance. These findings highlight a substantial gap between current LLM capabilities and enterprise demands, underscoring the need for advancements in multi-turn reasoning, confidentiality adherence, and versatile skill acquisition.

大模型评估商业智能多轮对话保密性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。