构建真实客服场景的评测数据集,挑战大模型在实际运营中的表现。
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios
- 用大模型生成模拟客服系统实体,覆盖知识库与对话流程。
- 五项任务测试显示顶尖模型文章搜索准确率仅68%,知识库优化F1仅0.3。
- 适合研究客服智能、RAG系统或工业级AI评估的开发者参考。
大型语言模型在客户体验管理(CXM)领域潜力巨大,尤其在客服中心应用中。然而,由于隐私限制导致数据稀缺,现有评测基准缺乏真实性,难以涵盖知识库集成、现实噪声和关键运营任务。为此,我们提出CXMArena——一个大规模合成基准数据集,专为评估人工智能在真实运营环境下的表现而设计。通过可控噪声注入与专家指导,构建了包括产品规格、问题分类体系和真实对话在内的品牌级实体。基于此,发布五个核心任务评测:知识库优化、意图识别、代理质量合规、文章检索及多轮带工具的RAG。基线实验表明,即使最先进的嵌入与生成模型,文章检索准确率也仅达68%,传统嵌入方法在知识库优化上F1仅为0.3,凸显当前模型面临严峻挑战,亟需复杂流水线而非常规方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) hold immense potential for revolutionizing Customer Experience Management (CXM), particularly in contact center operations. However, evaluating their practical utility in complex operational environments is hindered by data scarcity (due to privacy concerns) and the limitations of current benchmarks. Existing benchmarks often lack realism, failing to incorporate deep knowledge base (KB) integration, real-world noise, or critical operational tasks beyond conversational fluency. To bridge this gap, we introduce CXMArena, a novel, large-scale synthetic benchmark dataset specifically designed for evaluating AI in operational CXM contexts. Given the diversity in possible contact center features, we have developed a scalable LLM-powered pipeline that simulates the brand's CXM entities that form the foundation of our datasets-such as knowledge articles including product specifications, issue taxonomies, and contact center conversations. The entities closely represent real-world distribution because of controlled noise injection (informed by domain experts) and rigorous automated validation. Building on this, we release CXMArena, which provides dedicated benchmarks targeting five important operational tasks: Knowledge Base Refinement, Intent Prediction, Agent Quality Adherence, Article Search, and Multi-turn RAG with Integrated Tools. Our baseline experiments underscore the benchmark's difficulty: even state of the art embedding and generation models achieve only 68% accuracy on article search, while standard embedding methods yield a low F1 score of 0.3 for knowledge base refinement, highlighting significant challenges for current models necessitating complex pipelines and solutions over conventional techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。