arXiv:2502.17321cs.CL2025-02ACL被引 6

从对话中自动提取服务流程,提升AI客服一致性。

Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents

  • 用关键词检索相关对话,再通过问答式思维链生成流程
  • 在ABCD和SynthABCD数据集上准确率提升12.16%
  • 自动生成模拟客户与代理,可替代人工评估

自动化服务代理需要结构化的工作流程以提供一致且准确的回应。然而,这些流程往往未被记录,且从未有研究探索如何从历史交互中自动提取。本文提出一种新框架,用于从对话中提取并评估对话工作流程。提取过程包含两个关键阶段:(1) 基于关键流程元素检索相关对话;(2) 采用基于问答的思维链(QA-CoT)提示生成结构化流程。为全面评估提取流程的质量,我们引入一个自动化代理与客户机器人仿真框架,用于衡量其解决客户问题的有效性。在ABCD和SynthABCD数据集上的大量实验表明,我们的QA-CoT方法在平均宏准确率上比基线提升12.16%。此外,评估结果与人类评估高度一致,为未来研究提供了可靠且可扩展的框架。

原文摘要 · Abstract (English)

Automated service agents require well-structured workflows to provide consistent and accurate responses to customer queries. However, these workflows are often undocumented, and their automatic extraction from conversations remains unexplored. In this work, we present a novel framework for extracting and evaluating dialog workflows from historical interactions. Our extraction process consists of two key stages: (1) a retrieval step to select relevant conversations based on key procedural elements, and (2) a structured workflow generation process using a question-answer-based chain-of-thought (QA-CoT) prompting. To comprehensively assess the quality of extracted workflows, we introduce an automated agent and customer bots simulation framework that measures their effectiveness in resolving customer issues. Extensive experiments on the ABCD and SynthABCD datasets demonstrate that our QA-CoT technique improves workflow extraction by 12.16\% in average macro accuracy over the baseline. Moreover, our evaluation method closely aligns with human assessments, providing a reliable and scalable framework for future research.

对话系统工作流提取自动化评估服务AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。