arXiv:2411.00427cs.CLcs.AI2024-11被引 10

用多个专业代理协作,让对话系统更懂多领域任务。

DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems

  • 分域代理由中央管理器调度,各司其职处理不同任务。
  • 在MultiWOZ上提升信息覆盖率6.6%,任务成功率4.1%。
  • 适合需要跨领域灵活应答的客服与助手系统。

任务导向型对话系统在客户服务到个人助理等场景中至关重要,但构建跨多领域的有效系统仍面临挑战,因需处理多样用户意图、实体类型及领域知识。本文提出DARD(领域分配响应委派)框架,采用由中央对话管理器协调的分域代理结构。通过在多种代理建模方法间对比实验,结合小规模微调模型(Flan-T5-large、Mistral-7B)与大规模语言模型(Claude Sonnet 3.0)的优势,揭示了该多代理架构在灵活性与可组合性上的优势。在标准的MultiWOZ基准上评估,DARD实现领先性能:对话信息覆盖率提升6.6%,任务成功率达4.1%。此外,还讨论了MultiWOZ数据集及其评估体系中的标注差异与问题。

原文摘要 · Abstract (English)

Task-oriented dialogue systems are essential for applications ranging from customer service to personal assistants and are widely used across various industries. However, developing effective multi-domain systems remains a significant challenge due to the complexity of handling diverse user intents, entity types, and domain-specific knowledge across several domains. In this work, we propose DARD (Domain Assigned Response Delegation), a multi-agent conversational system capable of successfully handling multi-domain dialogs. DARD leverages domain-specific agents, orchestrated by a central dialog manager agent. Our extensive experiments compare and utilize various agent modeling approaches, combining the strengths of smaller fine-tuned models (Flan-T5-large & Mistral-7B) with their larger counterparts, Large Language Models (LLMs) (Claude Sonnet 3.0). We provide insights into the strengths and limitations of each approach, highlighting the benefits of our multi-agent framework in terms of flexibility and composability. We evaluate DARD using the well-established MultiWOZ benchmark, achieving state-of-the-art performance by improving the dialogue inform rate by 6.6% and the success rate by 4.1% over the best-performing existing approaches. Additionally, we discuss various annotator discrepancies and issues within the MultiWOZ dataset and its evaluation system.

对话系统多智能体任务导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。