arXiv:2411.10416cs.CLcs.AI2024-11被引 1

提出新评估方法,量化对话流程的质量与覆盖率。

Towards Automatic Evaluation of Task-Oriented Dialogue Flows

  • 基于模糊图编辑距离衡量对话流程结构复杂度
  • 能有效评估单个对话与流程的匹配程度
  • 适合对话系统设计者和自动化生成工具使用

任务型对话系统依赖预定义的对话流程(对话流),通常以有向无环图形式表示。这些流程可人工设计或从历史对话中自动生成。由于领域知识差异或训练数据不同,流程在图结构上可能显著不同。然而,目前尚无标准方法评估对话流程质量。本文提出FuDGE(模糊对话图编辑距离)这一新度量方法,通过评估流程的结构复杂性与对对话数据的表征覆盖能力来衡量其优劣。FuDGE计算单个对话与流程的对齐程度,进而反映一组对话被流程整体覆盖的效果。在人工配置流程及自动生成流程上进行大量实验,验证了FuDGE的有效性及其评估框架。通过标准化与优化对话流程,FuDGE使对话设计者与自动化技术实现更高效率与自动化水平。

原文摘要 · Abstract (English)

Task-oriented dialogue systems rely on predefined conversation schemes (dialogue flows) often represented as directed acyclic graphs. These flows can be manually designed or automatically generated from previously recorded conversations. Due to variations in domain expertise or reliance on different sets of prior conversations, these dialogue flows can manifest in significantly different graph structures. Despite their importance, there is no standard method for evaluating the quality of dialogue flows. We introduce FuDGE (Fuzzy Dialogue-Graph Edit Distance), a novel metric that evaluates dialogue flows by assessing their structural complexity and representational coverage of the conversation data. FuDGE measures how well individual conversations align with a flow and, consequently, how well a set of conversations is represented by the flow overall. Through extensive experiments on manually configured flows and flows generated by automated techniques, we demonstrate the effectiveness of FuDGE and its evaluation framework. By standardizing and optimizing dialogue flows, FuDGE enables conversational designers and automated techniques to achieve higher levels of efficiency and automation.

对话系统流程评估图编辑距离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。