新评估框架同时检查对话每轮和整体质量,更准识别错误。
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
- 分两步评估:先看每轮对话的连贯性、知识一致性和策略合规性
- 在MultiWOZ 2.4和τ-Bench上比传统方法更贴近人工判断
- 适合想提升对话系统评测精度的研究者与开发者
任务导向型对话(TOD)系统受大语言模型推动迎来革新,但评估方法仍不足以应对其日益复杂的特性。传统自动指标仅关注对话整体层面,无法发现用户与代理交互中可能出现的关键中间错误。本文提出TD-EVAL(Turn and Dialogue-level Evaluation),一种结合细粒度轮次级分析与整体对话级比较的两步评估框架。在轮次层面,从对话连贯性、后端知识一致性、策略合规性三个维度评估每条回复;同时设计了TOD Agent Arena,通过成对比较提供对话级质量度量。在MultiWOZ 2.4和τ-Bench数据集上的实验表明,TD-EVAL能有效识别传统指标遗漏的对话错误,并且在与人工判断的一致性上优于传统及基于LLM的评估方法。结果证明,TD-EVAL为TOD系统评估引入了新范式,可高效评估轮次与系统层级表现,且具备即插即用特性,适用于未来研究。
原文摘要 · Abstract (English)
Task-oriented dialogue (TOD) systems are experiencing a revolution driven by Large Language Models (LLMs), yet the evaluation methodologies for these systems remain insufficient for their growing sophistication. While traditional automatic metrics effectively assessed earlier modular systems, they focus solely on the dialogue level and cannot detect critical intermediate errors that can arise during user-agent interactions. In this paper, we introduce TD-EVAL (Turn and Dialogue-level Evaluation), a two-step evaluation framework that unifies fine-grained turn-level analysis with holistic dialogue-level comparisons. At turn level, we evaluate each response along three TOD-specific dimensions: conversation cohesion, backend knowledge consistency, and policy compliance. Meanwhile, we design TOD Agent Arena that uses pairwise comparisons to provide a measure of dialogue-level quality. Through experiments on MultiWOZ 2.4 and τ-Bench, we demonstrate that TD-EVAL effectively identifies the conversational errors that conventional metrics miss. Furthermore, TD-EVAL exhibits better alignment with human judgments than traditional and LLM-based metrics. These findings demonstrate that TD-EVAL introduces a new paradigm for TOD system evaluation, efficiently assessing both turn and system levels with a plug-and-play framework for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。