构建首个评估智能体式对话系统的基准与框架
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
- 设计合成对话生成管道,生成需长期推理的丰富标注对话
- 提出多维度评估框架,覆盖任务完成率、智能体能力与回复质量
- 适合研究对话系统智能体行为的学者与开发者
近期基于大语言模型(LLMs)并集成大量API和工具的任务导向对话(TOD)系统,使对话代理能够协调交错目标、保持长时上下文,并通过异步执行主动行动。这些能力超越了传统TOD系统,但现有基准缺乏对这类智能体行为的系统性评估支持。为此,我们提出ATOD,一个包含合成对话生成管道的基准,能生成需要长期推理的丰富标注对话。ATOD捕捉先进TOD的关键特征,包括多目标协调、依赖管理、记忆、适应性和主动性。基于ATOD,我们进一步提出ATOD-Eval,一个综合评估框架,将这些维度转化为细粒度指标,并支持可复现的离线与在线评估。我们还引入一种基于记忆的强效评估器用于在ATOD上基准测试。实验表明,ATOD-Eval能全面评估任务完成、智能体能力与回复质量,且该评估器相比现有基于记忆或LLM的方法,在准确率与效率间取得更优平衡。
原文摘要 · Abstract (English)
Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved goals, maintain long-horizon context, and act proactively through asynchronous execution. These capabilities extend beyond traditional TOD systems, yet existing benchmarks lack systematic support for evaluating such agentic behaviors. To address this gap, we introduce ATOD, a benchmark and synthetic dialogue generation pipeline that produces richly annotated conversations requiring long-term reasoning. ATOD captures key characteristics of advanced TOD, including multi-goal coordination, dependency management, memory, adaptability, and proactivity. Building on ATOD, we propose ATOD-Eval, a holistic evaluation framework that translates these dimensions into fine-grained metrics and supports reproducible offline and online evaluation. We further present a strong agentic memory-based evaluator for benchmarking on ATOD. Experiments show that ATOD-Eval enables comprehensive assessment across task completion, agentic capability, and response quality, and that the proposed evaluator offers a better accuracy-efficiency tradeoff compared to existing memory- and LLM-based approaches under this evaluation setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。