统一评测对话系统,揭示架构与提示策略影响
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
- 构建可插拔框架,统一用户模拟器与对话系统评测条件
- 在相同数据与约束下验证多系统,发现模型规模与提示策略显著影响性能
- 适合研发对话系统的研究者与工程师参考实际部署效果
指令微调的大语言模型(LLMs)推动了对话系统发展,使真实用户模拟和鲁棒的多轮对话代理成为可能。然而,现有研究常孤立评估单一用户模拟器或特定系统设计,限制了跨架构与配置的泛化结论。本文提出 clem todd(用于任务导向对话系统开发的聊天优化大语言模型),一个灵活的系统化评测框架,可在一致条件下对用户模拟器与对话系统组合进行详细评估。该框架支持即插即用集成,确保数据集、评估指标与计算约束的一致性。我们通过在统一设置中重新评估现有对话系统,并集成三个新提出的系统,展示了其灵活性。结果揭示了架构、模型规模与提示策略对对话性能的影响,为构建高效可靠的对话AI系统提供了实践指导。
原文摘要 · Abstract (English)
The emergence of instruction-tuned large language models (LLMs) has advanced the field of dialogue systems, enabling both realistic user simulations and robust multi-turn conversational agents. However, existing research often evaluates these components in isolation-either focusing on a single user simulator or a specific system design-limiting the generalisability of insights across architectures and configurations. In this work, we propose clem todd (chat-optimized LLMs for task-oriented dialogue systems development), a flexible framework for systematically evaluating dialogue systems under consistent conditions. clem todd enables detailed benchmarking across combinations of user simulators and dialogue systems, whether existing models from literature or newly developed ones. It supports plug-and-play integration and ensures uniform datasets, evaluation metrics, and computational constraints. We showcase clem todd's flexibility by re-evaluating existing task-oriented dialogue systems within this unified setup and integrating three newly proposed dialogue systems into the same evaluation pipeline. Our results provide actionable insights into how architecture, scale, and prompting strategies affect dialogue performance, offering practical guidance for building efficient and effective conversational AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。