arXiv:2508.20973cs.CLcs.AI2025-08ACL被引 10

构建统一评估框架,系统测试大模型主动对话能力。

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents

  • 拆解主动对话为目标规划与对话引导,设计跨领域评价指标。
  • 生成328个覆盖6大领域的评测环境,验证模型表现差异。
  • 发现推理能力影响主动行为,指导未来模型优化方向。

主动对话已成为推动大语言模型发展的重要且具有挑战性的研究问题。现有工作多集中于特定领域或任务导向场景,导致评估碎片化,限制了对模型主动对话能力的全面探索。本文提出ProactiveEval,一个专为评估大语言模型主动对话能力而设计的统一框架。该框架将主动对话分解为目标规划与对话引导两个维度,并在多个领域建立评价指标;同时支持自动构建多样且具挑战性的评测数据。基于此框架,我们构建了涵盖6个不同领域的328个评测环境。通过在22种不同类型的大型语言模型上进行实验,结果显示DeepSeek-R1在目标规划任务中表现优异,Claude-3.7-Sonnet在对话引导任务中表现突出。最后,我们探究了推理能力对主动行为的影响,并讨论其对未来模型发展的启示。

原文摘要 · Abstract (English)

Proactive dialogue has emerged as a critical and challenging research problem in advancing large language models (LLMs). Existing works predominantly focus on domain-specific or task-oriented scenarios, which leads to fragmented evaluations and limits the comprehensive exploration of models' proactive conversation abilities. In this work, we propose ProactiveEval, a unified framework designed for evaluating proactive dialogue capabilities of LLMs. This framework decomposes proactive dialogue into target planning and dialogue guidance, establishing evaluation metrics across various domains. Moreover, it also enables the automatic generation of diverse and challenging evaluation data. Based on the proposed framework, we develop 328 evaluation environments spanning 6 distinct domains. Through experiments with 22 different types of LLMs, we show that DeepSeek-R1 and Claude-3.7-Sonnet exhibit exceptional performance on target planning and dialogue guidance tasks, respectively. Finally, we investigate how reasoning capabilities influence proactive behaviors and discuss their implications for future model development.

对话系统大模型评估主动对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。