arXiv:2512.21120cs.CLcs.IR2025-12被引 12

构建多轮对话澄清评估基准,发现大模型常提前回答。

ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models

  • 基于五维模糊性分类与六类用户角色生成6120条对话
  • 十款大模型普遍提前回应,对话越深表现越差
  • 提出ClarifyAgent框架,提升多轮澄清鲁棒性

大型语言模型在开放域多轮对话中日益作为对话助手使用,但用户常提供不完整或模糊信息。现有聚焦于大模型的澄清评估基准多假设单轮交互或合作用户,难以评估真实场景下的澄清行为。我们提出 extbf{ClarifyMT-Bench},一个基于五维模糊性分类和六种行为多样化的模拟用户人格的多轮澄清评估基准。通过混合式大模型-人类流水线,构建了6,120条多轮对话,涵盖多种模糊来源与互动模式。评估十款代表性大模型发现,普遍存在未充分澄清的倾向:模型倾向于过早回答,且随着对话深度增加性能持续下降。为此,我们提出 extbf{ClarifyAgent},一种将澄清分解为感知、预测、追踪与规划的智能体方法,在各类模糊条件下显著提升鲁棒性。ClarifyMT-Bench为研究大模型何时应提问、何时应作答,以及如何应对现实人机交互中的模糊性,建立了可复现的基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as conversational assistants in open-domain, multi-turn settings, where users often provide incomplete or ambiguous information. However, existing LLM-focused clarification benchmarks primarily assume single-turn interactions or cooperative users, limiting their ability to evaluate clarification behavior in realistic settings. We introduce \textbf{ClarifyMT-Bench}, a benchmark for multi-turn clarification grounded in a five-dimensional ambiguity taxonomy and a set of six behaviorally diverse simulated user personas. Through a hybrid LLM-human pipeline, we construct 6,120 multi-turn dialogues capturing diverse ambiguity sources and interaction patterns. Evaluating ten representative LLMs uncovers a consistent under-clarification bias: LLMs tend to answer prematurely, and performance degrades as dialogue depth increases. To mitigate this, we propose \textbf{ClarifyAgent}, an agentic approach that decomposes clarification into perception, forecasting, tracking, and planning, substantially improving robustness across ambiguity conditions. ClarifyMT-Bench establishes a reproducible foundation for studying when LLMs should ask, when they should answer, and how to navigate ambiguity in real-world human-LLM interactions.

对话系统大模型评估多轮澄清智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。