arXiv:2609.02054cs.CL2026-09

用三个智能体模拟对话,评估大模型的追问能力。

A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models

  • 设计三智能体框架:提问者、用户模拟器、评价员协同评测。
  • 提出五项指标,量化评估澄清效果与对话效率。
  • 适合研发对话系统或需要精准理解用户意图的场景。

大型语言模型(LLMs)在交互式系统中日益重要,精准理解用户意图是关键。当用户问题模糊或不完整时,有效追问能力尤为关键。本文提出一种新型三智能体框架,用于稳健评估LLM的澄清对话能力。框架包含三个基于LLM的智能体:(1) 问答澄清代理(QCA),被评估的系统,负责识别歧义并提出澄清问题;(2) 回应代理(RA),模拟人类用户回复,可能包含无关或挑战性回答;(3) 评价代理(EA),作为大模型裁判,依据多项指标评估对话质量。我们以供应链领域为例,介绍合成数据生成方法,并提出评估歧义处理、问题质量、对话效率、语言得体性及最终意图对齐的综合指标。同时简要讨论了评价代理与人工判断的一致性验证。该工作为基准测试、验证和提升对话类大模型的澄清能力提供了结构化方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarification, especially when user queries are ambiguous or underspecified. This paper introduces a novel tri-agent framework for the robust evaluation of an LLM's ability to engage in clarifying dialogue. Our framework comprises three distinct LLM-based agents: (1) a Question Clarifying Agent (QCA), the system under evaluation, tasked with identifying ambiguities and posing clarifying questions; (2) a Respondent Agent (RA), designed to simulate human user responses, potentially including irrelevant or challenging replies; and (3) an Evaluator Agent (EA), an LLM-as-a-judge, which assesses the quality of the dialogue based on a comprehensive set of metrics. We detail a methodology for synthetic data generation in the supply chain domain as an example. We propose metrics evaluating ambiguity handling, question quality, dialogue efficiency, language appropriateness, and final intent alignment. We also briefly discuss the validation of the EA against human judgments. This work provides a structured approach to benchmark, validate, and improve the clarification capabilities of conversational LLM applications.

对话系统大模型评估意图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。