测试大模型在模糊对话中主动提问的能力,发现现有模型普遍不会问关键问题。
InfoQuest: Evaluating Multi-Turn Dialogue Agents for Open-Ended Conversations with Hidden Context
- 设计多轮对话评测集,模拟用户隐含需求场景
- 所有模型需多次追问才能理解意图,多数仍给出泛化回答
- 适合研究对话系统鲁棒性与主动澄清能力的学者
大型语言模型擅长执行明确指令,但在面对模糊或不完整用户请求时,常以冗长通用回复应对,而非主动寻求澄清。我们提出 InfoQuest,一个用于评估对话代理在开放式请求中处理隐藏上下文能力的多轮对话基准。该基准设计了有意模糊的场景,要求模型通过提问获取关键信息后再作答。对开放与闭源模型的评估显示,尽管专有模型表现更优,但所有当前助手均难以有效收集必要信息:往往需多轮交互才能推断用户意图,且频繁直接输出通用回应。我们提供了一套生成多样化场景和评估信息获取能力的系统方法,可用于自动生成数据以实现模型自我改进。同时,本研究揭示了语言模型在多轮互动中处理模糊请求的现存局限。
原文摘要 · Abstract (English)
Large language models excel at following explicit instructions, but they often struggle with ambiguous or incomplete user requests, defaulting to verbose, generic responses instead of seeking clarification. We introduce InfoQuest, a multi-turn chat benchmark designed to evaluate how dialogue agents handle hidden context in open-ended user requests. This benchmark presents intentionally ambiguous scenarios that require models to engage in information-seeking dialogue by asking clarifying questions before providing appropriate responses. Our evaluation of both open and closed models reveals that, while proprietary models generally perform better, all current assistants struggle to gather critical information effectively. They often require multiple turns to infer user intent and frequently default to generic responses without proper clarification. We provide a systematic methodology for generating diverse scenarios and evaluating models' information-seeking capabilities, which can be leveraged to automatically generate data for self-improvement. We also offer insights into the current limitations of language models in handling ambiguous requests through multi-turn interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。