大模型能识别问题模糊但很少主动追问,上下文越丰富越不提问。
Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions

- 通过三类任务测试模型对模糊问题的识别与回应行为
- 显式判断时识别率高,但问答场景下90%仍直接回答
- 检索到更多上下文反而降低提问意愿,暴露行为断层
用户问题常缺乏明确指代,可能有多个合理解释。理想的助手应识别模糊性并主动提问澄清,而非默认假设意图。这需要两项能力:识别问题模糊,以及基于识别采取澄清行动而非直接回答。我们通过三种设置评估模型表现:标准问答、显式模糊性判断、行为分析(由裁判模型将回复分类为直接回答、拒绝或澄清问题)。结果发现识别与行为间存在明显差距:模型在被要求判断模糊性时能有效识别,但在问答任务中却几乎全部选择直接回答。检索到的上下文进一步加剧这一差距——虽提升答案可得性,却使模型更少提出澄清问题。
原文摘要 · Abstract (English)
User queries are often underspecified and may admit multiple valid interpretations. Rather than silently making assumptions about the user's intent, a helpful assistant should surface such ambiguity by asking a clarifying question. Doing so requires two abilities: recognizing that a query is ambiguous, and acting on that recognition by seeking clarification instead of answering directly. To study these abilities, we evaluate models on ambiguous, unambiguous, and disambiguated questions in three settings: standard question answering, explicit ambiguity judgment, and behavioral analysis, where a judge model classifies responses as direct answers, refusals, or clarifying questions. We find a clear gap between recognition and behavior: models often identify ambiguity when explicitly asked to judge it, yet in the QA setting they overwhelmingly default to direct answers. Retrieved context further widens this gap by improving answerability while making models even less likely to ask clarifying questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。