arXiv:2606.27669cs.CL2026-06

评测大模型搜索代理在模糊查询下主动提问的能力

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

论文配图:When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search
图 1 · 摘自论文原文
  • 构建包含463处模糊点的多轮交互基准测试集
  • 发现当前模型误判率高,频繁搜索反而比直接猜测更差
  • 适合研究人机交互、智能搜索与大模型推理的学者

由大语言模型驱动的搜索代理日益用于解决复杂信息检索任务,需通过多步检索与推理达成用户目标。然而,现有基准常假设用户查询完整明确,忽略了真实搜索请求中普遍存在模糊、不完整甚至事实错误的问题。在深度搜索场景中,此类模糊性会沿多步推理链传播,导致代理走向错误路径。为此,我们提出DiscoBench,一个面向澄清意识的深度搜索评估基准,旨在检验搜索代理是否能主动识别模糊点、提出有效澄清问题,并通过用户交互恢复正确推理路径。DiscoBench涵盖11个真实领域,共211个样本和463个模糊实例,覆盖四种模糊类型。我们还设计了多轮交互用户模拟器,从任务效用、模糊检测、交互策略和成本效率四方面评估模型表现。对代表性LLM的实验表明,模糊检测与有效澄清是独立能力,重复搜索而非提问的表现往往劣于直接猜测,揭示当前搜索代理在检索能力与交互式解决问题之间存在显著差距。

原文摘要 · Abstract (English)

Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval and reasoning to fulfill user goals. However, existing benchmarks often assume that user queries are complete and explicit, overlooking the fact that real-world search requests are frequently vague, underspecified, or even factually incorrect. In deep search scenarios, such ambiguity can propagate along multi-step reasoning chains and lead agents toward incorrect search trajectories. To address this gap, we introduce DiscoBench, a benchmark for clarification-aware deep search, designed to evaluate whether search agents can proactively identify ambiguity, ask effective clarification questions, and recover correct reasoning paths through user interaction. DiscoBench contains 211 samples and 463 ambiguity instances across 11 real-world domains, covering four ambiguity types. We further design a user simulator for multi-turn interaction and evaluate model performance from four perspectives: task utility, ambiguity detection, interaction strategy, and cost efficiency. Experiments on representative LLMs show that ambiguity detection and effective clarification are distinct capabilities, and that repeatedly searching instead of asking for clarification often performs worse than direct guessing, highlighting a critical gap between retrieval ability and interactive problem-solving in current search agents.

搜索代理大模型交互评估模糊查询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。