arXiv:2601.06676cs.CLcs.AI2026-01被引 2

首个评测大模型研究交互能力的基准,强调边问边查比盲目推理更高效。

One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

  • 设计交互式研究流程,允许模型主动提问澄清用户意图。
  • 7个模型实验显示互动使研究质量与鲁棒性显著提升。
  • 适合关注人机协作、动态需求响应的研究者使用。

由大语言模型驱动的深度研究代理可执行多步推理、网络探索和长文报告生成。然而,现有系统基本为自主运行,假设用户意图完全明确,仅评估最终输出。实际上,研究目标常不清晰且在探索中演变,但当前基准既未模拟动态用户反馈,也未衡量交互成本。为此,我们提出IDRBench,首个系统评估深度研究代理交互能力的基准。IDRBench将深度研究建模为交互过程,允许代理通过提问来更好对齐用户意图。其集成模块化交互框架、可扩展的参考对齐用户模拟器,以及兼顾对齐增益与交互开销的评估套件。在七个代表性专有与开源大模型上的实验表明,交互能持续提升研究质量与鲁棒性,同时揭示各模型间显著的交互效率差异。这些发现确立了交互能力作为独立评估维度,并使IDRBench成为未来面向用户的深度研究代理的可复用基准。

原文摘要 · Abstract (English)

Deep research agents powered by Large Language Models (LLMs) can perform multi-step reasoning, web exploration, and long-form report generation. However, existing systems remain largely autonomous, assuming fully specified user intent and evaluating only final outputs. In practice, research goals are often underspecified and evolve during exploration, yet current benchmarks neither model dynamic user feedback nor measure interaction costs. To address this gap, we introduce IDRBench, the first Interactive Deep Research Benchmark for systematically evaluating the interactive capabilities of deep research agents. IDRBench formulates deep research as an interactive process where agents may solicit clarification to better align with user intent. It integrates a modular interactive framework, a scalable reference-grounded user simulator, and an interaction-aware evaluation suite that jointly measures alignment gains and interaction overhead. Experiments on seven representative proprietary and open-weight LLMs show that interaction consistently improves research quality and robustness, while revealing substantial differences in interaction efficiency across models. These findings establish interactive capability as a distinct evaluation dimension and position IDRBench as a reusable benchmark for future user-aligned deep research agents.

交互式研究大模型评测用户对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。