arXiv:2605.27882cs.CLcs.AI2026-05

构建真实搜索场景的基准,测试大模型长时主动搜索能力

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

  • 设计多轮对话、无固定模板的混合语言任务,模拟真实用户模糊意图
  • 7个前沿模型在专业与日常任务上最佳F1仅30.30,暴露长程推理短板
  • 适合研究多轮交互、主动探查与知识构建的AI搜索团队

基于大语言模型的智能体在现有搜索评测中表现良好,但真实用户反馈结果仍不理想,暴露出评估与体验之间的持续差距。我们归因于现有基准依赖过度明确的查询、单轮交互和固定模式评估,无法反映真实搜索中用户与代理通过多轮对话协同澄清模糊意图的行为。为此,我们提出新范式VibeSearch,并构建VibeSearchBench基准,包含200个手工标注的中英文双语任务,覆盖20个领域,分为专业(VibeSearch-Pro)与日常生活(VibeSearch-Daily)两个子集。每项任务配有一个用户角色与无结构的知识图谱真值,采用渐进披露用户模拟器与图匹配评估框架进行评测。我们在ReAct框架与OpenClaw代理系统下对7个前沿模型进行评测,结果显示所有模型在VibeSearch任务上仍显著不足(最高F1为30.30),凸显在长上下文推理、主动意图挖掘与结构化知识构建方面亟需根本性突破。

原文摘要 · Abstract (English)

LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries, single-turn interactions, and fixed-schema evaluation, none of which reflect real search behavior where users and agents collaboratively refine vague intent through multi-turn dialogue. We term this paradigm VibeSearch and introduce VibeSearchBench, a benchmark comprising 200 manually curated bilingual (Chinese and English) tasks across 20 domains, split into VibeSearch-Pro (professional) and VibeSearch-Daily (daily-life) subsets. Each task pairs a user persona with a schema-free ground-truth knowledge graph, and is evaluated through a progressive-disclosure user simulator and a graph-matching evaluation framework. We benchmark seven frontier models under both the ReAct framework and the OpenClaw agent harness. Results show that all models remain substantially inadequate for VibeSearch (best F1: 30.30), highlighting the need for fundamental advances in long-context reasoning, proactive intent elicitation, and structured knowledge construction.

搜索评测多轮对话长程推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。