GISA构建了真实信息搜索任务的基准,推动智能助手能力评估
GISA: A Benchmark for General Information-Seeking Assistant
- 人工设计373个真实搜索问题,避免传统反向构造偏差
- 最佳模型仅达19.3%准确率,复杂任务表现更差
- 提供完整搜索轨迹,适合过程监督与模仿学习
大语言模型的发展加速了能够通过多轮网络交互自主获取信息的搜索代理的进步。现有基准多从答案反推问题,生成不自然的任务,且常聚焦单一信息定位或聚合,依赖静态答案集易导致数据泄露。为此,我们提出GISA——一个通用信息搜索助手基准,包含373个真实场景的人工编写查询。GISA采用四种结构化答案格式(项、集合、列表、表格),支持确定性评估;统一任务中融合深度推理与广泛信息聚合,并设有定期更新的答案子集以防止记忆。尤为关键的是,每个查询均提供完整的用户搜索轨迹,为过程级监督与模仿学习提供黄金标准。在主流LLM及商用搜索产品上的实验表明,即使最优模型也仅达19.30%精确匹配得分,复杂规划与全面信息收集任务性能显著下降,凸显未来改进空间。
原文摘要 · Abstract (English)
The advancement of large language models (LLMs) has significantly accelerated the development of search agents capable of autonomously gathering information through multi-turn web interactions. Various benchmarks have been proposed to evaluate such agents. However, existing benchmarks often construct queries backward from answers, producing unnatural tasks misaligned with real-world needs. Moreover, these benchmarks tend to focus on either locating specific information or aggregating information from multiple sources, while relying on static answer sets prone to data contamination. To bridge these gaps, we introduce GISA, a benchmark for General Information-Seeking Assistants comprising 373 human-crafted queries that reflect authentic information-seeking scenarios. GISA features four structured answer formats (item, set, list, and table), enabling deterministic evaluation. It integrates both deep reasoning and broad information aggregation within unified tasks, and includes a live subset with periodically updated answers to resist memorization. Notably, GISA provides complete human search trajectories for every query, offering gold-standard references for process-level supervision and imitation learning. Experiments on mainstream LLMs and commercial search products reveal that even the best-performing model achieves only 19.30\% exact match score, with performance notably degrading on tasks requiring complex planning and comprehensive information gathering. These findings highlight substantial room for future improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。