构建复杂搜索任务基准,挑战智能体长程推理能力
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling

- 基于知识图谱自动生成高复杂度搜索问题
- 最强模型仅34.74%准确率,远低于人类上限
- 适合评估长期记忆与多步推理的智能体
现有搜索代理评测集(如BrowseComp)在一年内迅速饱和,最强模型准确率已超90%。由于这些评测主要由人工创建,标注者缺乏对实体统计的全局视角,无法系统化扩大搜索空间和结构复杂性,导致难度天花板难以突破。为此,我们提出LoHoSearch(长时程搜索代理基准),包含544个跨11个领域的经人工验证问题。该基准基于覆盖超过700万维基百科实体的知识图谱,通过选择具有大搜索空间的关系并组装成结构复杂的题目,确保答案唯一且经知识图谱验证。评估显示,即使最强模型准确率也仅为34.74%,现有上下文管理策略(最佳提升6.8%)带来的增益远小于以往评测。LoHoSearch为评估搜索代理的长时程推理与上下文管理能力提供了更严苛的标准。
原文摘要 · Abstract (English)
Search agent benchmarks exemplified by BrowseComp have rapidly saturated over the past year, with the strongest models surpassing 90% accuracy. Since these benchmarks are predominantly human-authored, annotators lack a global perspective on entity statistics and cannot systematically maximize search space size and structural complexity. This creates a difficulty ceiling that is hard to break. To address this, we introduce LoHoSearch (Long-Horizon Search Agents), a challenging benchmark comprising 544 human-verified questions across 11 domains. LoHoSearch is constructed via an automated pipeline built upon a knowledge graph covering over 7 million Wikipedia entities, which selects relations with large search spaces and assembles them into structurally complex questions with KG-verified unique answers. Our evaluation demonstrates that even the strongest model achieves only 34.74% accuracy, and existing context management strategies (best +6.8%) yield far smaller gains than on prior benchmarks. LoHoSearch provides a more demanding standard for evaluating long-horizon reasoning and context management in search agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。