arXiv:2508.09129cs.AI2025-08被引 14

用程序化代理对提升网页浏览效率,解决搜索广度与推理深度的矛盾。

BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair

  • 分拆规划器与执行器,分工协作实现高效搜索与连贯推理。
  • 在英/中文基准上分别达30.0和46.5分,显著优于现有模型。
  • 适合需要大规模信息检索与复杂推理的智能助手场景。

在庞大且持续增长的数字环境中高效获取信息,需平衡广泛搜索与策略性推理。当前基于大语言模型(LLM)的智能体因搜索覆盖面有限、推理深度不足而难以兼顾:慢速串行查询限制了相关资源的覆盖,原始输入噪声破坏了多步推理的连贯性。为此,我们提出BrowseMaster,一个由程序化增强的规划-执行代理对构成的可扩展框架。规划器根据任务约束制定并动态调整搜索策略,执行器则进行高效、精准的检索,为规划器提供简洁、相关的证据。这种分工机制在保持长程推理连贯性的同时,实现广泛且系统的探索,突破了现有智能体面临的权衡困境。在具有挑战性的英/中文基准上的大量实验表明,BrowseMaster持续优于开源及商用基线,在BrowseComp-en上取得30.0分,在BrowseComp-zh上取得46.5分,展现了其在大规模复杂推理型信息检索任务中的强大能力。

原文摘要 · Abstract (English)

Effective information seeking in the vast and ever-growing digital landscape requires balancing expansive search with strategic reasoning. Current large language model (LLM)-based agents struggle to achieve this balance due to limitations in search breadth and reasoning depth, where slow, serial querying restricts coverage of relevant sources and noisy raw inputs disrupt the continuity of multi-step reasoning. To address these challenges, we propose BrowseMaster, a scalable framework built around a programmatically augmented planner-executor agent pair. The planner formulates and adapts search strategies based on task constraints, while the executor conducts efficient, targeted retrieval to supply the planner with concise, relevant evidence. This division of labor preserves coherent, long-horizon reasoning while sustaining broad and systematic exploration, overcoming the trade-off that limits existing agents. Extensive experiments on challenging English and Chinese benchmarks show that BrowseMaster consistently outperforms open-source and proprietary baselines, achieving scores of 30.0 on BrowseComp-en and 46.5 on BrowseComp-zh, which demonstrates its strong capability in complex, reasoning-heavy information-seeking tasks at scale.

信息检索智能代理程序化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。