提升网页智能体搜索效率,让大模型更聪明地找信息
WebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking
- 将信息搜索建模为树状推理,扩大可覆盖的目标实体
- 通过三种合成任务策略,使搜索效率和准确率同步提升
- 仅保留高效且正确的训练轨迹,优化模型表现
基于大语言模型的智能体在开放问题求解中展现出变革性潜力,其中信息搜索(IS)是实现自主推理与决策的核心能力。现有研究多聚焦于检索深度,但当前IS智能体普遍存在搜索效率低下的问题,根源在于训练任务中目标实体稀疏,限制了高效搜索行为的学习与泛化。为此,我们提出WebLeaper框架,通过构建高覆盖率的信息搜索任务并生成高效求解轨迹来解决上述问题。我们将信息搜索形式化为树状结构推理问题,使更多目标实体能在有限上下文内被嵌入。利用精心筛选的维基百科表格,设计出基础、联合与反向联合三种任务合成方式,系统性提升搜索效率与有效性。最终,仅保留同时具备准确性和高效性的训练轨迹,确保模型在正确性与搜索性能上双重优化。在五个信息搜索基准(BrowserComp、GAIA、xbench-DeepSearch、WideSearch、Seal-0)上的大量实验表明,该方法在基础与综合设置下均持续优于强基线,在有效性和效率上实现双重提升。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based agents have emerged as a transformative approach for open-ended problem solving, with information seeking (IS) being a core capability that enables autonomous reasoning and decision-making. While prior research has largely focused on improving retrieval depth, we observe that current IS agents often suffer from low search efficiency, which in turn constrains overall performance. A key factor underlying this inefficiency is the sparsity of target entities in training tasks, which limits opportunities for agents to learn and generalize efficient search behaviors. To address these challenges, we propose WebLeaper, a framework for constructing high-coverage IS tasks and generating efficient solution trajectories. We formulate IS as a tree-structured reasoning problem, enabling a substantially larger set of target entities to be embedded within a constrained context. Leveraging curated Wikipedia tables, we propose three variants for synthesizing IS tasks, Basic, Union, and Reverse-Union, to systematically increase both IS efficiency and efficacy. Finally, we curate training trajectories by retaining only those that are simultaneously accurate and efficient, ensuring that the model is optimized for both correctness and search performance. Extensive experiments on both basic and comprehensive settings, conducted on five IS benchmarks, BrowserComp, GAIA, xbench-DeepSearch, WideSearch, and Seal-0, demonstrate that our method consistently achieves improvements in both effectiveness and efficiency over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。