让智能体像人一样分层浏览网页,高效获取深层信息
Nested Browser-Use Learning for Agentic Information Seeking
- 设计分层浏览器操作框架,分离控制与探索
- 在深度信息检索任务中显著提升准确率和效率
- 适合需要复杂网页交互的智能搜索系统开发者
信息检索(IS)智能体在广泛且深入的搜索任务中已取得优异表现,但其工具使用仍主要局限于API级片段获取和基于URL的页面抓取,难以触及真实浏览所蕴含的丰富信息。尽管完整浏览器交互可解锁更深层次能力,但其细粒度控制和冗长页面内容带来的复杂性,给基于ReAct风格函数调用的智能体带来挑战。为此,我们提出嵌套浏览器使用学习(NestBrowse),引入一种最小且完备的浏览器动作框架,通过嵌套结构解耦交互控制与页面探索。该设计简化了智能体推理过程,同时实现有效的深网信息获取。在具有挑战性的深度信息检索基准上的实证结果表明,NestBrowse在实践中展现出明显优势。进一步的深入分析也证实了其高效性与灵活性。
原文摘要 · Abstract (English)
Information-seeking (IS) agents have achieved strong performance across a range of wide and deep search tasks, yet their tool use remains largely restricted to API-level snippet retrieval and URL-based page fetching, limiting access to the richer information available through real browsing. While full browser interaction could unlock deeper capabilities, its fine-grained control and verbose page content returns introduce substantial complexity for ReAct-style function-calling agents. To bridge this gap, we propose Nested Browser-Use Learning (NestBrowse), which introduces a minimal and complete browser-action framework that decouples interaction control from page exploration through a nested structure. This design simplifies agentic reasoning while enabling effective deep-web information acquisition. Empirical results on challenging deep IS benchmarks demonstrate that NestBrowse offers clear benefits in practice. Further in-depth analyses underscore its efficiency and flexibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。