大规模实验重评网页代理中的工具使用效果,揭示其真实收益与设计原则。
The Tool Illusion: Rethinking Tool Use in Web Agents
- 在多种模型、框架和基准上系统测试工具使用效果。
- 发现工具并非总带来提升,且存在性能波动与副作用风险。
- 为工具设计提供可复用的实证依据,适合研究者参考。
随着网页代理的快速发展,越来越多研究超越传统的原子浏览器操作,探索将工具使用作为更高层次的动作范式。尽管先前研究展示了工具的潜力,但其结论往往基于有限的实验规模或不可比的设置。因此,一些基本问题仍不明确:工具是否能为网页代理带来持续收益?有效工具的设计原则是什么?工具使用可能引入哪些副作用?为建立更坚实的实证基础,本文通过跨多样工具源、主干模型、工具使用框架和评估基准的广泛而严谨的实验,重新审视网页代理中的工具使用。研究结果既修正了部分既有结论,也以更广泛的证据补充了其他发现。我们希望本研究能为未来工具使用型网页代理的研究提供更可靠的实证支撑。
原文摘要 · Abstract (English)
As web agents rapidly evolve, an increasing body of work has moved beyond conventional atomic browser interactions and explored tool use as a higher-level action paradigm. Although prior studies have shown the promise of tools, their conclusions are often drawn from limited experimental scales and sometimes non-comparable settings. As a result, several fundamental questions remain unclear: i) whether tools provide consistent gains for web agents, ii) what practical design principles characterize effective tools, and iii) what side effects tool use may introduce. To establish a stronger empirical foundation for future research, we revisit tool use in web agents through an extensive and carefully controlled study across diverse tool sources, backbone models, tool-use frameworks, and evaluation benchmarks. Our findings both revise some prior conclusions and complement others with broader evidence. We hope this study provides a more reliable empirical basis and inspires future research on tool-use web agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。