arXiv:2510.03285cs.AIcs.CR2025-10被引 12

测试网页代理在真实网络下的可靠性,发现现有模型表现大幅下降

WAREX: Web Agent Reliability Evaluation on Existing Benchmarks

  • 在真实网络环境中引入多种干扰因素测试代理表现
  • 三类主流基准任务成功率普遍下降30%以上
  • 适合关注网页代理鲁棒性与实际应用落地的研究者

基于浏览器的大型语言模型代理在表单填写、酒店预订和在线购物等任务中展现出潜力。当前基准测试多在可控环境(如容器或稳定网络)中进行,网站行为具有确定性。然而在真实世界中,用户通过存在多重不稳定的网络和HTTPS连接访问网站,包括客户端、服务器端问题或系统级故障。此外,活网站易受跨站脚本攻击及内容修改影响,导致意外弹窗或功能异常。为填补这一空白,我们提出WAREX:在现有基准上评估网页代理的可靠性。我们在WebArena、WebVoyager和REAL三个主流基准上测试了WAREX的影响,实验表明引入该框架后任务成功率显著下降,揭示了当前顶尖代理在真实场景中的脆弱性。

原文摘要 · Abstract (English)

Recent advances in browser-based LLM agents have shown promise for automating tasks ranging from simple form filling to hotel booking or online shopping. Current benchmarks measure agent performance in controlled environments, such as containers or stable networks, where websites behave deterministically. However, in the real world, users access websites over networks and HTTPS connections that introduce instability from multiple sources: client-side, server-side issues or broader system failures. Moreover, live websites are prone to web attacks such Cross-Site Scripting, as well as general site modifications which can cause unexpected or malicious pop-ups or improper functionality. To address this gap, we present WAREX: Web Agent Reliability Evaluation on Existing Benchmarks. We measure the impact of WAREX across three popular benchmarks: WebArena, WebVoyager, and REAL. Our experiments show that introducing WAREX leads to significant drops in task success rates, highlighting the limited robustness of state-of-the-art agents.

网页代理可靠性评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。