构建压力测试基准,评估网页智能体在真实交互波动下的鲁棒性。
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability

- 设计可控的扰动机制模拟布局变化、语义差异等现实交互波动。
- 实测显示,主流多模态网页智能体在压力下失败率显著上升。
- 适合关注智能体实际部署可靠性的研究人员和开发者。
基于大语言模型的网页智能体在真实交互任务中表现强劲,但现有评估多在稳定理想环境下进行,可能高估其鲁棒性。为弥补这一不足,本文提出一个诊断式压力测试基准。首先构建干净稳定的网页环境作为基准参考;随后引入结构化扰动,模拟布局漂移、交互语义改变及执行中断等现实交互变异性。通过对比智能体在干净与扰动环境下的行为,本框架可系统诊断其在假设性场景下的鲁棒性。对前沿多模态网页智能体的广泛评估表明,压力测试揭示了隐藏的失效模式和显著的鲁棒性差距。
原文摘要 · Abstract (English)
Large language model-based web agents have demonstrated strong performance on realistic web interaction tasks. However, existing evaluations are predominantly conducted under relatively stable and well-behaved interaction conditions, which may overestimate agent robustness. High task success in such idealized settings does not necessarily reflect performance under realistic web interaction. To address this limitation, we introduce a diagnostic stress-testing benchmark for web agents. We first construct realistic and controllable web environments that provide clean and stable interaction workflows as reference baselines. We then introduce structured and controlled perturbations that emulate interaction variability, including shifting layouts, altered interaction semantics, and execution disruptions. By comparing agent behavior between clean and perturbed settings, our framework enables systematic diagnosis of robustness under what-if interaction scenarios. Through extensive evaluation of state-of-the-art multimodal web agents, we show that stress-based evaluation exposes failure modes and substantial robustness gaps that remain hidden under clean benchmark conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。