自动化构建可复现、真实且可扩展的浏览器智能体评测基准
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark

- 四智能体流水线自动生成交互式网页环境,无需人工标注
- 构建934个任务的基准,覆盖7个领域和3级难度
- 多维难度控制可揭示模型真实能力差异,超越单一得分
现有浏览器智能体评测面临现实性、可复现性与可扩展性的三难困境:真实网站因内容漂移导致不可复现,可控环境又因缺乏真实网络噪声而失真,且两者均需高昂的人工维护。本文提出WebForge,首个完全自动化的框架,通过计划、生成、精炼、验证四智能体流水线,端到端生成无须人工标注的交互式、自包含网页环境。基于七维难度控制体系(导航深度、视觉复杂度、推理难度等),实现任务设计的系统化,支持超越单一综合评分的能力画像。利用WebForge构建了包含934个任务的WebForge-Bench基准,覆盖7个领域和3个难度层级。多模型实验表明,难度分层有效区分模型能力,跨领域分析揭示了聚合指标无法捕捉的能力偏差。结果证实,多维度评估能揭示单个综合分数无法反映的模型真实能力特征。代码与基准已公开于https://github.com/yuandaxia2001/WebForge。
原文摘要 · Abstract (English)
Existing browser agent benchmarks face a fundamental trilemma: real-website benchmarks lack reproducibility due to content drift, controlled environments sacrifice realism by omitting real-web noise, and both require costly manual curation that limits scalability. We present WebForge, the first fully automated framework that resolves this trilemma through a four-agent pipeline -- Plan, Generate, Refine, and Validate -- that produces interactive, self-contained web environments end-to-end without human annotation. A seven-dimensional difficulty control framework structures task design along navigation depth, visual complexity, reasoning difficulty, and more, enabling systematic capability profiling beyond single aggregate scores. Using WebForge, we construct WebForge-Bench, a benchmark of 934 tasks spanning 7 domains and 3 difficulty levels. Multi-model experiments show that difficulty stratification effectively differentiates model capabilities, while cross-domain analysis exposes capability biases invisible to aggregate metrics. Together, these results confirm that multi-dimensional evaluation reveals distinct capability profiles that a single aggregate score cannot capture. Code and benchmark are publicly available at https://github.com/yuandaxia2001/WebForge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。