新基准评估网页代理的安全与可信度,发现主流模型完成任务但常违规。
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- 构建222个企业级任务,每项配安全规则并多维度评分
- 提出新指标CuP,仅计遵守所有规则的完成任务
- 揭示主流代理安全达标率不足完成率的三分之二
自主网页代理能解决复杂浏览任务,但现有基准仅衡量任务完成度,忽视安全性与可信度。为使代理用于关键流程,安全与可信度(ST)是必要前提。我们提出 extbf{ extsc{ST-WebAgentBench}},一个可配置、易扩展的评测套件,用于评估真实企业场景中网页代理的ST表现。其222个任务均配有安全策略(政策),以简洁规则形式定义约束,并按六个正交维度(如用户同意、鲁棒性)评分。除原始任务成功率外,我们引入 extit{Completion Under Policy}(CuP)指标,仅奖励遵守所有适用政策的完成;提出 extit{Risk Ratio}量化各维度的违规风险。对三个开源先进代理的评估显示,其平均CuP低于名义完成率的三分之二,暴露出严重安全隐患。通过开放代码、评测模板及策略编写界面, extsc{ST-WebAgentBench}为大规模部署可信网页代理提供了可操作的第一步。
原文摘要 · Abstract (English)
Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprises can trust. To integrate these agents into critical workflows, safety and trustworthiness (ST) are prerequisite conditions for adoption. We introduce \textbf{\textsc{ST-WebAgentBench}}, a configurable and easily extensible suite for evaluating web agent ST across realistic enterprise scenarios. Each of its 222 tasks is paired with ST policies, concise rules that encode constraints, and is scored along six orthogonal dimensions (e.g., user consent, robustness). Beyond raw task success, we propose the \textit{Completion Under Policy} (\textit{CuP}) metric, which credits only completions that respect all applicable policies, and the \textit{Risk Ratio}, which quantifies ST breaches across dimensions. Evaluating three open state-of-the-art agents reveals that their average CuP is less than two-thirds of their nominal completion rate, exposing critical safety gaps. By releasing code, evaluation templates, and a policy-authoring interface, \href{https://sites.google.com/view/st-webagentbench/home}{\textsc{ST-WebAgentBench}} provides an actionable first step toward deploying trustworthy web agents at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。