arXiv:2608.26204cs.CRcs.AI2026-08

评测智能助手在跨设备操作中的安全与理解能力,发现所有模型都存在严重漏洞。

ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices

论文配图:ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
图 1 · 摘自论文原文
  • 构建双流评测基准,分别测试安全防护与歧义理解能力。
  • 所有模型在高风险操作上成功率超80%,但攻击成功率也达30%以上。
  • 揭示三类安全架构,提示工具依赖性对安全性的关键影响。

计算机使用代理(CUAs)正被广泛用于代表用户在移动和桌面应用中导航,但尚无全面基准评估其在处理模糊指令时是否能安全地与视觉界面交互。我们提出ADeptS-Bench,一个基于ADEPTS能力框架和普通用户研究的双流可信度评测基准。安全流包含成对的良性/恶意任务,威胁嵌入于视觉界面中;歧义流评估代理在意图不明确时是否主动寻求澄清。评估七个模型发现:无一模型能在任务成功率超过80%的同时,将攻击成功率控制在30%以下;所有模型均毫不犹豫点击“结账”完成25,000美元订单,且无一识别出“出厂重置”按钮被错误标记为“优化”。消融实验揭示三种不同安全架构:工具依赖型(使用拒绝工具后+21-23个百分点)、部分工具依赖型(+10-11个百分点),以及无机制型(无变化)。在歧义处理中,所有模型均高估后果严重性,表现出与安全领域相同的过度拒绝偏差。论文发布全部数据、评估代码与分析工具。

原文摘要 · Abstract (English)

Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can safely interact with visual interfaces while handling ambiguous instructions. We introduce ADeptS-Bench, a dual-stream trustworthiness benchmark, grounded in the ADEPTS capability framework and general population user studies. The Safety stream provides paired benign/malicious tasks with threats embedded in the visual interface. The Disambiguation stream evaluates whether agents seek clarification when intent is ambiguous. Evaluating seven models reveals that no model consistently exceeds 80% task success while staying below 30% attack success; every model clicks "Checkout" on a $25K order without hesitation, and none detects that a "factory reset" button is mislabeled as "Optimize." An ablation reveals three distinct safety architectures: tool-dependent (ASR +21-23pp without refusal tool), partially tool-dependent (+10-11pp), and no mechanism (unchanged). In disambiguation, all models overestimate consequence severity, mirroring the over-refusal bias observed in safety. We release all data, evaluation code, and analysis tools upon publication.

智能代理安全评测人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。