arXiv:2604.24348cs.CL2026-04

OS-SPEAR为操作系统代理提供安全、性能、效率与鲁棒性的全面评估工具。

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

论文配图:OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents
图 1 · 摘自论文原文
  • 构建四维评估体系,涵盖安全、性能、效率与跨模态鲁棒性
  • 22个主流代理测试显示效率与安全存在显著权衡
  • 适合开发可信高效操作代理的研究者与工程师使用

多模态大模型的演进使研究重点从文本生成转向主动行为执行,尤其体现在通过操作系统代理在复杂图形界面中导航。然而,这些代理成为可信赖日常伙伴的过程中,受限于对安全性、效率和多模态鲁棒性缺乏严谨评估。现有基准存在安全场景狭窄、轨迹标注噪声大、鲁棒性度量有限等问题。为此,我们提出OS-SPEAR,一个系统化分析操作系统代理在安全、性能、效率和鲁棒性四个维度上的综合工具包。OS-SPEAR包含四个专项子集:(1) 安全子集,覆盖环境与人为引发的多样化危害;(2) 性能子集,基于轨迹价值估计与分层采样构建;(3) 效率子集,从时间延迟与令牌消耗双重角度量化表现;(4) 鲁棒性子集,对视觉与文本输入施加跨模态干扰。此外,提供自动化分析工具生成可读诊断报告。我们对22个主流操作系统代理进行了广泛评估,结果揭示关键洞察:效率与安全/鲁棒性之间普遍存在权衡,专用代理在性能上优于通用模型,且不同模态存在差异化的鲁棒性脆弱点。通过提供多维排名与标准化框架,OS-SPEAR为下一代可靠高效操作系统代理的发展奠定基础。数据集与代码已开源:https://github.com/Wuzheng02/OS-SPEAR。

原文摘要 · Abstract (English)

The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS agents navigating complex GUIs. However, the transition of these agents into trustworthy daily partners is hindered by a lack of rigorous evaluation regarding safety, efficiency, and multi-modal robustness. Current benchmarks suffer from narrow safety scenarios, noisy trajectory labeling, and limited robustness metrics. To bridge this gap, we propose OS-SPEAR, a comprehensive toolkit for the systematic analysis of OS agents across four dimensions: Safety, Performance, Efficiency, and Robustness. OS-SPEAR introduces four specialized subsets: (1) a S(afety)-subset encompassing diverse environment- and human-induced hazards; (2) a P(erformance)-subset curated via trajectory value estimation and stratified sampling; (3) an E(fficiency)-subset quantifying performance through the dual lenses of temporal latency and token consumption; and (4) a R(obustness)-subset that applies cross-modal disturbances to both visual and textual inputs. Additionally, we provide an automated analysis tool to generate human-readable diagnostic reports. We conduct an extensive evaluation of 22 popular OS agents using OS-SPEAR. Our empirical results reveal critical insights into the current landscape: notably, a prevalent trade-off between efficiency and safety or robustness, the performance superiority of specialized agents over general-purpose models, and varying robustness vulnerabilities across different modalities. By providing a multidimensional ranking and a standardized evaluation framework, OS-SPEAR offers a foundational resource for developing the next generation of reliable and efficient OS agents. The dataset and codes are available at https://github.com/Wuzheng02/OS-SPEAR.

操作系统代理评估工具多模态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。