评测计算机代理在长期任务中识别安全风险的能力,涵盖正常与恶意场景。
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
- 构建多智能体自动化数据生成管道,覆盖65种场景
- 发现现有代理在长期规划中普遍存在安全意识不足问题
- 适合研究智能体安全、长程规划或对抗防御的学者使用
能与真实计算机系统交互的计算机使用代理(CUAs)可执行自动化任务,但面临重大安全风险。模糊指令可能引发有害行为,而攻击者可通过操纵工具执行达成恶意目的。现有基准多聚焦短周期或基于GUI的任务,仅评估执行时错误,忽视规划阶段的风险预判能力。为此,我们提出LPS-Bench,一个评估基于MCP的CUA在长期任务中规划期安全意识的基准,涵盖7个任务领域、9类风险及65种情景,包含良性与对抗性交互。我们设计了多智能体自动化数据生成流程,并采用大模型作为裁判的评估协议,通过规划轨迹判断安全意识。实验揭示现有CUAs存在显著安全缺陷。我们进一步分析风险并提出缓解策略,以提升MCP驱动的CUA系统在长期规划中的安全性。代码已开源:https://github.com/tychenn/LPS-Bench。
原文摘要 · Abstract (English)
Computer-use agents (CUAs) that interact with real computer systems can perform automated tasks but face critical safety risks. Ambiguous instructions may trigger harmful actions, and adversarial users can manipulate tool execution to achieve malicious goals. Existing benchmarks mostly focus on short-horizon or GUI-based tasks, evaluating on execution-time errors but overlooking the ability to anticipate planning-time risks. To fill this gap, we present LPS-Bench, a benchmark that evaluates the planning-time safety awareness of MCP-based CUAs under long-horizon tasks, covering both benign and adversarial interactions across 65 scenarios of 7 task domains and 9 risk types. We introduce a multi-agent automated pipeline for scalable data generation and adopt an LLM-as-a-judge evaluation protocol to assess safety awareness through the planning trajectory. Experiments reveal substantial deficiencies in existing CUAs' ability to maintain safe behavior. We further analyze the risks and propose mitigation strategies to improve long-horizon planning safety in MCP-based CUA systems. We open-source our code at https://github.com/tychenn/LPS-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。