arXiv:2503.10809cs.CRcs.LG2025-03NeurIPS被引 11

恶意图像补丁可诱骗多模态系统代理执行危险操作。

MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents

  • 用对抗性扰动的图像区域欺骗系统代理
  • 攻击可在不同提示和界面下跨场景生效
  • 适合关注AI安全与系统防护的研究者

近期操作系统(OS)代理的发展使视觉-语言模型(VLMs)能直接控制用户计算机。与传统被动输出文本的VLM不同,OS代理能根据单一用户指令自主执行基于屏幕的计算机任务,通过捕获、解析和分析截图,并调用应用程序接口(API)如鼠标点击和键盘输入来执行低层操作。这种与操作系统直接交互显著提升了风险,因故障或操控可能带来即时且真实的后果。本文揭示了一种新型攻击向量:恶意图像补丁(MIPs),即经过对抗性扰动的屏幕区域,当被OS代理捕获时,会利用特定API诱导其执行有害行为。例如,嵌入桌面壁纸或社交平台中的MIP可导致代理窃取敏感用户数据。我们证明,MIPs在不同用户指令和屏幕配置间具有泛化能力,甚至可在执行良性指令时劫持多个OS代理。这些发现暴露了OS代理中亟需解决的关键安全漏洞,必须在广泛部署前加以重视。

原文摘要 · Abstract (English)

Recent advances in operating system (OS) agents have enabled vision-language models (VLMs) to directly control a user's computer. Unlike conventional VLMs that passively output text, OS agents autonomously perform computer-based tasks in response to a single user prompt. OS agents do so by capturing, parsing, and analysing screenshots and executing low-level actions via application programming interfaces (APIs), such as mouse clicks and keyboard inputs. This direct interaction with the OS significantly raises the stakes, as failures or manipulations can have immediate and tangible consequences. In this work, we uncover a novel attack vector against these OS agents: Malicious Image Patches (MIPs), adversarially perturbed screen regions that, when captured by an OS agent, induce it to perform harmful actions by exploiting specific APIs. For instance, a MIP can be embedded in a desktop wallpaper or shared on social media to cause an OS agent to exfiltrate sensitive user data. We show that MIPs generalise across user prompts and screen configurations, and that they can hijack multiple OS agents even during the execution of benign instructions. These findings expose critical security vulnerabilities in OS agents that have to be carefully addressed before their widespread deployment.

AI安全对抗攻击系统代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。