通过注意力操控诱导电脑代理选择特定商品,实现隐蔽攻击
Preference Redirection via Attention Concentration: An Attack on Computer Use Agents
- 利用隐蔽对抗补丁引导模型注意力,改变内部偏好决策
- 在在线购物平台成功使代理选择指定目标商品,攻击成功率高
- 可迁移至微调版本模型,对基于开源模型的代理构成重大威胁
多模态大模型的发展催生了能够自主与图形用户界面交互的电脑使用代理(CUAs)。由于CUAs不局限于特定工具,可执行更复杂的智能任务,但也引入了新的安全漏洞。现有研究主要关注语言模态的脆弱性,而视觉模态的威胁尚未得到足够重视。本文提出PRAC攻击,不同于直接攻击视觉语言模型输出的方法,该攻击通过将模型注意力引向隐蔽的对抗性补丁,操纵其内部偏好。实验表明,PRAC能有效操控在线购物平台上CUA的商品选择过程,使其偏向指定目标产品。虽然攻击需白盒访问以生成对抗样本,但其在同模型的微调版本上仍具泛化能力,凸显了多个企业基于开源权重模型构建专用代理时所面临的严重风险。
原文摘要 · Abstract (English)
Advancements in multimodal foundation models have enabled the development of Computer Use Agents (CUAs) capable of autonomously interacting with GUI environments. As CUAs are not restricted to certain tools, they allow to automate more complex agentic tasks but at the same time open up new security vulnerabilities. While prior work has concentrated on the language modality, the vulnerability of the vision modality has received less attention. In this paper, we introduce PRAC, a novel attack that, unlike prior work targeting the VLM output directly, manipulates the model's internal preferences by redirecting its attention toward a stealthy adversarial patch. We show that PRAC is able to manipulate the selection process of a CUA on an online shopping platform towards a chosen target product. While we require white-box access to the model for the creation of the attack, we show that our attack generalizes to fine-tuned versions of the same model, presenting a critical threat as multiple companies build specific CUAs based on open weights models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。