arXiv:2602.15927cs.CVcs.LG2026-02被引 1

攻击者用伪造图像诱导大模型在对话中输出指定内容。

Visual Memory Injection Attacks for Multi-Turn Conversations

  • 通过隐蔽图像注入记忆,触发后强制输出预设信息。
  • 多轮对话下仍能生效,攻击成功率超90%。
  • 适合关注视觉安全与对抗样本的研究者。

生成式大视觉语言模型(LVLMs)近期性能大幅提升,用户数量迅速增长。然而,其在长上下文多轮对话场景下的安全性尚未充分研究。本文考虑真实场景:攻击者上传经篡改的图像至网络/社交媒体,正常用户下载并将其作为输入使用LVLM。我们提出新型隐蔽的视觉记忆注入(VMI)攻击:在常规提示下模型表现正常,但一旦用户发出触发性提示,模型便会输出特定预设目标信息,用于恶意营销或政治煽动。相比以往单轮攻击,VMI在长时间多轮对话后依然有效。我们在多个近期开源权重的LVLM上验证了该攻击,表明通过被污染图像在多轮对话中大规模操纵用户是可行的,亟需提升LVLM对此类攻击的鲁棒性。源代码已公开于https://github.com/chs20/visual-memory-injection。

原文摘要 · Abstract (English)

Generative large vision-language models (LVLMs) have recently achieved impressive performance gains, and their user base is growing rapidly. However, the security of LVLMs, in particular in a long-context multi-turn setting, is largely underexplored. In this paper, we consider the realistic scenario in which an attacker uploads a manipulated image to the web/social media. A benign user downloads this image and uses it as input to the LVLM. Our novel stealthy Visual Memory Injection (VMI) attack is designed such that on normal prompts the LVLM exhibits nominal behavior, but once the user gives a triggering prompt, the LVLM outputs a specific prescribed target message to manipulate the user, e.g. for adversarial marketing or political persuasion. Compared to previous work that focused on single-turn attacks, VMI is effective even after a long multi-turn conversation with the user. We demonstrate our attack on several recent open-weight LVLMs. This article thereby shows that large-scale manipulation of users is feasible with perturbed images in multi-turn conversation settings, calling for better robustness of LVLMs against these attacks. We release the source code at https://github.com/chs20/visual-memory-injection

视觉安全对抗攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。