arXiv:2604.09024cs.CVcs.AI2026-04ACL被引 1

用微小干扰让大模型拒绝分析图片,保护隐私

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

  • 在图片中嵌入几乎看不见的干扰,诱导模型拒绝分析
  • 六种大模型测试均有效,触发拒绝回复率超90%
  • 适合担心图像被滥用的普通用户和隐私敏感场景

多模态大语言模型(MLLMs)虽能高效分析海量图像,但也带来隐私泄露风险。本文提出ImageProtector,一种用户侧防护方法:在分享前向图像注入精心设计的、几乎不可见的扰动,作为视觉提示攻击。当攻击者使用开放权重的MLLM分析受保护图像时,模型会持续生成拒绝回应,如“我无法协助此请求”。实验验证了该方法在六种MLLM和四个数据集上的有效性。同时评估了高斯噪声、DiffPure和对抗训练三种防御手段,发现它们虽部分削弱攻击效果,但会显著降低模型精度或效率。研究聚焦于开放权重模型与大规模自动化图像分析的实际场景,揭示了基于扰动的隐私保护的潜力与局限。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but also raising critical safety and societal concerns. In particular, open-weight MLLMs may be misused to extract sensitive information from personal images at scale, such as identities, locations, or other private details. In this work, we propose ImageProtector, a user-side method that proactively protects images before sharing by embedding a carefully crafted, nearly imperceptible perturbation that acts as a visual prompt injection attack on MLLMs. As a result, when an adversary analyzes a protected image with an MLLM, the MLLM is consistently induced to generate a refusal response such as "I'm sorry, I can't help with that request." We empirically demonstrate the effectiveness of ImageProtector across six MLLMs and four datasets. Additionally, we evaluate three potential countermeasures, Gaussian noise, DiffPure, and adversarial training, and show that while they partially mitigate the impact of ImageProtector, they simultaneously degrade model accuracy and/or efficiency. Our study focuses on the practically important setting of open-weight MLLMs and large-scale automated image analysis, and highlights both the promise and the limitations of perturbation-based privacy protection.

隐私保护多模态对抗攻击图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。