神经启发的多模态模型能更好抵御隐私泄露攻击。
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- 引入拓扑正则化框架,让多模态模型更接近生物神经机制。
- 在COCO数据集上,隐私攻击成功率降低24%(ROC-AUC)。
- 提升隐私防护能力,同时保持生成文本质量不下降。
在智能代理兴起的时代,多模态模型(MMs)的广泛应用带来了敏感训练数据泄露的新风险。本文研究了针对多模态视觉-语言模型(VLMs)的黑箱隐私攻击——成员推断攻击(MIA)。现有研究主要聚焦于单模态模型,而近期表明多模态模型也面临隐私威胁。尽管生物启发的神经表征可增强单模态模型对抗对抗攻击的能力,但其对隐私攻击的鲁棒性仍未知。本文提出一种受神经科学启发的拓扑正则化(tau)框架,评估三种VLMs(BLIP、PaliGemma 2、ViT-GPT2)在三个基准数据集(COCO、CC3M、NoCaps)上的隐私韧性。对比基线与神经型VLM(tau > 0),在COCO数据集上,神经型BLIP模型的MIA攻击成功度(平均ROC-AUC)下降24%,同时在MPNet和ROUGE-2指标上保持相近的生成质量。在另外两个模型与数据集上的验证进一步确认了结果的一致性。本研究揭示了多模态模型的隐私风险,并为神经型模型的隐私韧性提供了实证支持。
原文摘要 · Abstract (English)
In the age of agentic AI, the growing deployment of multi-modal models (MMs) has introduced new attack vectors that can leak sensitive training data in MMs, causing privacy leakage. This paper investigates a black-box privacy attack, i.e., membership inference attack (MIA) on multi-modal vision-language models (VLMs). State-of-the-art research analyzes privacy attacks primarily to unimodal AI-ML systems, while recent studies indicate MMs can also be vulnerable to privacy attacks. While researchers have demonstrated that biologically inspired neural network representations can improve unimodal model resilience against adversarial attacks, it remains unexplored whether neuro-inspired MMs are resilient against privacy attacks. In this work, we introduce a systematic neuroscience-inspired topological regularization (tau) framework to analyze MM VLMs resilience against image-text-based inference privacy attacks. We examine this phenomenon using three VLMs: BLIP, PaliGemma 2, and ViT-GPT2, across three benchmark datasets: COCO, CC3M, and NoCaps. Our experiments compare the resilience of baseline and neuro VLMs (with topological regularization), where the tau > 0 configuration defines the NEURO variant of VLM. Our results on the BLIP model using the COCO dataset illustrate that MIA attack success in NEURO VLMs drops by 24% mean ROC-AUC, while achieving similar model utility (similarities between generated and reference captions) in terms of MPNet and ROUGE-2 metrics. This shows neuro VLMs are comparatively more resilient against privacy attacks, while not significantly compromising model utility. Our extensive evaluation with PaliGemma 2 and ViT-GPT2 models, on two additional datasets: CC3M and NoCaps, further validates the consistency of the findings. This work contributes to the growing understanding of privacy risks in MMs and provides evidence on neuro VLMs privacy threat resilience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。