arXiv:2512.03121cs.CRcs.AI2025-12中稿 · ESANN 2026

首次评估文本攻击在多模态模型中的效果,发现视觉信息会隐藏数据泄露信号。

Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models

  • 用对数概率法测试多模态模型的数据泄露风险
  • 分布内场景下攻击效果相似,视觉输入略优
  • 分布外时视觉输入干扰攻击,提升隐私保护

大型多模态语言模型(MLLMs)正成为众多应用的核心工具,因此理解其训练数据泄露问题至关重要。基于对数概率的成员推理攻击(MIAs)已被广泛用于评估大语言模型(LLMs)中的数据暴露风险,但在多模态模型中的有效性尚不明确。本文首次系统评估了将文本型MIA方法扩展到多模态场景的效果。我们在DeepSeek-VL与InternVL模型族上,在视觉-文本(V+T)和纯文本(T-only)条件下进行实验,结果表明:在分布内设置中,各类配置下的对数激活值(logit-based)MIA表现相近,略有视觉-文本优势;而在分布外设置中,视觉输入起到正则化作用,有效掩盖了成员身份信号。

原文摘要 · Abstract (English)

Large Multimodal Language Models (MLLMs) are emerging as one of the foundational tools in an expanding range of applications. Consequently, understanding training-data leakage in these systems is increasingly critical. Log-probability-based membership inference attacks (MIAs) have become a widely adopted approach for assessing data exposure in large language models (LLMs), yet their effect in MLLMs remains unclear. We present the first comprehensive evaluation of extending these text-based MIA methods to multimodal settings. Our experiments under vision-and-text (V+T) and text-only (T-only) conditions across the DeepSeek-VL and InternVL model families show that in in-distribution settings, logit-based MIAs perform comparably across configurations, with a slight V+T advantage. Conversely, in out-of-distribution settings, visual inputs act as regularizers, effectively masking membership signals.

多模态模型成员推理攻击隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。