arXiv:2503.19338cs.LGcs.CR2025-03综述被引 25

首次系统分析大模型隐私漏洞,揭示数据是否被训练的攻击方法

Membership Inference Attacks on Large-Scale Models: A Survey

  • 按模型类型和训练阶段分类,梳理大模型成员推理攻击策略
  • 发现预训练、微调、对齐等阶段均存在隐私泄露风险
  • 适合关注大模型安全与隐私的研究者和开发者参考

随着大规模模型如大语言模型(LLMs)和大多模态模型(LMMs)的广泛应用,其隐私风险仍缺乏深入研究。成员推理攻击(MIAs)能判断某数据点是否曾用于训练目标模型,是评估隐私风险的重要手段,在多种机器学习算法中已被证实有效。然而,尽管已有大量关于经典模型的MIAs研究,针对大模型的系统性综述仍属空白。为此,本文首次全面回顾了针对LLMs和LMMs的MIAs,从模型类型、攻击者知识和策略三个维度进行分析。不同于以往综述,本文进一步考察了攻击在模型流水线多个阶段的表现,包括预训练、微调、对齐及检索增强生成(RAG)。最后,识别出当前开放挑战,并提出未来强化大模型隐私鲁棒性的研究方向。

原文摘要 · Abstract (English)

As large-scale models such as Large Language Models (LLMs) and Large Multimodal Models (LMMs) see increasing deployment, their privacy risks remain underexplored. Membership Inference Attacks (MIAs), which reveal whether a data point was used in training the target model, are an important technique for exposing or assessing privacy risks and have been shown to be effective across diverse machine learning algorithms. However, despite extensive studies on MIAs in classic models, there remains a lack of systematic surveys addressing their effectiveness and limitations in large-scale models. To address this gap, we provide the first comprehensive review of MIAs targeting LLMs and LMMs, analyzing attacks by model type, adversarial knowledge, and strategy. Unlike prior surveys, we further examine MIAs across multiple stages of the model pipeline, including pre-training, fine-tuning, alignment, and Retrieval-Augmented Generation (RAG). Finally, we identify open challenges and propose future research directions for strengthening privacy resilience in large-scale models.

隐私安全大模型成员推理攻击分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。