arXiv:2412.17249cs.ROcs.CR2024-12中稿 · ICASSP 2025 Main被引 7

融合多种攻击方法,提升大模型成员推理攻击精度

EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling

  • 用XGBoost集成LOSS、参考基、min-k等攻击方法
  • 在多个大模型和数据集上显著提高AUC-ROC与准确率
  • 适合研究模型隐私风险与防御机制的学者参考

随着大语言模型(LLM)的广泛应用,其训练数据的隐私泄露问题日益受到关注。成员推理攻击(MIAs)已成为评估此类模型隐私风险的关键工具。尽管现有方法如LOSS、基于参考、min-k和zlib在特定场景下表现良好,但在大规模预训练语言模型上,尤其在单轮训练和大规模数据集下,其效果常接近随机猜测。为此,本文提出一种新型集成攻击方法EM-MIAs,将LOSS、参考基、min-k和zlib等多种现有MIAs技术整合到基于XGBoost的模型中,以提升整体攻击性能。实验结果表明,该集成模型在多个大语言模型和数据集上均显著优于单一攻击方法,大幅提升了AUC-ROC和准确率。这说明通过融合不同方法的优势,可更有效地识别模型训练数据中的成员,为评估大模型隐私风险提供了更强大的工具。本研究为大模型隐私保护领域提供了新方向,并强调了开发更强隐私审计方法的必要性。

原文摘要 · Abstract (English)

With the widespread application of large language models (LLM), concerns about the privacy leakage of model training data have increasingly become a focus. Membership Inference Attacks (MIAs) have emerged as a critical tool for evaluating the privacy risks associated with these models. Although existing attack methods, such as LOSS, Reference-based, min-k, and zlib, perform well in certain scenarios, their effectiveness on large pre-trained language models often approaches random guessing, particularly in the context of large-scale datasets and single-epoch training. To address this issue, this paper proposes a novel ensemble attack method that integrates several existing MIAs techniques (LOSS, Reference-based, min-k, zlib) into an XGBoost-based model to enhance overall attack performance (EM-MIAs). Experimental results demonstrate that the ensemble model significantly improves both AUC-ROC and accuracy compared to individual attack methods across various large language models and datasets. This indicates that by combining the strengths of different methods, we can more effectively identify members of the model's training data, thereby providing a more robust tool for evaluating the privacy risks of LLM. This study offers new directions for further research in the field of LLM privacy protection and underscores the necessity of developing more powerful privacy auditing methods.

成员推理攻击大模型隐私集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。