arXiv:2509.05449cs.LGcs.AI2025-09Conference of the …被引 5

通过分析模型内部状态和注意力模式,发现训练数据泄露痕迹。

Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis

  • 从变换器隐藏状态与注意力分布中提取内部信号
  • 在多个模型上实现平均AUC 0.85的成员推断效果
  • 适合关注大模型隐私安全的研究者

成员推断攻击(MIAs)可判断特定数据是否用于训练机器学习模型,是隐私审计与合规评估的重要工具。近期研究指出,针对大型语言模型(LLMs)的MIAs表现仅略优于随机猜测,暗示现代大规模预训练可能无隐私泄露风险。本文提出不同视角:通过分析模型内部表示而非仅输出,可揭示潜在的成员推断信号。我们提出的框架memTrace利用“神经足迹”机制,从候选序列处理过程中的隐藏状态与注意力模式中提取信息。通过层间表征动态、注意力分布特征及跨层转换模式分析,检测传统基于损失的方法难以捕捉的记忆指纹。该方法在多个模型家族上取得平均AUC 0.85的优异性能,表明即使输出层面信号受保护,内部行为仍可能暴露训练数据,亟需进一步研究成员隐私与更鲁棒的隐私保护训练技术。

原文摘要 · Abstract (English)

Membership inference attacks (MIAs) reveal whether specific data was used to train machine learning models, serving as important tools for privacy auditing and compliance assessment. Recent studies have reported that MIAs perform only marginally better than random guessing against large language models, suggesting that modern pre-training approaches with massive datasets may be free from privacy leakage risks. Our work offers a complementary perspective to these findings by exploring how examining LLMs' internal representations, rather than just their outputs, may provide additional insights into potential membership inference signals. Our framework, \emph{memTrace}, follows what we call \enquote{neural breadcrumbs} extracting informative signals from transformer hidden states and attention patterns as they process candidate sequences. By analyzing layer-wise representation dynamics, attention distribution characteristics, and cross-layer transition patterns, we detect potential memorization fingerprints that traditional loss-based approaches may not capture. This approach yields strong membership detection across several model families achieving average AUC scores of 0.85 on popular MIA benchmarks. Our findings suggest that internal model behaviors can reveal aspects of training data exposure even when output-based signals appear protected, highlighting the need for further research into membership privacy and the development of more robust privacy-preserving training techniques for large language models.

隐私安全大模型成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。