大规模语言模型成员推断攻击效果受数据分布影响,需多角度验证其可靠性。
A Statistical and Multi-Perspective Revisiting of the Membership Inference Attack in Large Language Models
- 通过上千次实验从多个角度重检成员推断攻击方法
- 攻击成功率随模型规模提升,但整体仍较低,存在明显异常样本
- 阈值设定和解码动态差异是被忽视的关键挑战,适合隐私安全研究者
大型语言模型(LLMs)的数据透明性缺失凸显了成员推断攻击(MIA)的重要性,该攻击用于区分训练过的(成员)与未训练过的(非成员)数据。尽管以往研究显示其有效,但近期工作在不同设置下报告接近随机的表现,表明性能存在显著不一致。我们假设单一设置无法代表庞大语料库的分布,导致成员与非成员样本采样分布不同,引发不一致。本研究不再依赖单一设置,而是对每种MIA方法进行数千次实验,从文本特征、嵌入表示、阈值决策到解码动态等多个维度进行统计重检。发现:(1) MIA性能随模型规模提升且因领域而异,但多数方法未显著优于基线;(2) 尽管整体表现低,仍存在大量可区分的成员与非成员异常样本,且不同方法间差异明显;(3) 阈值设定是被忽视的关键挑战;(4) 文本差异性大和长文本更有利于提高攻击性能;(5) 是否可区分体现在模型嵌入中;(6) 成员与非成员表现出不同的解码动态。
原文摘要 · Abstract (English)
The lack of data transparency in Large Language Models (LLMs) has highlighted the importance of Membership Inference Attack (MIA), which differentiates trained (member) and untrained (non-member) data. Though it shows success in previous studies, recent research reported a near-random performance in different settings, highlighting a significant performance inconsistency. We assume that a single setting doesn't represent the distribution of the vast corpora, causing members and non-members with different distributions to be sampled and causing inconsistency. In this study, instead of a single setting, we statistically revisit MIA methods from various settings with thousands of experiments for each MIA method, along with study in text feature, embedding, threshold decision, and decoding dynamics of members and non-members. We found that (1) MIA performance improves with model size and varies with domains, while most methods do not statistically outperform baselines, (2) Though MIA performance is generally low, a notable amount of differentiable member and non-member outliers exists and vary across MIA methods, (3) Deciding a threshold to separate members and non-members is an overlooked challenge, (4) Text dissimilarity and long text benefit MIA performance, (5) Differentiable or not is reflected in the LLM embedding, (6) Member and non-members show different decoding dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。