用分布视角分析模型记忆,发现自影响会低估近似重复数据的风险
Counterfactual Influence as a Distributional Quantity
- 将反事实影响视为整体分布,考察所有训练样本对目标样本的影响
- 自影响值低的样本仍可能被近似提取,因近似重复样本会稀释自影响
- 该方法可有效识别图像和文本数据中的近似重复项,适合隐私风险研究
机器学习模型会记忆训练数据,引发隐私与泛化担忧。反事实自影响是常用指标,衡量样本是否因出现在训练集中而改变预测结果。但近期研究发现,记忆还受其他样本(尤其是近似重复)影响。本文将反事实影响视为分布量,分析小语言模型中所有训练样本间的完整影响分布。结果显示,仅看自影响会严重低估风险:近似重复样本的存在显著降低自影响,但这些样本仍可被近似提取。在图像分类任务中,通过分析影响分布也发现了CIFAR-10中的近似重复。结论表明,记忆源于训练数据间的复杂交互,应以全影响分布替代单一自影响来刻画。
原文摘要 · Abstract (English)
Machine learning models are known to memorize samples from their training data, raising concerns around privacy and generalization. Counterfactual self-influence is a popular metric to study memorization, quantifying how the model's prediction for a sample changes depending on the sample's inclusion in the training dataset. However, recent work has shown memorization to be affected by factors beyond self-influence, with other training samples, in particular (near-)duplicates, having a large impact. We here study memorization treating counterfactual influence as a distributional quantity, taking into account how all training samples influence how a sample is memorized. For a small language model, we compute the full influence distribution of training samples on each other and analyze its properties. We find that solely looking at self-influence can severely underestimate tangible risks associated with memorization: the presence of (near-)duplicates seriously reduces self-influence, while we find these samples to be (near-)extractable. We observe similar patterns for image classification, where simply looking at the influence distributions reveals the presence of near-duplicates in CIFAR-10. Our findings highlight that memorization stems from complex interactions across training data and is better captured by the full influence distribution than by self-influence alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。