arXiv:2410.07582cs.CLcs.AI2024-10Conference of the …被引 15

无需标注非成员数据,用迭代优化提升模型训练数据归属检测精度。

Detecting Training Data of Large Language Models via Expectation Maximization

  • 通过期望最大化策略迭代优化提示词效果和归属得分
  • 在分布可区分场景下性能超越现有基线方法
  • 适合关注模型隐私泄露风险的研究者使用

成员推理攻击(MIAs)旨在判断特定样本是否用于训练给定语言模型。尽管已有研究探索基于提示的攻击方法(如ReCALL),但这些方法严重依赖于使用已知非成员作为提示能有效抑制模型对非成员查询的响应这一假设。本文提出EM-MIA,一种无需标签非成员示例的新型成员推理方法,利用期望最大化策略迭代优化前缀有效性与成员得分。为支持可控评估,我们构建了OLMoMIA基准,可系统分析在不同分布重叠和难度下的MIA鲁棒性。在WikiMIA和OLMoMIA上的实验表明,EM-MIA在分布明显可区分的场景中优于现有基线。我们揭示了在部分分布重叠的实际场景中EM-MIA仍有效,而失败案例暴露了当前MIA方法在几乎相同条件下的根本局限。代码与评估流程已公开,以促进可复现且稳健的MIA研究。

原文摘要 · Abstract (English)

Membership inference attacks (MIAs) aim to determine whether a specific example was used to train a given language model. While prior work has explored prompt-based attacks such as ReCALL, these methods rely heavily on the assumption that using known non-members as prompts reliably suppresses the model's responses to non-member queries. We propose EM-MIA, a new membership inference approach that iteratively refines prefix effectiveness and membership scores using an expectation-maximization strategy without requiring labeled non-member examples. To support controlled evaluation, we introduce OLMoMIA, a benchmark that enables analysis of MIA robustness under systematically varied distributional overlap and difficulty. Experiments on WikiMIA and OLMoMIA show that EM-MIA outperforms existing baselines, particularly in settings with clear distributional separability. We highlight scenarios where EM-MIA succeeds in practical settings with partial distributional overlap, while failure cases expose fundamental limitations of current MIA methods under near-identical conditions. We release our code and evaluation pipeline to encourage reproducible and robust MIA research.

隐私安全成员推理语言模型机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。