arXiv:2506.06003cs.LGcs.CR2025-06NeurIPS被引 4

数据投毒可让成员推断测试失效,连语义相似数据也骗过检测。

What Really is a Member? Discrediting Membership Inference via Poisoning

  • 通过投毒训练集,让成员推断测试误判目标数据是否在训练集中
  • 攻击可使现有测试性能降至随机水平以下,准确率大幅下降
  • 揭示了准确性与抗投毒能力的内在矛盾,适合关注模型安全的研究者

成员推断测试旨在判断某个数据点是否曾出现在语言模型的训练集中。然而,近期研究发现,基于精确匹配的严格定义下,此类测试常会失败,并建议放宽定义,将语义相近的数据也视为成员。本文表明,即使在该放宽定义下,成员推断测试依然不可靠:可通过投毒训练数据,使测试对特定目标点产生错误判断。我们理论揭示了测试准确率与抗投毒鲁棒性之间的权衡关系,并提出一种具体投毒攻击实例,实证验证其有效性。结果表明,该攻击可使现有测试性能显著下降至随机水平以下。

原文摘要 · Abstract (English)

Membership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often fail under the strict definition of membership based on exact matching, and have suggested relaxing this definition to include semantic neighbors as members as well. In this work, we show that membership inference tests are still unreliable under this relaxation - it is possible to poison the training dataset in a way that causes the test to produce incorrect predictions for a target point. We theoretically reveal a trade-off between a test's accuracy and its robustness to poisoning. We also present a concrete instantiation of this poisoning attack and empirically validate its effectiveness. Our results show that it can degrade the performance of existing tests to well below random.

成员推断数据投毒模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。