用合成数据评估模型记忆会误判,因攻击把合成文本当训练数据。
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
- 将合成数据作非成员数据,导致成员推理攻击误判
- 不同模型大小和架构下,合成文本均被误判为训练数据
- 提醒警惕合成数据在模型评估中的误导风险
近期研究显示,大语言模型的成员推理攻击(MIAs)结果不明确,部分源于难以创建无时间偏移的非成员数据集。研究人员转向合成数据作为替代,但我们发现该方法可能从根本上产生误导。实验表明,MIAs实际上充当机器生成文本检测器,无论数据来源如何,均错误将合成数据识别为训练样本。此现象在不同模型架构与规模(从开源模型到GPT-3.5等商用模型)中持续存在。即使由更大型模型生成的合成文本,也会被目标模型判定为训练数据。这揭示了严重隐患:使用合成数据进行成员评估可能导致对模型记忆与数据泄露的错误结论。我们警告,此类问题也可能影响其他依赖模型信号(如损失值)的评估,当合成或机器翻译数据替代真实样本时。
原文摘要 · Abstract (English)
Recent work shows membership inference attacks (MIAs) on large language models (LLMs) produce inconclusive results, partly due to difficulties in creating non-member datasets without temporal shifts. While researchers have turned to synthetic data as an alternative, we show this approach can be fundamentally misleading. Our experiments indicate that MIAs function as machine-generated text detectors, incorrectly identifying synthetic data as training samples regardless of the data source. This behavior persists across different model architectures and sizes, from open-source models to commercial ones such as GPT-3.5. Even synthetic text generated by different, potentially larger models is classified as training data by the target model. Our findings highlight a serious concern: using synthetic data in membership evaluations may lead to false conclusions about model memorization and data leakage. We caution that this issue could affect other evaluations using model signals such as loss where synthetic or machine-generated translated data substitutes for real-world samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。