arXiv:2512.13352cs.LGcs.CL2025-12中稿 · publication at the…

用会员推理攻击提升大模型数据提取效率,验证其实际威胁。

On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models

  • 将多种会员推理方法整合进数据提取流程,系统评估效果。
  • 实测显示部分攻击在真实场景下可有效识别训练数据。
  • 为隐私风险研究提供可复现的评估框架,适合安全与隐私方向读者。

大型语言模型容易记忆训练数据,带来严重隐私风险。主要担忧包括训练数据提取和会员推理攻击(MIAs)。先前研究指出这两类威胁相互关联:攻击者可通过大量查询模型生成文本,再利用会员推理攻击验证特定数据点是否在训练集中。本研究将多种会员推理技术集成至数据提取流程,系统性地基准测试其有效性,并将其在集成环境下的表现与传统会员推理基准结果对比,从而评估其在真实世界数据提取场景中的实际应用价值。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are prone to memorizing training data, which poses serious privacy risks. Two of the most prominent concerns are training data extraction and Membership Inference Attacks (MIAs). Prior research has shown that these threats are interconnected: adversaries can extract training data from an LLM by querying the model to generate a large volume of text and subsequently applying MIAs to verify whether a particular data point was included in the training set. In this study, we integrate multiple MIA techniques into the data extraction pipeline to systematically benchmark their effectiveness. We then compare their performance in this integrated setting against results from conventional MIA benchmarks, allowing us to evaluate their practical utility in real-world extraction scenarios.

隐私安全会员推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。