arXiv:2605.26133cs.CLcs.AI2026-05中稿 · NLDB 2025综述被引 2

首次系统梳理大模型训练数据泄露风险,涵盖隐私与数据污染问题。

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications

  • 从数据泄露角度统一分析成员推断与数据污染
  • 梳理攻击与防御方法,揭示现有模型安全短板
  • 适合关注AI隐私与模型安全的研究者阅读

大规模语言模型(LLMs)已成为自然语言处理领域的主流范式,推动了科研与产业进步。随着模型规模和预训练数据量的增长,预训练数据暴露(PDE)问题日益突出,源于训练数据集的规模与不透明性。PDE 指的是判断特定数据是否出现在 LLM 的预训练语料中,对保障评估完整性与保护隐私至关重要,关联数据污染与成员推断两个关键领域。尽管概念相关,但二者常被孤立研究。本文首次在 PDE 框架下提供统一综述,形式化不同暴露层级,回顾攻击与防御方法,整合实证发现,并指出开放挑战与未来方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become the predominant paradigm in NLP, advancing both research and industry. As model sizes and pretraining data grow, concerns about Pretraining Data Exposure (PDE) increase due to the scale and opacity of training datasets. PDE refers to determining whether specific data appeared in an LLM's pretraining corpus. It is critical for ensuring evaluation integrity and protecting privacy, intersecting two key areas: data contamination and membership inference. Though conceptually related, these areas have often been studied in isolation. This paper offers the first unified survey of both under the PDE framework. We formalize PDE across exposure levels, review attack and defense methods, synthesize empirical findings, and highlight open challenges and future research directions.

大模型安全数据隐私成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。