arXiv:2503.12072cs.CL2025-03NAACL被引 14

无需模型权重,用信息引导探针识别大模型记忆的训练数据。

Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models

  • 通过高意外性文本作为探测线索,定位模型记忆内容。
  • 在不访问模型参数下,成功识别出大量被记忆的训练文本。
  • 适合关注数据版权、模型可解释性与数据泄露风险的研究者。

高质量训练数据对构建高性能大语言模型至关重要,但商业LLM提供商极少披露训练数据细节。这种透明度缺失导致外部监督困难,影响版权审查、数据作者权益,并阻碍对数据污染与数据选择等关键问题的研究。如何恢复大模型所知的训练数据?本文提出一种新方法,仅通过信息引导探针即可识别如GPT-4等专有大模型的记忆内容,无需访问模型权重或概率输出。核心思路是:高意外性(surprisal)的文本片段是理想的记忆探测材料。通过评估模型重建这些高意外性词元的能力,我们成功识别出大量被模型记忆的文本,揭示了模型中隐藏的数据印记。

原文摘要 · Abstract (English)

High-quality training data has proven crucial for developing performant large language models (LLMs). However, commercial LLM providers disclose few, if any, details about the data used for training. This lack of transparency creates multiple challenges: it limits external oversight and inspection of LLMs for issues such as copyright infringement, it undermines the agency of data authors, and it hinders scientific research on critical issues such as data contamination and data selection. How can we recover what training data is known to LLMs? In this work, we demonstrate a new method to identify training data known to proprietary LLMs like GPT-4 without requiring any access to model weights or token probabilities, by using information-guided probes. Our work builds on a key observation: text passages with high surprisal are good search material for memorization probes. By evaluating a model's ability to successfully reconstruct high-surprisal tokens in text, we can identify a surprising number of texts memorized by LLMs.

大模型数据记忆隐私安全探针检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。