arXiv:2507.16414cs.AI2025-07

通过神经元激活差异,精准识别大模型训练数据来源。

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework

  • 基于神经元激活模式差异检测训练数据
  • 在多个模型和基准上显著超越现有方法
  • 适用于数据合规性审查与版权保护场景

大型语言模型(LLMs)的性能与其训练数据密切相关,而这些数据可能包含受版权保护的内容或隐私信息,引发法律与伦理问题。此外,数据集污染和模型内化偏见也受到广泛关注。为此,提出了预训练数据检测(PDD)任务,用于判断特定数据是否被纳入模型的预训练语料库。然而,现有PDD方法多依赖预测置信度、损失等表面特征,性能有限。本文提出NA-PDD,一种分析训练数据与非训练数据在推理时神经元激活模式差异的新算法。该方法基于观察:不同数据类型会激活不同的神经元。同时,构建了时间无偏的基准CCNewsPDD,通过严格的文本变换确保训练与非训练数据的时间分布一致。实验表明,NA-PDD在三个基准及多种大模型上均显著优于现有方法。

原文摘要 · Abstract (English)

The performance of large language models (LLMs) is closely tied to their training data, which can include copyrighted material or private information, raising legal and ethical concerns. Additionally, LLMs face criticism for dataset contamination and internalizing biases. To address these issues, the Pre-Training Data Detection (PDD) task was proposed to identify if specific data was included in an LLM's pre-training corpus. However, existing PDD methods often rely on superficial features like prediction confidence and loss, resulting in mediocre performance. To improve this, we introduce NA-PDD, a novel algorithm analyzing differential neuron activation patterns between training and non-training data in LLMs. This is based on the observation that these data types activate different neurons during LLM inference. We also introduce CCNewsPDD, a temporally unbiased benchmark employing rigorous data transformations to ensure consistent time distributions between training and non-training data. Our experiments demonstrate that NA-PDD significantly outperforms existing methods across three benchmarks and multiple LLMs.

数据检测神经元分析模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。