arXiv:2510.09717cs.LGcs.AI2025-10被引 2

提出可证明的训练数据识别方法,严格控制误识别率。

Provable Training Data Identification for Large Language Models

  • 将数据识别转为集合级推断,用置信p值与修正边界估计比例。
  • 在多个模型和数据集上实现更高检测力且严格控制错误率。
  • 适合版权诉讼、隐私审计等需可靠证据的场景。

大模型训练数据的识别对版权纠纷、隐私审计和公平评估至关重要。现有方法通常以个体为单位进行识别,无法控制识别集合的错误率,难以提供统计上可靠的证据。本文将训练数据识别形式化为集合级推断问题,提出无分布假设的可证明训练数据识别(PTDI)方法,实现严格可控的误识别率(FIR)。具体而言,利用一组已知未见数据计算每个数据点的分位数p值,通过新型杰克刀校正贝塔边界(JKBB)估计测试集中的训练数据占比,并据此缩放p值;再应用贝尼尼-霍赫伯格(BH)过程筛选出具有可证明且严格控制误识别率的数据子集。大量实验表明,PTDI在多种模型与数据集上均显著提升检测能力,同时严格控制误识别率。

原文摘要 · Abstract (English)

Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. However, existing works typically treat this task as an instance-wise identification without controlling the error rate of the identified set, which cannot provide statistically reliable evidence. In this work, we formalize training data identification as a set-level inference problem and propose Provable Training Data Identification (PTDI), a distribution-free approach that enables provable and strict false identification rate control. Specifically, our method computes conformal p-values for each data point using a set of known unseen data and then develops a novel Jackknife-corrected Beta boundary (JKBB) estimator to estimate the training-data proportion of the test set, which allows us to scale these p-values. By applying the Benjamini-Hochberg (BH) procedure to the scaled p-values, we select a subset of data points with provable and strict false identification control. Extensive experiments across various models and datasets demonstrate that PTDI achieves higher power than prior methods while strictly controlling the FIR.

数据识别可证明性大模型隐私审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。