arXiv:2605.30462cs.LGcs.AI2026-05

通过模型内部语义关联指纹,可精准识别训练数据来源。

idSCD: Identifying Training Datasets through Semantic Correlation Descriptors

论文配图:idSCD: Identifying Training Datasets through Semantic Correlation Descriptors
图 1 · 摘自论文原文
  • 用语义相关性描述符提取模型内部学习的特有语义结构。
  • 在三类任务中平均性能超越现有方法,最大提升超60%。
  • 适用于数据集成员身份推断,尤其对语义差异明显的数据有效。

能否从模型训练中产生的偶然相关性识别其训练数据?我们提出,数据集会在模型学习到的语义相关结构中留下特定痕迹:某些在数据集中具有预测性但非任务因果性的偶然规律,会在训练过程中被内化。基于此,我们研究了数据集层面的成员身份推断,突破了依赖置信度、损失、生成样本等行为或分布证据的现有方法。提出一种白盒语义指纹方法——语义相关性描述符(SCD),可捕捉模型学习到的语义相关结构,并实现跨数据集混合的可比性。在留一数据集排除的控制实验中,SCDs能完美区分匹配与非匹配数据集对。进一步提出基于SCD的实用成员身份得分,仅需目标数据集的独立SCD和模型的SCD,无需构建留一模型。在自然语言推理、情绪分类和医疗文本分类三类实验设置中,测试了不同语义分离度和关键词支持下的优劣。该得分分类器平均表现最佳,标准差最低,显著优于黑盒基线RMIA、Attack-P、LiRA及白盒基线SIF。结果表明,通过内部语义关联可追踪数据集成员身份,当数据组具明显语义特征时,ROC-AUC相对提升超过60%。

原文摘要 · Abstract (English)

Can a dataset be recognized from the spurious correlations it induces during training? We argue that datasets leave dataset-specific traces in a model's learned semantic correlation structure: incidental regularities that are predictive within a dataset, but not causal for the underlying task, can be internalized during training. We use this insight to study dataset-level membership inference, moving beyond existing methods that rely on behavioral or distributional evidence such as confidence scores, losses, margins, generated samples, or query responses. We introduce a white-box semantic fingerprinting approach based on semantic correlation descriptors (SCDs), which capture the semantic correlation structure learned by a model and make it comparable across dataset mixtures. In a controlled leave-one-dataset-out diagnostic, SCDs recover dataset-specific changes and perfectly separate matching from non-matching dataset pairs. We then propose a practical SCD-based membership score that tests whether a target dataset is part of a model's training mixture using only the model's SCD and the target dataset's standalone SCD, without requiring leave-one-dataset-out models. Across three diverse experimental settings, with dataset groups for natural language inference, emotion classification, and medical text classification, we test both the advantages and limitations of SCD-based membership inference with different degrees of semantic separation and keyword support between dataset splits. On average, the classifier based on this score achieves the highest performance and the lowest std, outperforming black-box baselines RMIA, Attack-P, and LiRA, as well as the white-box SIF baseline. These results show that dataset membership can be traced through internal semantic correlations, with the largest relative gain exceeding 60% in ROC-AUC when dataset groups expose distinct semantic particularities.

数据溯源语义指纹成员身份推断模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。