提出分层检测框架,识别合成数据中隐藏的语义污染问题。
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
- 从词符、语义、推理模式到性能突降四层检测污染
- 在MMLU等数据集上,新方法F1达0.76,优于现有方法26.5%
- 适合关注合成数据安全性的模型训练与审计人员
合成数据已成为训练基础模型的关键,但基准数据污染威胁评估可靠性。现有方法仅能检测词符级重叠,无法发现语义级污染——即合成数据在概念上接近基准,却无词汇重合。这一漏洞尤为关键,因基础模型越来越多地使用可能隐含基准知识的合成数据进行训练。本文提出一种四层级污染检测框架:词符级、语义级、推理模式级和性能突降检测。在MMLU、GSM8K和HumanEval上的受控实验表明,语义级污染可规避现有方法(F1=0.17-0.49),而本框架有效识别,F1达0.76,平均比当前最优基线提升26.5%。该框架为从业者提供实用审计工具,支持合成数据的负责任部署。
原文摘要 · Abstract (English)
Synthetic data has become essential for training foundation models, yet benchmark contamination threatens evaluation integrity. Although existing detection methods identify token-level overlap, they fail to detect semantic-level contamination where synthetic data conceptually resemble benchmarks without lexical overlap. This gap is critical as foundation models increasingly train on synthetic data that may implicitly encode benchmark knowledge. We propose a hierarchical contamination detection framework operating at four levels: token level, semantic level, reasoning pattern, and performance cliff detection. Through controlled experiments on MMLU, GSM8K and HumanEval, we demonstrate that semantic-level contamination evades existing methods (F1=0.17-0.49) but is effectively detected by our hierarchical approach (F1 = 0.76), with an average improvement of 26. 5\% over state-of-the-art baselines. Our framework provides practitioners with practical tools for audit pipelines and enables responsible deployment of synthetic training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。