arXiv:2502.17259cs.CRcs.AI2025-02被引 7

用可检测水印标记测试集,防模型作弊评估

Detecting Benchmark Contamination Through Watermarking

  • 用带水印的LLM重写题目,保持测试集可用性
  • 水印残留可被统计方法检测,$p$值达$10^{-3}$
  • 适合关注评测公平性的模型开发者

基准数据集污染严重威胁大语言模型评估的可靠性,因难以判断模型是否在测试集上训练过。我们提出一种解决方案:在发布前对基准数据集进行水印标记。该方法通过带有水印的LLM重新表述原始问题,不改变基准实用性。评估时,可使用理论基础坚实的统计检验检测模型训练过程中留下的“放射性”痕迹。我们通过从头预训练10亿参数模型(共100亿词元)并控制基准污染程度进行测试,在ARC-Easy、ARC-Challenge和MMLU上验证了该方法的有效性。结果显示,水印后基准实用性基本不变,当模型因污染显著提升性能时(如在ARC-Easy上提升5%),检测成功率高,$p$-值低至$10^{-3}$。

原文摘要 · Abstract (English)

Benchmark contamination poses a significant challenge to the reliability of Large Language Models (LLMs) evaluations, as it is difficult to assert whether a model has been trained on a test set. We introduce a solution to this problem by watermarking benchmarks before their release. The embedding involves reformulating the original questions with a watermarked LLM, in a way that does not alter the benchmark utility. During evaluation, we can detect ``radioactivity'', \ie traces that the text watermarks leave in the model during training, using a theoretically grounded statistical test. We test our method by pre-training 1B models from scratch on 10B tokens with controlled benchmark contamination, and validate its effectiveness in detecting contamination on ARC-Easy, ARC-Challenge, and MMLU. Results show similar benchmark utility post-watermarking and successful contamination detection when models are contaminated enough to enhance performance, \eg $p$-val $=10^{-3}$ for +5$\%$ on ARC-Easy.

模型评估水印技术基准污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。