用知识蒸馏伪装能力,人为抬高模型评测分数
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
- 通过伪合法训练步骤转移特定评测知识
- 2层BERT模型在GPQA上准确率提升75%无真实推理能力
- 提醒研究者警惕误用导致结果虚高,适合评估安全研究者
本文揭示知识蒸馏可被滥用以操纵语言模型评测分数,暴露当前评估体系的关键漏洞。我们提出“数据清洗”(Data Laundering)方法,通过看似正当的中间训练步骤,隐蔽传递针对特定基准的知识。在2层BERT学生模型上,大量实验表明该方法可显著提升基准准确率(如GPQA最高达75%),但模型并未获得真正的推理能力。此方法可能被有意或无意地使用,研究人员可能在未察觉的情况下夸大成绩。本研究旨在警示评估完整性的重要性,呼吁建立更可靠的评测机制。代码已开源。
原文摘要 · Abstract (English)
In this paper, we show that knowledge distillation can be subverted to manipulate language model benchmark scores, revealing a critical vulnerability in current evaluation practices. We introduce "Data Laundering," a process that enables the covert transfer of benchmark-specific knowledge through seemingly legitimate intermediate training steps. Through extensive experiments with a 2-layer BERT student model, we show how this approach can achieve substantial improvements in benchmark accuracy (up to 75\% on GPQA) without developing genuine reasoning capabilities. Notably, this method can be exploited intentionally or even unintentionally, as researchers may inadvertently adopt this method and inflate scores without realising the implications. While our findings demonstrate the effectiveness of this technique, we present them as a cautionary tale highlighting the urgent need for more robust evaluation methods in AI. This work aims to contribute to the ongoing discussion about evaluation integrity in AI development and the need for benchmarks that more accurately reflect true model capabilities. The code is available at https://github.com/mbzuai-nlp/data_laundering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。