检测7个主流多语言基准在7个大模型中的数据污染情况。
Contamination Report for Multilingual Benchmarks
- 用黑盒测试法分析多语言基准数据是否出现在模型训练中。
- 发现几乎所有模型都存在几乎全部基准的数据污染迹象。
- 为多语言模型评估提供更可靠的基准选择建议。
基准污染指测试数据出现在大型语言模型(LLM)的预训练或微调数据中。这会导致基准分数虚高,影响评估结果,难以真实反映模型能力。本文研究支持多种语言的主流多语言基准在7个常见开源与闭源大模型中的污染情况。采用黑盒测试方法,检验7个常用多语言基准在7个主流模型中的污染状况,发现几乎所有模型均表现出几乎所有基准的污染迹象。该研究有助于社区选择更可信的多语言评估基准。
原文摘要 · Abstract (English)
Benchmark contamination refers to the presence of test datasets in Large Language Model (LLM) pre-training or post-training data. Contamination can lead to inflated scores on benchmarks, compromising evaluation results and making it difficult to determine the capabilities of models. In this work, we study the contamination of popular multilingual benchmarks in LLMs that support multiple languages. We use the Black Box test to determine whether $7$ frequently used multilingual benchmarks are contaminated in $7$ popular open and closed LLMs and find that almost all models show signs of being contaminated with almost all the benchmarks we test. Our findings can help the community determine the best set of benchmarks to use for multilingual evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。