9个模型开发者公开训练测试数据重叠情况,呼吁提升评估透明度
Language model developers should report train-test overlap
- 调研30家开发机构,仅9家披露训练测试数据重叠情况
- 4家开源训练数据可直接测算重叠率,5家公布测算方法与数据
- 建议所有评估结果均应公开重叠统计或训练数据以增强可信度
语言模型广泛评估,但正确解读结果需了解训练-测试数据重叠程度。当前公众缺乏相关透明信息:多数模型未公开重叠数据,第三方无法直接测量因无训练数据访问权。我们调研了30家模型开发者,发现仅9家报告重叠情况:4家开放训练数据供社区直接测算,5家发布重叠计算方法与统计数据。通过协作,我们为另外3家开发者获得新信息。我们主张,模型开发者在公开测试集上报告评估结果时,应同时提供训练-测试重叠统计或训练数据。此举旨在提升评估透明度,增强社区对模型评价的信任。
原文摘要 · Abstract (English)
Language models are extensively evaluated, but correctly interpreting evaluation results requires knowledge of train-test overlap which refers to the extent to which the language model is trained on the very data it is being tested on. The public currently lacks adequate information about train-test overlap: most models have no public train-test overlap statistics, and third parties cannot directly measure train-test overlap since they do not have access to the training data. To make this clear, we document the practices of 30 model developers, finding that just 9 developers report train-test overlap: 4 developers release training data under open-source licenses, enabling the community to directly measure train-test overlap, and 5 developers publish their train-test overlap methodology and statistics. By engaging with language model developers, we provide novel information about train-test overlap for three additional developers. Overall, we take the position that language model developers should publish train-test overlap statistics and/or training data whenever they report evaluation results on public test sets. We hope our work increases transparency into train-test overlap to increase the community-wide trust in model evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。