arXiv:2501.09653cs.CLcs.AI2025-01中稿 · FORGE 2025 Dataset…被引 7

构建无污染多语言代码数据集,助力大模型公平评估

The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models

  • 从多语言代码中去重,避免与其他开源数据重复
  • 覆盖57种编程语言,支持跨语言模型评估
  • 适合关注模型泛化与评估公平性的研究者

大型语言模型的兴起推动了大规模代码数据集的开发。然而,这导致可用于下游行为研究或模型评估的代码资源有限,且易受数据污染影响。为解决此问题,我们发布了The Heap——一个涵盖57种编程语言的大规模多语言代码数据集,已对其他公开代码数据集进行去重处理,使研究人员能在几乎无需数据清洗的前提下,开展公正的大型语言模型评估。

原文摘要 · Abstract (English)

The recent rise in the popularity of large language models has spurred the development of extensive code datasets needed to train them. This has left limited code available for collection and use in the downstream investigation of specific behaviors, or evaluation of large language models without suffering from data contamination. To address this problem, we release The Heap, a large multilingual dataset covering 57 programming languages that has been deduplicated with respect to other open datasets of code, enabling researchers to conduct fair evaluations of large language models without significant data cleaning overhead.

代码生成多语言数据集评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。