用翻译基准评估多语言代码模型,发现结果可复现但存在不一致
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
- 用OctoPack和MultiPL-E翻译HumanEval构建多语言测试集
- 模型在翻译基准上表现与训练时困惑度一致,验证有效性
- 不同基准间性能差异大,结果难以复现,需更多实证研究
评估代码语言模型(CLMs)在多语言及低资源编程语言环境下的软件工程任务表现面临挑战,主要源于缺乏高质量跨语言基准以及训练语料的不平衡性。尽管近年通过不同方法引入翻译后的HumanEval基准,在代码生成任务中展现潜力,但尚无实证研究评估这些基准的有效性。为此,我们开展初步研究,评估Poly-Coder这一开创性的开源多语言代码生成模型。利用OctoPack和MultiPL-E两项最新研究提供的HumanEval翻译版本,结果表明:翻译基准上的表现与训练阶段的困惑度指标高度吻合,验证了其作为性能估测工具的有效性。然而,我们也发现模型在不同翻译基准间表现不一致,且存在结果难以复现的问题。这些初步发现凸显了深入研究翻译基准方法学、局限性和可复现性的必要性,以确保其在广泛采用前具备可靠性。
原文摘要 · Abstract (English)
Evaluating the performance of Code Language Models (CLMs) for software engineering tasks, especially in multilingual and low-resource programming language settings, poses significant challenges. These challenges are primarily due to the lack of high-quality benchmarks across various programming languages and the imbalanced nature of the CLMs training corpus. Although recent advances in one of the common downstream tasks, code generation, have shown promise by introducing translated benchmarks using different methodologies, there is a current lack of empirical evidence assessing these benchmarks. To address this gap, we conducted a preliminary study to evaluate the performance of Poly-Coder, a pioneering open-source, multilingual CLM built for code generation. We utilized two existing state-of-the-art translations of the popular code generation benchmark, HumanEval, facilitated by the OctoPack and MultiPL-E studies. Our results suggest that the outcomes observed in these translated benchmarks align well with evaluation metrics used during the training phase, such as perplexity, thereby validating their effectiveness in estimating the performance of CLMs. However, we identified several inconsistencies in the CLMs' performance across the translated benchmarks and encountered challenges in replicating the results. These initial insights highlight the need for more comprehensive empirical studies to fully understand translated benchmarks' methodological approaches, limitations, and reproducibility. Such studies are essential to ensure their reliability before they are widely adopted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。