构建多语言代码生成评测集,揭示大模型跨语言能力短板
mHumanEval -- A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
- 用200+语言构建代码生成测试集,覆盖低资源语种
- 引入人工校对翻译,确保15种语言数据质量
- 发现主流模型在非英语提示下性能显著下降
大语言模型在自然语言转代码方面取得显著进展,但现有评测集如HumanEval仍存在任务多样性、覆盖率和语言范围不足的问题。当前评估主要集中在英语到Python的转换,且测试用例有限,可能高估模型表现。尽管已有研究扩展了测试覆盖率和编程语言种类,但低资源语言提示下的代码生成仍鲜有探索。为此,我们提出mHumanEval,一个支持超过200种自然语言提示的多语言评测基准。通过成熟的机器翻译方法构建数据集,并辅以专家人工校对,确保15种多样语言的数据质量。最后分析了当前最先进(SOTA)代码大模型在多语言场景下的表现,揭示了跨语言代码生成的现状与挑战。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations, particularly in task diversity, test coverage, and linguistic scope. Current evaluations primarily focus on English-to-Python conversion tasks with limited test cases, potentially overestimating model performance. While recent works have addressed test coverage and programming language (PL) diversity, code generation from low-resource language prompts remains largely unexplored. To address this gap, we introduce mHumanEval, an extended benchmark supporting prompts in over 200 natural languages. We employ established machine translation methods to compile the benchmark, coupled with a quality assurance process. Furthermore, we provide expert human translations for 15 diverse natural languages (NLs). We conclude by analyzing the multilingual code generation capabilities of state-of-the-art (SOTA) Code LLMs, offering insights into the current landscape of cross-lingual code generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。