首个针对Mojo语言的代码生成框架与评测基准。
MojoBench: Language Modeling and Benchmarks for Mojo
- 构建Mojo代码生成专用评测集HumanEval-Mojo和Mojo-Coder模型。
- Mojo-Coder在多项任务上比GPT-4o和Claude-3.5-Sonnet提升30%-35%。
- 揭示大模型对新兴编程语言的适应性,助力未来代码生成系统优化。
由Modular推出的新型编程语言Mojo因其宣称的显著性能提升,在科学界引发广泛关注。尽管代码大语言模型(LLM)在多种编程语言上取得进展,但Mojo尚未被系统研究。为此,我们提出MojoBench,首个针对Mojo代码生成的框架。该框架包含为评估代码LLM而设计的HumanEval-Mojo基准数据集,以及首个在Mojo上预训练并微调的LLM——Mojo-Coder,支持五种自然语言指令。实验表明,Mojo-Coder在多项任务中较GPT-4o和Claude-3.5-Sonnet提升30%-35%。此外,研究还揭示了大模型在低频及未见编程语言上的行为特征,为提升模型泛化能力提供策略参考。MojoBench有助于深入理解大模型在新兴编程范式中的能力与局限,推动更鲁棒的代码生成系统发展。
原文摘要 · Abstract (English)
The recently introduced Mojo programming language (PL) by Modular, has received significant attention in the scientific community due to its claimed significant speed boost over Python. Despite advancements in code Large Language Models (LLMs) across various PLs, Mojo remains unexplored in this context. To address this gap, we introduce MojoBench, the first framework for Mojo code generation. MojoBench includes HumanEval-Mojo, a benchmark dataset designed for evaluating code LLMs on Mojo, and Mojo-Coder, the first LLM pretrained and finetuned for Mojo code generation, which supports instructions in 5 natural languages (NLs). Our results show that Mojo-Coder achieves a 30-35% performance improvement over leading models like GPT-4o and Claude-3.5-Sonnet. Furthermore, we provide insights into LLM behavior with underrepresented and unseen PLs, offering potential strategies for enhancing model adaptability. MojoBench contributes to our understanding of LLM capabilities and limitations in emerging programming paradigms fostering more robust code generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。