评测大模型全栈编程能力,覆盖16种语言和多领域真实任务。
FullStack Bench: Evaluating LLMs as Full Stack Coders
- 构建涵盖16种语言的全栈编程评测集,包含真实任务指令与单元测试。
- 在全栈任务上,主流大模型平均通过率不足50%,暴露其实际能力短板。
- 开源支持多语言的SandboxFusion执行工具,提升评估效率与真实性。
随着代码大模型能力持续扩展,其在各类代码智能场景中的应用迅速增长。然而,现有数据集大多仅覆盖有限的应用领域。为填补这一空白,我们构建了一个面向全栈编程的综合性代码评估数据集FullStack Bench,涵盖基础编程、数据分析、软件工程、数学及机器学习等多个应用领域。为评估多语言编程能力,我们在FullStack Bench中设计了来自16种主流编程语言的真实任务指令与对应的单元测试用例,反映真实使用场景而非简单翻译。此外,我们还发布了高效的代码沙箱执行工具SandboxFusion,支持多种编程语言和包,可高效评估FullStack Bench的表现。在该数据集上的综合实验结果表明,FullStack Bench与SandboxFusion具有必要性和有效性。
原文摘要 · Abstract (English)
As the capabilities of code large language models (LLMs) continue to expand, their applications across diverse code intelligence domains are rapidly increasing. However, most existing datasets only evaluate limited application domains. To address this gap, we have developed a comprehensive code evaluation dataset FullStack Bench focusing on full-stack programming, which encompasses a wide range of application domains (e.g., basic programming, data analysis, software engineering, mathematics, and machine learning). Besides, to assess multilingual programming capabilities, in FullStack Bench, we design real-world instructions and corresponding unit test cases from 16 widely-used programming languages to reflect real-world usage scenarios rather than simple translations. Moreover, we also release an effective code sandbox execution tool (i.e., SandboxFusion) supporting various programming languages and packages to evaluate the performance of our FullStack Bench efficiently. Comprehensive experimental results on our FullStack Bench demonstrate the necessity and effectiveness of our FullStack Bench and SandboxFusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。