arXiv:2410.03981cs.SEcs.LG2024-10中稿 · ACM Transactions o…综述被引 146

系统梳理大模型在低资源与专用编程语言中的代码生成挑战与进展

A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages

  • 聚焦低资源与专用语言的代码生成难题,分析数据稀缺与语法特异性问题
  • 筛选111篇论文,总结六类提升方法与四大评估技术,揭示性能瓶颈
  • 为研究者提供领域指南,推动专用语言大模型应用落地

大语言模型在主流编程语言的代码生成方面表现卓越,但在低资源编程语言(LRPLs)和领域特定语言(DSLs)上的表现仍面临重大挑战,影响数百万开发者——仅Rust用户就达350万——无法充分受益于大模型能力。LRPLs和DSLs存在数据稀缺、语法特殊等问题,难以在通用语料中得到良好表征。解决这些挑战至关重要,因为它们能显著提升金融、科学等专业领域的开发效率。尽管已有多个关于大模型在软件工程中应用的综述,但尚无专门针对LRPLs和DSLs的系统性综述。本调查通过筛选2020至2024年间超过2.7万篇论文中的111篇,系统回顾了大模型在该领域的现状、方法与挑战。我们总结了使用的模型、评测基准与指标,以及数据集构建与清洗策略;识别出四种主要评估方法和若干关键评价指标;将改进方法归为六大类,并归纳了研究人员提出的新型架构。尽管已有多种技术和指标,但当前仍缺乏统一的评估标准和基准数据集。本综述为大模型、软件工程与专用语言交叉领域的研究者与实践者提供了重要参考,为未来该方向的发展奠定基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive capabilities in code generation for popular programming languages. However, their performance on Low-Resource Programming Languages (LRPLs) and Domain-Specific Languages (DSLs) remains a significant challenge, affecting millions of developers-3.5 million users in Rust alone-who cannot fully utilize LLM capabilities. LRPLs and DSLs encounter unique obstacles, including data scarcity and, for DSLs, specialized syntax that is poorly represented in general-purpose datasets. Addressing these challenges is crucial, as LRPLs and DSLs enhance development efficiency in specialized domains, such as finance and science. While several surveys discuss LLMs in software engineering, none focus specifically on the challenges and opportunities associated with LRPLs and DSLs. Our survey fills this gap by systematically reviewing the current state, methodologies, and challenges in leveraging LLMs for code generation in these languages. We filtered 111 papers from over 27,000 published studies between 2020 and 2024 to evaluate the capabilities and limitations of LLMs in LRPLs and DSLs. We report the LLMs used, benchmarks, and metrics for evaluation, strategies for enhancing performance, and methods for dataset collection and curation. We identified four main evaluation techniques and several metrics for assessing code generation in LRPLs and DSLs. Our analysis categorizes improvement methods into six groups and summarizes novel architectures proposed by researchers. Despite various techniques and metrics, a standard approach and benchmark dataset for evaluating code generation in LRPLs and DSLs are lacking. This survey serves as a resource for researchers and practitioners at the intersection of LLMs, software engineering, and specialized programming languages, laying the groundwork for future advancements in code generation for LRPLs and DSLs.

大模型代码生成低资源语言领域语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。