梳理代码大模型的分类体系,揭示其在编程任务中的应用与局限。
Code LLMs: A Taxonomy-based Survey
- 构建基于分类框架的代码大模型分析体系
- 系统归纳模型架构、训练方法与应用模式
- 适合关注AI编程工具研发的研究者与开发者
大型语言模型(LLMs)在自然语言处理任务中表现出色,并已拓展至编程领域,弥合了自然语言(NL)与编程语言(PL)之间的鸿沟。本综述提出一种基于分类的框架,对NL-PL领域的LLMs进行全面分析,探讨其在编码任务中的应用方式,考察其方法论、架构设计及训练过程。该框架对相关概念进行分类,提供统一的分类体系,有助于深入理解这一快速发展的领域。本文还总结了当前研究现状与未来发展方向,涵盖应用场景与现存挑战。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities across various NLP tasks and have recently expanded their impact to coding tasks, bridging the gap between natural languages (NL) and programming languages (PL). This taxonomy-based survey provides a comprehensive analysis of LLMs in the NL-PL domain, investigating how these models are utilized in coding tasks and examining their methodologies, architectures, and training processes. We propose a taxonomy-based framework that categorizes relevant concepts, providing a unified classification system to facilitate a deeper understanding of this rapidly evolving field. This survey offers insights into the current state and future directions of LLMs in coding tasks, including their applications and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。