arXiv:2503.17502cs.SEcs.AI2025-03被引 26

综述大模型在代码分析中的应用、模型与数据集现状

Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets

  • 系统梳理大模型在代码分析中的各类应用场景
  • 归纳常用模型与代表性数据集,揭示研究趋势
  • 适合关注AI辅助编程的开发者与研究者阅读

大型语言模型(LLMs)和基于Transformer的架构正被越来越多地应用于源代码分析。随着软件系统复杂度不断提升,将大模型融入代码分析流程对于提高效率、准确性和自动化水平至关重要。本文探讨了大模型在不同代码分析任务中的作用,重点关注三个核心方面:1)其可分析的任务类型及具体应用;2)所采用的模型;3)使用的数据集及其面临的挑战。为实现研究目标,我们分析了多篇相关学术论文,旨在揭示该新兴领域的研究进展、当前趋势及知识结构。此外,本文还总结了现有局限性,强调了对后续研究具有价值的关键工具、数据集与核心挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) and transformer-based architectures are increasingly utilized for source code analysis. As software systems grow in complexity, integrating LLMs into code analysis workflows becomes essential for enhancing efficiency, accuracy, and automation. This paper explores the role of LLMs for different code analysis tasks, focusing on three key aspects: 1) what they can analyze and their applications, 2) what models are used and 3) what datasets are used, and the challenges they face. Regarding the goal of this research, we investigate scholarly articles that explore the use of LLMs for source code analysis to uncover research developments, current trends, and the intellectual structure of this emerging field. Additionally, we summarize limitations and highlight essential tools, datasets, and key challenges, which could be valuable for future work.

代码分析大模型综述AI编程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。