arXiv:2506.13932cs.SEcs.AI2025-06综述被引 1

系统梳理代码推理技术,揭示其对软件工程任务的提升作用

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

  • 聚焦代码特有的推理方法,结合结构与执行反馈提升性能
  • 实证发现代码专属信号能显著改善生成、测试等任务表现
  • 适合关注大模型在编程中应用的研究者与开发者

大型语言模型(LLMs)在自然语言任务中取得显著进展,通过引入测试时推理技术可进一步提升性能。这些推理机制已被应用于代码领域,支持代码生成、测试生成和缺陷修复等复杂软件工程(SWE)任务。然而,不同推理技术对代码相关SWE任务的影响尚未系统研究。本文综述了支撑这些能力的代码推理技术,重点关注测试时计算与推理范式。分析多种代码专用推理方法,并逐步构建到融合规划、工具使用和多步交互的SWE智能体。同时,在常用模型与基准上比较不同技术的影响,揭示其相对重要性,并提出开放挑战与未来方向。结果表明,利用代码特有信号(如结构信息与执行反馈)的方法通常带来性能提升,推动针对代码推理的专门研究超越自然语言推理范畴。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) has led to dramatic improvements across a wide range of natural language tasks. Their performance on certain tasks can be further enhanced by incorporating test-time reasoning techniques. These inference-time advances have been adopted into the code domain, enabling complex software engineering (SWE) tasks such as code generation, test generation and issue resolution. However, the impact of different reasoning techniques on code-centric SWE tasks has not been systematically explored. In this work, we survey code reasoning techniques that underpin these capabilities, with a focus on test-time compute and inference-time reasoning paradigms. We examine a variety of code-specific reasoning methods and progressively build up to SWE agents, which combine planning, tool use, and multi-step interaction. We also compare the impact of different techniques on coding tasks, highlighting their relative importance and outlining open challenges and future research directions. Across commonly used models and benchmarks, we find that approaches exploiting code-specific signals (e.g., structure and execution feedback) are frequently associated with improved performance, motivating a dedicated study of code reasoning beyond natural-language reasoning.

代码推理大模型软件工程LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。