arXiv:2601.04526cs.SEcs.AI2026-01中稿 · ICSE 2026

提升代码大模型的性能,从数据、架构到推理全面优化。

Advancing Language Models for Code-related Tasks

  • 用代码差异引导的对抗增强和去噪技术提升数据质量
  • 引入语法引导的编码架构,显著改善模型对代码的理解能力
  • 结合提示工程与智能体机制,增强模型的复杂逻辑推理能力

近期语言模型(LMs)在软件工程任务中取得显著进展,但现有模型在复杂编程场景下仍表现不佳,受限于数据质量、模型架构与推理能力。本研究从三个互补方向系统性解决这些问题:(1) 通过代码差异引导的对抗增强技术(CODA)和代码去噪技术(CodeDenoise)提升代码数据质量;(2) 采用语法引导的代码语言模型(LEAM 和 LEAM++)改进模型架构;(3) 利用提示技术(muFiX)与基于智能体的方法(Specine)增强模型推理能力。这些方法旨在推动语言模型在软件开发中的实际应用,进一步促进智能软件工程的发展。

原文摘要 · Abstract (English)

Recent advances in language models (LMs) have driven significant progress in various software engineering tasks. However, existing LMs still struggle with complex programming scenarios due to limitations in data quality, model architecture, and reasoning capability. This research systematically addresses these challenges through three complementary directions: (1) improving code data quality with a code difference-guided adversarial augmentation technique (CODA) and a code denoising technique (CodeDenoise); (2) enhancing model architecture via syntax-guided code LMs (LEAM and LEAM++); and (3) advancing model reasoning with a prompting technique (muFiX) and an agent-based technique (Specine). These techniques aim to promote the practical adoption of LMs in software development and further advance intelligent software engineering.

代码生成语言模型软件工程推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。