arXiv:2601.15188cs.SEcs.AI2026-01被引 1

评测大模型生成ABAP代码能力,发现强模型经编译反馈可达75%成功率

Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback

  • 用180个任务测试大模型写ABAP代码,结合编译器反馈迭代优化
  • 强模型经多轮迭代后成功率达75%,弱模型表现明显更差
  • 适合关注企业级系统开发与自动纠错的工程师参考

本文研究大型语言模型(LLMs)在生成ABAP代码方面的表现。尽管生成式AI已在多种编程语言中取得成功,但针对ABAP代码生成的系统性分析仍很少。本研究旨在实证分析不同LLMs生成语法正确且功能有效的ABAP代码的能力,评估其利用编译器反馈进行迭代改进的效果,并识别特殊挑战的任务类型。为此,构建了一个包含180个任务的基准测试,涵盖改编的HumanEval任务和实际SAP应用场景。结果表明,模型间表现差异显著:更强的LLMs经过多轮迭代后成功率可达约75%,并极大受益于编译器反馈;而较小模型表现明显较差。整体而言,研究凸显了强大LLMs在ABAP开发流程中的高潜力,尤其在迭代错误修正方面。

原文摘要 · Abstract (English)

This work investigates the performance of Large Language Models (LLMs) in generating ABAP code. Despite successful applications of generative AI in many programming languages, there are hardly any systematic analyses of ABAP code generation to date. The aim of the study is to empirically analyze to what extent various LLMs can generate syntactically correct and functional ABAP code, how effectively they use compiler feedback for iterative improvement, and which task types pose special challenges. For this purpose, a benchmark with 180 tasks is conducted, consisting of adapted HumanEval tasks and practical SAP scenarios. The results show significant performance differences between the models: more powerful LLMs achieve success rates of around 75% after several iterations and benefit greatly from compiler feedback, while smaller models perform significantly weaker. Overall, the study highlights the high potential of powerful LLMs for ABAP development processes, especially in iterative error correction.

ABAP生成大模型代码纠错迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。