用语言化过程监督提升代码生成智能体的推理能力
Verbal Process Supervision Elicits Better Coding Agents
- 通过语言化过程监督让模型在编码时显式表达思考步骤
- 在BigCodeBench上比基线提升3.65%准确率
- 适合需要高可靠性代码生成的研究与工程场景
大型语言模型及其作为AI智能体的应用已显著推动代码生成基准的发展,重塑现代软件工程任务。然而,即使采用测试时计算的推理模型,这些系统仍难以应对复杂的软件工程挑战。本文提出CURA,一种结合语言化过程监督(VPS)的代码理解与推理智能体系统,在BigCodeBench等挑战性基准上相比基线模型提升了3.65%。此外,当与o3-mini模型配合使用时,CURA达到当前最优性能。该工作推进了基于大模型的代码生成与推理驱动架构的融合,使语言模型具备解决复杂软件工程任务的智能体推理能力。
原文摘要 · Abstract (English)
The emergence of large language models and their applications as AI agents have significantly advanced state-of-the-art code generation benchmarks, transforming modern software engineering tasks. However, even with test-time computed reasoning models, these systems still struggle with complex software engineering challenges. This work introduces CURA, a code understanding and reasoning agent system enhanced with verbal process supervision (VPS), achieving a 3.65\% improvement over baseline models on challenging benchmarks like BigCodeBench. Furthermore, CURA, when paired with the o3-mini model and VPS techniques, attains state-of-the-art performance. This work represents a step forward in integrating reasoning-driven architectures with LLM-based code generation, enabling agentic reasoning for language models to solve complex software engineering tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。