arXiv:2503.15129cs.AI2025-03被引 44

用众包反馈优化代码生成模型,提升AI编程助手的准确性。

Aligning Crowd-sourced Human Feedback for Reinforcement Learning on Code Generation by Large Language Models

  • 通过贝叶斯优化分配反馈任务,降低人工标注负担。
  • 实验证明该方法显著提升大模型生成代码的质量。
  • 适用于多种领域特定语言,适合开发AI辅助编程工具的团队。

本文研究了大型语言模型(LLM)如何通过AI工具(如GitHub Copilot和Amazon CodeWhisperer)提升开发者能力,并结合众包计算的人类反馈来增强强化学习(RLHF),以改进文本到代码的生成。我们提出一种贝叶斯优化框架,有效分担反馈收集压力,突出高质量人类反馈的价值。实证评估表明,该方法可显著提升LLM代理在代码生成上的表现。该框架可扩展至通用领域特定语言,促进大模型能力与人类反馈在AI辅助编程中的对齐。

原文摘要 · Abstract (English)

This paper studies how AI-assisted programming and large language models (LLM) improve software developers' ability via AI tools (LLM agents) like Github Copilot and Amazon CodeWhisperer, while integrating human feedback to enhance reinforcement learning (RLHF) with crowd-sourced computation to enhance text-to-code generation. Additionally, we demonstrate that our Bayesian optimization framework supports AI alignment in code generation by distributing the feedback collection burden, highlighting the value of collecting human feedback of good quality. Our empirical evaluations demonstrate the efficacy of this approach, showcasing how LLM agents can be effectively trained for improved text-to-code generation. Our Bayesian optimization framework can be designed for general domain-specific languages, promoting the alignment of large language model capabilities with human feedback in AI-assisted programming for code generation.

代码生成强化学习众包反馈AI编程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。