arXiv:2411.00005cs.SEcs.AI2024-11NAACL被引 4

系统梳理代码大模型数据合成技术,助你高效训练更优代码生成模型。

Mastering the Craft of Data Synthesis for CodeLLMs

  • 构建代码数据合成与筛选的技术分类体系
  • 总结近期进展并指出关键挑战与改进方向
  • 为新研究者提供实用入门指南

大型语言模型(LLMs)在代码理解与生成方面表现卓越,使编程任务成为研究热点,兼具实际应用价值和作为大模型评估基准的意义。数据合成与过滤技术在此领域被广泛采用,且已被证明极为有效。本文系统地回顾并构建了这些技术的分类体系,重点聚焦最新进展。我们揭示了关键技术挑战,探讨了未来研究方向,并为初入该领域的研究者提供了实用指导。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown impressive performance in \emph{code} understanding and generation, making coding tasks a key focus for researchers due to their practical applications and value as a testbed for LLM evaluation. Data synthesis and filtering techniques have been widely adopted and shown to be highly effective in this context. In this paper, we present a focused survey and taxonomy of these techniques, emphasizing recent advancements. We highlight key challenges, explore future research directions, and offer practical guidance for new researchers entering the field.

代码生成数据合成大模型训练综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。