通过原子分解与重组生成更难、更新颖的代码任务,提升大模型编程能力训练效果。
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

- 将代码任务拆解为原子元素并重新组合,生成新任务
- 在算法、工具使用等场景中显著提升模型编码能力
- 适合关注代码生成与强化学习训练的研究者
强化学习结合可验证奖励(RLVR)已成为塑造大语言模型(LLM)强大编程能力的核心方法。然而,RLVR的可扩展性受限于接近模型能力边界且足够挑战的可验证代码任务稀缺。以往研究多依赖启发式种子扩展生成数据,严重限制了任务的新颖性与难度,导致合成数据的训练价值无法随规模增长。为此,我们提出原子分解与重组(ADR)框架,通过将代码任务分解为原子单元并进行可控重组,实现真正新颖且具有挑战性的可验证代码任务生成。实验与分析表明,ADR在原创性、难度、多样性及测试质量上均优于现有基线,在算法编程、工具使用和数据科学等多个下游领域,持续带来更大的代码能力提升。本工作揭示了新型代码任务合成与可扩展RLVR训练的新范式。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However, the scalability of RLVR is severely constrained by the scarcity of sufficiently challenging verifiable code tasks that target near the model's edge of competence. Prior studies often rely on heuristic seed expansions for data synthesis, which severely limits both novelty and difficulty. Consequently, the training value of such data fails to scale proportionally with the size of its synthesis. To this end, we propose Atomic Decomposition and Recombination (ADR), a novel framework that generates verifiable code tasks via decomposition into atomic elements and controlled recombination, thereby enabling the generation of genuinely novel and challenging verifiable code tasks. Experiments and analysis demonstrate that ADR achieves superior originality, difficulty, diversity, and test quality over existing baselines, and consistently delivers greater improvements in code ability across RLVR in diverse downstream domains, including algorithmic programming, tool usage, and data science. Our work sheds light on a new paradigm for novel code task synthesis and scalable RLVR training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。