arXiv:2606.03800cs.LGcs.AI2026-06

用合成数据替代人工标注任务,可大幅降低RLVR训练成本且不影响效果。

Trading Human Curation for Synthetic Augmentation in RLVR

论文配图:Trading Human Curation for Synthetic Augmentation in RLVR
图 1 · 摘自论文原文
  • 通过预设筛选的合成数据增强替代部分人工设计任务
  • 在十项基准测试中保持了整体泛化能力
  • 合成任务与人工任务的成本比在1.4到11.6倍之间

强化学习从可验证奖励(RLVR)中训练智能体语言模型的核心瓶颈在于高质量训练任务的供给。每个任务需独立沙箱环境、提示词及人工编写的奖励函数,且仅通过质量筛选的任务才能提供有效信号。人工标注难以经济地扩展至大规模训练所需的任务量,而自动生成任务与人工任务之间的替代率尚不明确。本文研究以少量人工构建的基础任务为起点,通过预设门控筛选的合成增强方式,替代额外的人工标注。我们定义并测量了成本调整后的替代率ρ_{cost},通过控制变量实验分析不同合成比例下的训练语料表现,刻画了整个增强流程的端到端经济性。结果表明,在涵盖代码生成、指令遵循、推理和多轮调用等任务的十项基准测试中,使用合成内容替代人工任务仍能保持总体泛化性能。在合理的人工与合成任务成本比范围内,ρ_{cost}维持在[1.4×, 11.6×]区间。

原文摘要 · Abstract (English)

The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models. Each task requires a sandboxed setup, a prompt, and a hand-authored reward function, and only tasks that pass a quality bar produce useful training signal. Hand-curation at this quality bar does not scale economically to the task counts effective RL training requires, and the substitution rate between automatically generated task variants and human-authored ones is not yet established. We investigate using pre-specified, gate-filtered augmentations of a small hand-authored base as a substitute for additional human curation during RLVR. We formalize the cost-adjusted trade rate $ρ_{\text{cost}}$ between augmented and human-authored tasks, measure it through a controlled ablation across training corpora with varying augmentation share, and characterize the end-to-end economics of the augmentation pipeline. Substituting augmented content for additional human-authored tasks retains aggregate held-out generalization on a ten-benchmark suite spanning code, instruction following, reasoning, and multi-turn agentic function-calling. The cost-adjusted trade rate $ρ_{\text{cost}}$ between gated synthetic and human-authored RLVR tasks stays in $[1.4\times, 11.6\times]$ across the plausible $c_{\text{human}}/c_{\text{aug}}$ range.

强化学习合成数据任务生成成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。