用互联网文本自动生成海量可验证强化学习任务,突破数据瓶颈
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
- 将填空题改造成多选题,从不可验证文本中提取推理步骤生成新任务
- 构建超70万条任务的数据集GooseReason,覆盖数学、编程和科学领域
- 在多个基准上实现新SOTA,尤其在网络安全领域超越专用大模型
强化学习与可验证奖励(RLVR)已成为激发大语言模型复杂推理能力的核心方法。然而,现有可验证数据有限,导致训练效果逐渐饱和。为此,本文提出Golden Goose,一种从不可验证的互联网文本中合成无限RLVR任务的简单技巧:通过提示大模型识别并遮蔽关键推理步骤,再生成多样且合理的干扰项,将原文转化为多选题形式。该方法使原本被排除的丰富推理文本(如科学教科书)得以利用,构建出包含超过70万条任务的GooseReason-0.7M大规模数据集,涵盖数学、编程和通用科学领域。实验表明,该数据能有效恢复在现有数据上陷入饱和的模型性能,在持续强化学习中实现稳健提升,并在15个多样化基准上刷新1.5B与4B-Instruct模型的SOTA。进一步地,将Golden Goose应用于真实场景,从原始FineWeb数据中合成网络安全领域的RLVR任务,训练的Qwen3-4B-Instruct模型在该领域达到新SOTA,超越经过大量领域预训练和后训练的7B专用模型,凸显了自动扩展推理丰富型不可验证文本为可验证任务的巨大潜力。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has become a cornerstone for unlocking complex reasoning in Large Language Models (LLMs). Yet, scaling up RL is bottlenecked by limited existing verifiable data, where improvements increasingly saturate over prolonged training. To overcome this, we propose Golden Goose, a simple trick to synthesize unlimited RLVR tasks from unverifiable internet text by constructing a multiple-choice question-answering version of the fill-in-the-middle task. Given a source text, we prompt an LLM to identify and mask key reasoning steps, then generate a set of diverse, plausible distractors. This enables us to leverage reasoning-rich unverifiable corpora typically excluded from prior RLVR data construction (e.g., science textbooks) to synthesize GooseReason-0.7M, a large-scale RLVR dataset with over 0.7 million tasks spanning mathematics, programming, and general scientific domains. Empirically, GooseReason effectively revives models saturated on existing RLVR data, yielding robust, sustained gains under continuous RL and achieving new state-of-the-art results for 1.5B and 4B-Instruct models across 15 diverse benchmarks. Finally, we deploy Golden Goose in a real-world setting, synthesizing RLVR tasks from raw FineWeb scrapes for the cybersecurity domain, where no prior RLVR data exists. Training Qwen3-4B-Instruct on the resulting data GooseReason-Cyber sets a new state-of-the-art in cybersecurity, surpassing a 7B domain-specialized model with extensive domain-specific pre-training and post-training. This highlights the potential of automatically scaling up RLVR data by exploiting abundant, reasoning-rich, unverifiable internet text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。