用数据配方提升大模型长文本推理能力,无需复杂奖励设计
Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

- 构建8个数据集,覆盖检索、多证据融合和推理三类任务
- 在7个基准上平均提升7.2~6.4分,优于现有强化学习训练集
- 适用于智能体任务,可显著提升GAIA和BrowseComp表现
长文本推理是大语言模型作为自主智能体时的关键能力。尽管强化学习(RL)已成为提升该能力的主要方法,但现有研究多聚焦于奖励工程,而多样化的训练数据仍十分稀缺。本文从数据中心视角重新审视该问题,证明仅通过一个简单有效的数据配方,配合最小化的基于结果的GRPO设置,即可显著提升长文本推理能力。我们针对检索、多证据融合和推理三类互补任务,构建并整理了共计约14,000个样本的八套数据集。在Qwen3-4B/8B/30B-A3B三个模型上的实验显示,在七个长文本基准上平均提升7.2、3.2和6.4分,超越此前的强化学习训练集。进一步实验证明,这些改进可迁移至智能体任务:在已调优的智能体模型上继续使用本数据配方进行强化训练,使GAIA提升4.8分,BrowseComp提升7.0分。相关数据集将公开以促进后续研究。
原文摘要 · Abstract (English)
Long-context reasoning is an essential capability for large language models, particularly when they are deployed as autonomous agents that must reason over lengthy trajectories. Reinforcement learning (RL) has recently emerged as a dominant paradigm for improving this ability, yet existing work largely focuses on reward engineering while diverse training data remains scarce. We revisit this problem from a data-centric perspective and show that a simple yet effective data recipe alone, paired with a minimal outcome-based GRPO setup, suffices to substantially improve long-context reasoning. Our recipe targets three complementary task families -- retrieval, multi-evidence synthesis, and reasoning -- for which we construct and curate eight datasets totaling ~14K examples. Experiments on three models (Qwen3-4B/8B/30B-A3B) yield average gains of +7.2/+3.2/+6.4 points across seven long-context benchmarks, surpassing prior RL training sets. We further demonstrate that these gains transfer to agentic tasks, where continuing RL training on an agent-tuned model with our data recipe improves GAIA by +4.8 and BrowseComp by +7.0 points. We will release our datasets to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。