在数据和算力有限时,用可生成的程序化数据提升小模型推理能力。
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes

- 用可控规模、多样性和复杂度的程序化数据评估小模型在低数据下的表现。
- 低复杂度任务训练的模型能泛化到高复杂度任务,混合复杂度数据最高效,样本效率最高提升5倍。
- 适合关注高效微调、数据稀缺场景或想构建智能评估数据集的研究者。
微调大语言模型通常依赖大量高质量标注数据,或在可验证奖励强化学习(RLVR)中需要有明确答案的问题。尽管已有研究探讨了扩大数据与计算资源对模型推理能力的提升,但这些成果在标注数据和算力匮乏的真实场景中难以应用。本文对开源小语言模型(SLM)在低数据条件下进行RLVR后的性能进行了全面实证研究。我们在三个新构建的数据集上开展实验,涵盖数数、图推理和空间推理任务,系统分析了模型性能随数据集规模、多样性及复杂度的变化规律。结果表明:(1)程序化数据支持对训练集进行细粒度设计,具备可控的大小、多样性和复杂度;(2)在RLVR下,于低复杂度任务训练的模型可有效泛化至高复杂度任务;(3)在低数据条件下,混合复杂度数据训练带来的收益最大,相比仅使用简单任务,样本效率最高提升5倍。这些发现为未来探索RLVR的数据扩展规律以及利用程序化生成器优化高效微调数据开发提供了启示。
原文摘要 · Abstract (English)
Fine-tuning Large Language Models (LLMs) typically relies on large quantities of high-quality annotated data, or questions with well-defined ground truth answers in the case of Reinforcement Learning with Verifiable Rewards (RLVR). While previous work has explored the benefits to model reasoning capabilities by scaling both data and compute used for RLVR, these results lack applicability in many real-world settings where annotated data and accessible compute may be scarce. In this work, we present a comprehensive empirical study of open-source Small Language Model (SLM) performance after RLVR in low data regimes. Across three novel datasets covering number counting problems, graph reasoning, and spatial reasoning, we characterize how model performance scales with dataset size, diversity, and complexity. We demonstrate that (1) procedural datasets allow for fine-grained evaluation and training dataset development with controllable properties (size, diversity, and complexity), (2) under RLVR, models trained on lower complexity tasks can generalize to higher complexity tasks, and (3) training on mixed complexity datasets is associated with the greatest benefits in low data regimes, providing up to 5x sample efficiency versus training on easy tasks. These findings inspire future work on the development of data scaling laws for RLVR and the use of procedural data generators to further understand effective data development for efficient LLM fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。