不依赖标注数据,智能挑选最有价值的样本训练推理模型。
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

- 通过追踪推理过程中的元认知关键点,动态评估样本不确定性。
- 仅用29.3%标注数据,达到全监督模型性能,收敛速度提升2.75倍。
- 适合资源有限时高效训练大模型推理能力的研究者使用。
强化学习结合可验证奖励(RLVR)显著提升了大推理模型(LRMs)的表现,但需在大规模完全标注数据集上训练。现有方法主要从两方面优化:一是数据选择法,从已有标注数据中筛选少量“黄金”样本以接近全量数据效果,但依赖预标注池;二是无监督RLVR方法,利用模型自身内部信号在海量未标注数据上训练,但性能较差。为此,本文研究“暗中选样”场景下的RLVR,即在无先验监督下选择最有利于训练且值得标注的未标注样本。系统分析表明,智能选样依赖于校准良好的不确定性估计器,以实现数据的策略性分组和自适应训练。基于此,提出PivotTrace——一种三路数据筛选框架,利用注意力动态追踪推理过程中的元认知关键点。通过枢纽密度精准量化不确定性,实现自动数据分流,协同提升标注与训练效率。实验显示,PivotTrace仅需29.3%标注样本即可超越全监督大模型,且收敛速度提升2.75倍。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this end, data-efficient RLVR methods have been widely studied from two perspectives: (i) data selection methods identify a small subset of "golden" samples that yield near-full-data performance, but they rely on a pre-existing pool of labeled data. (ii) unsupervised RLVR methods train the model using its own internal supervision signals on large-scale unlabeled data, yet they exhibit suboptimal performance. Accordingly, we investigate the "pick in the dark" setup for RLVR, which aims to select, without prior supervision, unlabeled samples that are most beneficial for training and worthy of annotation. Through systematic analysis, we demonstrate that smart picks hinge on a well-calibrated uncertainty estimator to enable strategic partitioning of data for adaptive training regimes. Building on this insight, we propose PivotTrace, a three-way data triage framework that leverages attention dynamics to trace metacognitive pivots during reasoning. By precisely quantifying uncertainty through pivot density, PivotTrace achieves automated data routing to synergistically maximize both annotation and training efficiency. Empirically, PivotTrace surpasses the fully supervised LRM with only 29.3% annotated samples and 2.75 faster convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。