用几何分布建模提升少样本强化学习的推理能力
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling

- 通过建模标注数据的全局特征分布,识别正确与错误推理路径差异
- 仅用10%标注数据,性能超越全监督模型,提升4.1%
- 适合标注成本高、需高效利用无标签数据的LLM推理场景
基于可验证奖励的强化学习(RLVR)显著提升了大语言模型的推理能力,但面临标注成本高与无监督方法易崩溃的两难。现有半监督方法虽以少量标注数据引导无标注数据,仍因依赖粗糙性能启发式而存在严重数据效率瓶颈,大量有效样本未被充分利用。为此,我们提出GeoMin,通过在标注数据上建模全局特征分布,揭示正确与错误轨迹间的结构差异,建立可靠的自奖励信号可信度评估先验,充分释放无标注数据潜力。实验证明,GeoMin相比最强基线提升4.1%,仅用10%标注数据即超越全监督模型,展现卓越数据效率。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation costs, while unsupervised alternatives suffer from severe model collapse. Recent semi-supervised RLVR methods address this by using a small labeled set to guide unlabeled data, achieving a promising trade-off between training efficacy and annotation cost. However, they suffer from a severe data-efficiency bottleneck due to the reliance on coarse performance heuristics, leaving a vast majority of valuable instances underutilized. To this end, we propose GeoMin, which models global feature distributions on labeled data to decode the structural discrepancy between correct and incorrect rollouts, thereby establishing a robust prior to assess the reliability of self-reward signals and fully unleash the potential of unlabeled data. Empirically, GeoMin outperforms the strongest baselines by +4.1% and even surpasses fully supervised models with only 10% of the annotations, demonstrating remarkable data efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。