用不确定性一致性筛选更少但更有效的查询,降低数学推理强化学习的成本
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
- 基于主观与客观不确定性的对齐度设计查询选择策略
- 仅用30%数据即可达到全量数据性能,显著减少标注开销
- 适用于资源受限下的强化学习推理任务,尤其适合高成本标注场景
大语言模型通过可验证奖励的强化学习(RLVR)提升了数学推理能力。然而现有RLVR算法需大量查询,导致标注成本高昂。本文探究是否可通过更少但更具信息量的查询实现相当或更优性能,将主动学习(AL)引入RLVR。发现经典主动学习策略因忽略客观不确定性而表现不及随机采样。为此提出不确定性一致性度量,评估主观不确定性与客观不确定性的一致程度:离线场景下使用点二列相关系数(PBC)衡量;在线训练中因采样有限且输出分布动态变化,难以估计PBC,故提出基于归一化优势与主观不确定性的新变体。理论证明该在线变体与离线PBC严格负相关,支持更优样本选择。实验表明,该方法持续优于随机及经典主动学习基线,在仅使用30%数据情况下即达成全数据集性能,有效降低推理任务的RLVR成本。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require large query budgets, making annotation costly. We investigate whether fewer but more informative queries can yield similar or superior performance, introducing active learning (AL) into RLVR. We identify that classic AL sampling strategies fail to outperform random selection in this setting, due to ignoring objective uncertainty when only selecting by subjective uncertainty. This work proposes an uncertainty consistency metric to evaluate how well subjective uncertainty aligns with objective uncertainty. In the offline setting, this alignment is measured using the Point-Biserial Correlation Coefficient (PBC). For online training, because of limited sampling and dynamically shifting output distributions, PBC estimation is difficult. Therefore, we introduce a new online variant, computed from normalized advantage and subjective uncertainty. Theoretically, we prove that the online variant is strictly negatively correlated with offline PBC and supports better sample selection. Experiments show our method consistently outperforms random and classic AL baselines, achieving full-dataset performance while training on only 30% of the data, effectively reducing the cost of RLVR for reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。