arXiv:2510.17923cs.LGcs.AI2025-10被引 3

让AI在无监督下自主学习推理,提升解题能力。

Rewarding the Journey, Not Just the Destination: A Composite Path and Answer Self-Scoring Reward Mechanism for Test-Time Reinforcement Learning

  • 通过双重校准答案与决策路径,自动生成可信奖励信号。
  • 在数学和代码任务中,性能显著优于现有自评分方法。
  • 适合希望模型持续自我进化、减少人工标注的开发者。

强化学习(RL)已成为推动大语言模型(LLMs)在数学与代码生成等复杂推理领域进步的关键范式。然而,现有方法严重依赖人工标注的偏好数据或标签数据进行奖励建模,面临可扩展性瓶颈。为突破此限制,我们探索基于无标签数据的测试时强化学习,使模型从连续经验流中自主学习。核心挑战在于无真实监督下的可靠奖励估计。现有方法如测试时强化学习通过自洽共识解决,但可能强化由多数投票产生的错误伪标签。为此,我们提出COMPASS(复合路径与答案自评分机制),一种无需外部监督的新颖测试时奖励机制。COMPASS融合两个互补组件:双校准答案奖励(DCAR),通过置信度与可信度校准建立可靠的伪标签;决定性路径奖励(DPR),直接优化推理过程质量,超越仅依赖结果的监督。通过联合强化可信共识答案与高确定性推理链,COMPASS系统性提升模型分析能力。大量实验表明,COMPASS在多种推理任务与模型架构上均实现显著且一致的性能提升,为大模型从持续经验中学习提供了更可扩展的方向。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has emerged as a powerful paradigm for advancing Large Language Models (LLMs), achieving remarkable performance in complex reasoning domains such as mathematics and code generation. However, current RL methods face a fundamental scalability bottleneck due to their heavy reliance on human-curated preference data or labeled datasets for reward modeling. To overcome this limitation, we explore RL on unlabeled data where models learn autonomously from continuous experience streams. The core challenge in this setting lies in reliable reward estimation without ground-truth supervision. Existing approaches like Test-Time RL address this through self-consistent consensus, but risk reinforcing incorrect pseudo-labels derived from majority voting. We introduce COMPASS (Composite Path and Answer Self-Scoring), a novel test-time reward mechanism that operates without external supervision. COMPASS integrates two complementary components: the Dual-Calibration Answer Reward (DCAR), which stabilizes training by establishing trustworthy pseudo-labels through confidence and credibility calibration, and the Decisive Path Reward (DPR), which directly optimizes the reasoning process quality beyond mere outcome supervision. By jointly reinforcing trustworthy consensus answers and highly decisive reasoning chains, the COMPASS systematically enhances the model's analytical capabilities. Extensive experiments show that COMPASS achieves significant and consistent performance gains across diverse reasoning tasks and model architectures, advancing a more scalable direction for LLMs to learn from continuous experience.

强化学习推理增强自评分无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。