arXiv:2603.16500cs.LGcs.CL2026-03

通过渐进式优化置信度分布提升模型校准效果

From the Inside Out: Progressive Distribution Refinement for Confidence Calibration

  • 用模型内部置信度分布做奖励信号,逐步优化
  • 在多个模型和基准上显著提升校准性能
  • 有效缓解投票策略导致的奖励欺骗问题

利用强化学习中的模型内部信息作为自奖励信号,因其无需标签而受到广泛关注。尽管已有研究在测试时缩放(TTS)策略应用于强化学习方面取得进展,但测试与训练阶段内部信息差异仍缺乏有效处理。此外,基于投票的TTS策略在测试时训练中常出现奖励欺骗问题。为此,我们提出DistriTTRL,利用模型置信度的分布先验,在强化学习中逐步优化奖励信号,而非依赖单次查询的回溯。同时,通过面向多样性的惩罚机制,缓解由投票策略引发的一致性奖励欺骗现象。得益于模型能力与自奖励信号相互促进的训练机制,以及对奖励欺骗的有效抑制,DistriTTRL在多个模型和基准上均实现显著性能提升。

原文摘要 · Abstract (English)

Leveraging the model's internal information as the self-reward signal in Reinforcement Learning (RL) has received extensive attention due to its label-free nature. While prior works have made significant progress in applying the Test-Time Scaling (TTS) strategies to RL, the discrepancy in internal information between test and training remains inadequately addressed. Moreover, Test-Time Training based on voting-based TTS strategies often suffers from reward hacking problems. To address these issues, we propose DistriTTRL, which leverages the distribution prior of the model's confidence during RL to progressively optimize the reward signal, rather than relying solely on single-query rollouts. Additionally, we mitigate the phenomenon of consistent reward hacking caused by the voting-based TTS strategies through diversity-targeted penalties. Benefiting from this training mechanism where model capability and self-reward signals complement each other, and the mitigation of reward hacking, DistriTTRL has achieved significant performance improvements across multiple models and benchmarks.

置信度校准强化学习奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。