arXiv:2605.27934cs.CL2026-05

无需领域验证器,用答案概率优化推理过程

GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization

论文配图:GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization
图 1 · 摘自论文原文
  • 用答案似然值指导推理路径优化,实现逐标记评估
  • 在11个基准上平均表现最佳,数学与理工类任务显著提升
  • 适合需要精细反馈的模型训练场景,如复杂推理系统

基于可验证奖励的强化学习能提升语言模型推理能力,但依赖领域特定验证器、稀疏结果奖励和粗粒度信用分配,限制了其应用范围。我们提出GeneralThinker,一种在线策略框架,将推理监督重构为密集的答案条件优化,无需领域特定验证器即可实现响应级评估和标记级信用分配。GeneralThinker通过真实答案的似然值评估生成的推理轨迹,并推导出细粒度的标记兼容信号以实现精确信用分配。为稳定优化过程,采用裁剪和方向保持调制约束标记级更新。在涵盖数学、理工及通用推理的11个基准上,GeneralThinker实现了最佳平均性能。进一步分析表明,不受控的标记级调制会破坏训练稳定性,而受控调制则使细粒度信用分配始终有效。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards improves language model reasoning, but its reliance on domain-specific verifiers, sparse outcome rewards, and coarse-grained credit assignment limits its applicability. We introduce GeneralThinker, an on-policy framework that reformulates reasoning supervision as dense answer-conditioned optimization, enabling response-level evaluation and token-level credit assignment without domain-specific verifiers. GeneralThinker evaluates generated reasoning trajectories using the likelihood of the ground-truth answer and derives token-wise compatibility signals for fine-grained credit assignment. To stabilize optimization, it constrains token-level updates through clipping and direction-preserving modulation. Across 11 benchmarks spanning mathematics, STEM, and general reasoning, GeneralThinker achieves the best average performance. Further analyses show that uncontrolled token-level modulation can destabilize training, whereas controlled modulation makes fine-grained credit assignment consistently effective.

推理增强强化学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。