通过可验证结果精调推理路径,提升大模型逻辑能力。
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning

- 对比成功与失败轨迹,识别关键策略片段的贡献度。
- 在多模型、多任务上显著提升推理准确率,最高增益达12.3%。
- 适合需要可信推理的AI系统,如智能助手与决策代理。
基于可验证奖励的强化学习(RLVR)已成为提升大语言模型推理能力的有效后训练范式。然而,现有方法通常仅依赖最终答案正确性分配轨迹级奖励,导致监督信号稀疏,且对所有词元一视同仁,忽略其实际推理贡献。尽管近期研究引入过程奖励、高熵词元和语义不确定性等中间信号,但这些信号往往不可直接验证,难以区分有益策略与有害模式。为此,我们提出STRIDE(基于判别性估计的战略轨迹推理),一种细粒度的RLVR框架,从可验证结果中提取战略推理监督信号。STRIDE在每组响应内对比成功与失败轨迹,估计每个n-gram策略模式的结果判别偏好,并结合推理显著性熵识别关键决策模式。这些模式在强化学习优化中被赋予差异化优势值,实现更精准的信用分配,同时保持RLVR的可验证性。大量实验表明,STRIDE在多种模型、任务及扩展场景(包括视觉语言模型与基于代理的系统)中均一致提升推理性能。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities of large language models. However, existing RLVR methods typically rely on final-answer correctness to assign trajectory-level rewards, providing sparse supervision and treating all tokens uniformly regardless of their actual contribution to reasoning. Although recent studies introduce intermediate signals such as process rewards, high-entropy tokens, and semantic uncertainty, these signals are often not inherently verifiable and may fail to distinguish beneficial strategic patterns from harmful ones. To address this limitation, we propose STRIDE (Strategic Trajectory Reasoning with Discriminative Estimation), a fine-grained RLVR framework that derives strategic reasoning supervision from verifiable outcomes. STRIDE contrasts successful and failed trajectories within each response group to estimate the outcome-discriminative preference of each $n$-gram strategic pattern, and further combines this signal with reasoning saliency entropy to identify decision-relevant strategic patterns. These patterns are assigned differentiated advantage values during RL optimization, enabling more precise credit assignment while preserving the verifiability of RLVR. Extensive experiments demonstrate that STRIDE consistently improves reasoning performance across diverse models, tasks, and extended settings, including VLMs and agent-based systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。