arXiv:2603.01641cs.AI2026-03被引 1

通过可控轨迹引导,让大模型学会多样化的复杂推理路径。

Learning Structured Reasoning via Tractable Trajectory Control

  • 用可控轨迹控制框架主动引导推理过程,探索多样化模式。
  • 在数学推理任务上,多模型实现稳定提升,验证了方法有效性。
  • 适合希望提升模型推理多样性与可靠性的研究人员使用。

大型语言模型可表现出涌现式推理行为,常以重复的词汇模式(如“wait”表示验证)体现。然而,在无约束采样下,复杂推理路径仍稀疏,标准强化学习难以保证多样推理行为的获取。本文提出结构化推理范式,通过在强化学习过程中有目标地探索特定推理模式,系统发现并强化多样推理模式。为此,我们提出 Ctrl-R 框架,基于可计算的轨迹控制,主动引导生成过程,激励探索对复杂问题求解至关重要的多样化推理路径。所获行为策略支持准确的重要性采样估计,实现无偏的在线优化。我们进一步引入重要性权重的幂缩放因子,使策略能选择性学习探索性、分布外的轨迹,同时保持优化稳定。实验表明,Ctrl-R 能有效探索并内化此前无法获得的推理模式,在数学推理任务中显著提升语言与视觉-语言模型性能。

原文摘要 · Abstract (English)

Large language models can exhibit emergent reasoning behaviors, often manifested as recurring lexical patterns (e.g., "wait," indicating verification). However, complex reasoning trajectories remain sparse in unconstrained sampling, and standard RL often fails to guarantee the acquisition of diverse reasoning behaviors. We propose a systematic discovery and reinforcement of diverse reasoning patterns through structured reasoning, a paradigm that requires targeted exploration of specific reasoning patterns during the RL process. To this end, we propose Ctrl-R, a framework for learning structured reasoning via tractable trajectory control that actively guides the rollout process, incentivizing the exploration of diverse reasoning patterns that are critical for complex problem-solving. The resulting behavior policy enables accurate importance-sampling estimation, supporting unbiased on-policy optimization. We further introduce a power-scaling factor on the importance-sampling weights, allowing the policy to selectively learn from exploratory, out-of-distribution trajectories while maintaining stable optimization. Experiments demonstrate that Ctrl-R enables effective exploration and internalization of previously unattainable reasoning patterns, yielding consistent improvements across language and vision-language models on mathematical reasoning tasks.

推理建模强化学习轨迹控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。