arXiv:2601.08310cs.LGcs.AI2026-01被引 4

让大模型按需选择推理深度,灵活控制算力与准确率的平衡。

ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning

  • 通过多阶段强化学习发现不同算力下的最优推理策略
  • 在各推理模式下均保持高推理密度,且模式间分离清晰
  • 将多种策略融合为单一模型,适合部署场景动态调整

近期的大规模推理模型(LRMs)通过长链式思维(CoT)推理获得优异表现,但在推理时统一使用过长的推理过程会带来巨大且常不必要的计算开销。以往方法尝试从输入推断合适的推理预算,但这类方法在最坏情况下不可靠,因最小推理需求估计本身极难,且训练中隐式固定了推理成本与准确率的权衡,限制了部署场景的灵活性。为此,我们提出ORBIT:一种可控制的多预算推理框架,根据输入触发不同推理模式。ORBIT采用多阶段强化学习,在每种努力水平下发现帕累托最优推理行为,并通过在线策略蒸馏将这些行为融合为单一统一模型。实验表明,ORBIT实现了(1)多模式可控推理行为,(2)各模式内具有竞争力的推理密度,(3)将前沿策略集成到单一学生模型中,同时保持清晰的模式分离和高单模式性能。

原文摘要 · Abstract (English)

Recent Large Reasoning Models (LRMs) achieve strong performance by leveraging long-form Chain-of-Thought (CoT) reasoning, but uniformly applying overlong reasoning at inference time incurs substantial and often unnecessary computational cost. To address this, prior work explores various strategies to infer an appropriate reasoning budget from the input. However, such approaches are unreliable in the worst case, as estimating the minimal required reasoning effort is fundamentally difficult, and they implicitly fix the trade-off between reasoning cost and accuracy during training, limiting flexibility under varying deployment scenarios. Motivated by these limitations, we propose ORBIT, a controllable multi-budget reasoning framework with well-separated reasoning modes triggered by input. ORBIT employs multi-stage reinforcement learning to discover Pareto-optimal reasoning behaviors at each effort, followed by on-policy distillation to fuse these behaviors into a single unified model. Experiments show that ORBIT achieves (1) controllable reasoning behavior over multiple modes, (2) competitive reasoning density within each mode, and (3) integration of these frontier policies into a single unified student model while preserving clear mode separation and high per-mode performance.

推理优化强化学习多预算大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。