arXiv:2606.11172cs.LG2026-06

用未来行为预测提升大模型推理控制效果,几乎不损失输出质量。

Predicting Future Behaviors in Reasoning Models Enables Better Steering

论文配图:Predicting Future Behaviors in Reasoning Models Enables Better Steering
图 1 · 摘自论文原文
  • 通过中间推理步骤预测未来行为,而非依赖已生成文本的检测特征。
  • 预测准确率达64%-91%,显著优于传统方法的干预效果。
  • 适合需要精准控制且保持输出质量的研究者与应用开发者。

部署的大规模推理模型(LRM)常表现出意料之外的行为。测试时的控制方法通过干预隐藏表示来调整输出,但可能降低输出质量。我们指出,以往的控制方法隐含依赖于从已生成文本中检测行为的内部特征,而这些特征对未来的行為结果预测能力较差,不适合作为干预目标。相反,我们训练激活探针,从中间推理步骤预测未来行为的可能性。这些探针对最可能行为的预测准确率为64%至91%,揭示了一类独立的内部预测特征。基于此类预测特征,我们提出一种文本级控制方法——未来探针控制生成(FPCG)。FPCG 采样多个候选句子,并根据预测未来行为可能性的探针选择最优句。该方法实现控制的同时几乎不降低输出质量,并在激活控制失效的多个评估中仍有效。结果表明,区分检测与预测特征可带来更精细的模型行为控制策略。

原文摘要 · Abstract (English)

Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output quality. We argue that prior steering work implicitly relies on internal features that detect behavior in already generated text. We show that these detection features are poor predictors of future behavioral outcomes, and thus not the natural intervention target. Instead, we train activation probes to predict future behavior likelihoods from intermediate reasoning steps. These probes predict the most likely behavior with 64%-91% accuracy, revealing a separate type of internal prediction features. Building on these prediction features, we introduce a text-level steering method, Future Probe Controlled Generation. FPCG samples multiple candidate sentences and chooses the best one according to a probe predicting the future behavior likelihood. This enables steering with almost no output quality degradation. FPCG also enables steering in several evaluations where activation steering fails. These results show that distinguishing detection and prediction features enables a more nuanced approach to controlling LRM behaviors.

大模型控制推理优化行为预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。