通过分析推理过程中的隐藏状态轨迹,提升对大模型最终行为的预测能力。
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

- 构建推理过程中概念概率的动态轨迹,捕捉思维演变过程。
- 轨迹特征使未来行为区分度显著提升,最高达95% AUROC。
- 适合关注模型安全监控与可解释性的研究者使用。
大型推理模型(LRM)通过其思维链(CoT)推理为安全监控提供了新机会。然而,CoT并不总能忠实反映模型最终输出,削弱了其作为监控工具的可靠性。为此,我们研究了LRM的隐藏表示,探究能否从提示和CoT表示中预测未来行为。通过在每个生成标记上评估探测器,我们构建了探测轨迹,即概念概率在推理过程中的连续演化。发现从完整轨迹而非单一静态预测中分析未来行为更具可区分性。为刻画这些时序动态,我们提取了捕捉波动性、趋势和稳态行为的信号处理特征,显著提升了未来模型状态的分离效果。还提出两个方法洞察:首先,基于模板的训练数据可达到与动态生成模型响应相媲美的效果,无需昂贵的初始推理与标注;其次,池化操作选择至关重要:平均池化与最后标记方法退化至近随机性能,而最大池化达到最高95% AUROC,并产生稳定探测轨迹。在四个数据集和四种跨安全与数学领域的推理模型上验证,轨迹特征编码了任务特异性动态,增强结果可分性。这些发现确立探测轨迹作为监控LRM行为的互补框架。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's final output, undermining its reliability as a monitoring tool. To address this, we investigate the hidden representations of LRMs to determine whether future behavior can be predicted from prompt and CoT representations. By evaluating a probe at each generated token, we construct a probe trajectory, the continuous evolution of a concept's probability across the reasoning process. We find that future model behavior is more distinguishable when examined over the full trajectory than from a single static prediction. To characterize these temporal dynamics, we extract signal-processing features that capture volatility, trend, and steady-state behavior, significantly improving the separation of future model states. We also present two methodological insights. First, template-based training data achieves near-parity with dynamically generated model responses, eliminating the need for a costly initial inference and labeling. Second, the choice of pooling operation is critical: average-pooling and last-token methods collapse to near-random performance, while max-pooling achieves up to 95% AUROC and yields stable probe trajectories. Using four datasets and four reasoning models across the domains of safety and mathematics, we demonstrate that trajectory features encode task-specific dynamics that improve outcome separability. These findings establish probe trajectories as a complementary framework for monitoring LRM behavior. Warning: This article contains potentially harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。