arXiv:2602.15332cs.LG2026-02被引 1

定位推理模型中的关键决策点,揭示哪些上下文真正影响了推理方向。

Directional Reasoning Trajectory Change (DRTC): Identifying Critical Trace Segments in Reasoning Models

  • 通过不确定性和分布变化信号识别推理转折点
  • 干预转折点后发现影响集中,前5%片段贡献23%-28%影响力
  • 适合研究模型决策机制或提升推理可解释性的研究人员

理解语言模型如何执行长序列推理仍是开放挑战。现有可解释性方法多聚焦于与答案相关的词元,却很少揭示关键推理转折点的位置、其由早期上下文触发的因果关系,以及被突出内容是否真正引导了推理路径。本文提出方向性推理轨迹变化(DRTC),一种过程-因果方法:(i) 通过不确定性与分布偏移信号检测转折点;(ii) 在转折点实施接收端干预,在不重采样延续路径的前提下,仅阻断选定早期片段的信息流。DRTC衡量每项干预相对于实际推理方向的对数概率轨迹偏移,生成带符号的逐块归因;同时计算逻辑空间曲率变化及曲率特征作为补充几何诊断。在四个推理模型中,影响高度集中(基尼系数约0.50–0.58,前5%质量占比约0.23–0.28),学习得到的转折段比随机匹配段效果更强。在500题MATH基准上,使用R1-Distill-Qwen-1.5B模型的规模研究显示,学习段持续优于随机段(中位差值=0.409,500题中355题为正;p=2.3e-21),且曲率影响与DRTC在轨迹内共定位。与基于梯度和扰动的块归因对比表明,顶级DRTC块在嵌入插值编辑下对教师强制黄金答案对数概率的降低程度高于严格位置匹配的随机块(在稳定性过滤子集上)。总体而言,DRTC为特定上下文元素如何引导策略内推理轨迹提供了因果基础视角。

原文摘要 · Abstract (English)

Understanding how language models carry out long-horizon reasoning remains an open challenge. Existing interpretability methods often highlight tokens correlated with an answer, but rarely reveal where consequential reasoning turns occur, which earlier context triggers them under causal intervention, or whether highlighted text actually steers the rollout. We introduce Directional Reasoning Trajectory Change (DRTC), a process-causal method that (i) detects pivot decision points via uncertainty and distribution-shift signals and (ii) applies receiver-side interventions that preserve the realized continuation without resampling while blocking information flow from selected earlier chunks only at a pivot. DRTC measures how each intervention redirects the log-probability trajectory relative to the realized rollout direction, yielding signed per-chunk attributions; we also compute logit-space curvature changes and curvature signatures as a complementary geometric diagnostic. Across four reasoning models, influence is sharply concentrated (Gini approximately 0.50-0.58, top-5% mass approximately 0.23-0.28), and learned pivots induce stronger effects than matched random spans. In a 500-problem MATH scaling study with R1-Distill-Qwen-1.5B, learned spans continue to outperform matched random spans (median Delta=0.409, 355/500 positive; p=2.3e-21), and curvature-impact co-localizes with DRTC within traces as a diagnostic. We benchmark against gradient- and perturbation-based chunk attributions and show graded outcome linkage: under embedding-interpolation edits, top-ranked DRTC chunks reduce teacher-forced gold-answer log-probability more than strict position-matched random chunks on a stability-filtered subset. Overall, DRTC provides a causally grounded view of how specific context elements steer on-policy reasoning trajectories.

推理可解释性因果分析模型归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。