arXiv:2602.02313cs.AIcs.CL2026-02被引 3

通过反向传播推理准确率,精准定位大模型的思维路径。

Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient

  • 用累积结果信号反向追踪模型推理过程,定位关键内部组件。
  • 在多个推理模型上实现更精确的行为定位与可调控制。
  • 适合研究模型可解释性及可控生成的开发者与研究人员。

大型语言模型在解决复杂现实问题时展现出强大的推理能力,但其内部驱动这些复杂推理行为的机制仍不透明。现有可解释性方法要么识别与特定文本模式相关的组件(如神经元),要么依赖人工标注的对比样本推导控制向量,因此难以精确定位复杂推理机制或捕捉模型内部运作对推理输出的序列影响。本文基于目标导向与序列影响感知原则,聚焦于识别对推理行为具有序列贡献且结果由长程效应累积的模型组件。我们提出集成策略梯度(IPG)框架,通过将推理后准确率等复合结果信号沿模型推理轨迹反向传播,归因于模型内部组件。实验表明,该方法实现了更精确的定位,并可在多种推理模型上可靠调控推理能力与推理强度。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong reasoning abilities in solving complex real-world problems. Yet, the internal mechanisms driving these complex reasoning behaviors remain opaque. Existing interpretability approaches targeting reasoning either identify components (e.g., neurons) correlated with special textual patterns, or rely on human-annotated contrastive pairs to derive control vectors. Consequently, current methods struggle to precisely localize complex reasoning mechanisms or capture sequential influence from model internal workings to the reasoning outputs. In this paper, built on outcome-oriented and sequential-influence-aware principles, we focus on identifying components that have sequential contribution to reasoning behavior where outcomes are cumulated by long-range effects. We propose Integrated Policy Gradient (IPG), a novel framework that attributes reasoning behaviors to model's inner components by propagating compound outcome-based signals such as post reasoning accuracy backward through model inference trajectories. Empirical evaluations demonstrate that our approach achieves more precise localization and enables reliable modulation of reasoning behaviors (e.g., reasoning capability, reasoning strength) across diverse reasoning models.

大模型可解释性推理控制策略梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。