arXiv:2604.19775cs.AIcs.CL2026-04中稿 · the Mechanistic In…被引 1

用统计方法解析大模型决策过程中的成败轨迹,让智能体行为可解释。

From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

论文配图:From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
图 1 · 摘自论文原文
  • 通过逐步奖励建模与置信区间预测,标记每一步的成败状态。
  • 在两个模拟环境中发现成功/失败概念可线性分离,结构清晰。
  • 适合研究可信自主智能体、模型调试与早期失效预警的人看。

大型语言模型(LLMs)正被部署为能在交互环境中进行推理、规划和行动的自主智能体。尽管其具备多步推理与决策能力,但其序列行为的内部机制仍不透明。本文提出一种基于分步置信区间的时序概念可解释性框架,将每一步的内部表示通过奖励建模与置信预测,统计标记为成功或失败。随后在这些表示上训练线性探测器,识别出对应任务成功、失败或推理偏移的潜在方向。在两个模拟交互环境ScienceWorld和AlfWorld上的实验表明,这些时序概念具有线性可分性,且与任务成功高度对齐。此外,初步结果表明可通过引导模型内识别出的成功方向提升智能体性能。该方法为基于LLM的智能体提供了可解释的早期失效检测与干预手段,推动复杂交互场景中可信自主语言模型的发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments. Despite their growing capability to perform multi-step reasoning and decision-making tasks, internal mechanisms guiding their sequential behavior remain opaque. This paper presents a framework for interpreting the temporal evolution of concepts in LLM agents through a step-wise conformal lens. We introduce the conformal interpretability framework for temporal tasks, which combines step-wise reward modeling with conformal prediction to statistically label model's internal representation at each step as successful or failing. Linear probes are then trained on these representations to identify directions of temporal concepts - latent directions in the model's activation space that correspond to consistent notions of success, failure or reasoning drift. Experimental results on two simulated interactive environments, namely ScienceWorld and AlfWorld, demonstrate that these temporal concepts are linearly separable, revealing interpretable structures aligned with task success. We further show preliminary results on improving an LLM agent's performance by leveraging the proposed framework for steering the identified successful directions inside the model. The proposed approach, thus, offers a principled method for early failure detection as well as intervention in LLM-based agents, paving the path towards trustworthy autonomous language models in complex interactive settings.

可解释性智能体置信预测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。