arXiv:2601.15075cs.AIcs.CL2026-01被引 7

揭秘智能体行为背后的内在动因,提升系统可解释性与问责能力。

The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution

  • 分层级分析:先定位关键交互步骤,再精确定位驱动文本
  • 在多种任务中精准识别影响决策的历史事件与句子
  • 适合关注AI可解释性、安全治理的研究者与开发者

基于大语言模型的智能体广泛应用于客服、网页导航和软件工程等实际场景。随着其自主性增强和规模化部署,理解智能体采取特定行为的原因对问责与治理至关重要。现有研究多聚焦于失败归因,难以解释行为背后的真实动因。为此,我们提出通用智能体归因框架,旨在识别无论任务成败都驱动行为的内部因素。该框架分层运作:在组件层面,利用时序似然动态识别关键交互步骤;在句子层面,通过扰动分析精确定位具体文本证据。我们在包括标准工具使用及记忆偏差等隐蔽可靠性风险在内的多样化场景中验证了该框架,实验表明其能可靠定位影响智能体行为的关键历史事件与语句,为构建更安全、可问责的智能体系统迈出关键一步。代码已开源。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents are widely used in real-world applications such as customer service, web navigation, and software engineering. As these systems become more autonomous and are deployed at scale, understanding why an agent takes a particular action becomes increasingly important for accountability and governance. However, existing research predominantly focuses on \textit{failure attribution} to localize explicit errors in unsuccessful trajectories, which is insufficient for explaining \textbf{the reason behind agent behaviors}. To bridge this gap, we propose a novel framework for \textbf{general agentic attribution}, designed to identify the internal factors driving agent actions regardless of the task outcome. Our framework operates hierarchically to manage the complexity of agent interactions. Specifically, at the \textit{component level}, we employ temporal likelihood dynamics to identify critical interaction steps; then at the \textit{sentence level}, we refine this localization using perturbation-based analysis to isolate the specific textual evidence. We validate our framework across a diverse suite of agentic scenarios, including standard tool use and subtle reliability risks like memory-induced bias. Experimental results demonstrate that the proposed framework reliably pinpoints pivotal historical events and sentences behind the agent behavior, offering a critical step toward safer and more accountable agentic systems. Codes are available at https://github.com/AI45Lab/AgentDoG.

智能体可解释性归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。