用执行轨迹生成可审计的AI代理行为解释,提升透明度。
Explaining AI Agents Through Execution Traces

- 基于执行轨迹生成结构化报告与自然语言解释
- 能准确识别不合理行为和证据缺失,优于直接用大模型生成的解释
- 适用于不同架构、环境和任务,无需修改代理设计
AI代理在现实场景中日益普及,它们与外部工具交互并做出多步决策,缺乏人类监督,亟需可靠且可审计的行为解释。传统可解释AI方法难以提供此类系统所需的过程级透明性,促使研究转向专为代理设计的新范式。本文提出一种后验XAI框架,将代理的长序列执行轨迹转化为结构化报告和忠实于可观测行为的自然语言解释。该框架仅依赖执行轨迹,可适用于不同代理架构、环境和任务。跨多个基准和架构的人类与自动化评估表明,该框架能生成高质量、轨迹忠实的解释,可靠识别出不支持的主张、无依据的行为及证据缺口,显著优于简单的大型语言模型生成解释。
原文摘要 · Abstract (English)
AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short of providing the process-level transparency required for such interactive, multi-step systems, motivating a paradigm shift toward approaches specifically designed for AI Agents. To address this gap, we present a post-hoc XAI framework that transforms a lengthy agent's execution trace into a structured report and a faithful natural-language explanation explicitly grounded in its observable behavior. Because it relies solely on execution traces, the framework applies across different agent architectures, environments, and tasks. Human and automated evaluations across multiple benchmarks and architectures show that our framework produces high-quality, trace-faithful explanations while reliably identifying unsupported claims, unjustified actions, and evidence gaps, outperforming naive LLM-generated explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。