arXiv:2505.20127cs.AI2025-05被引 8

通过轨迹分析发现LLM智能体行为差异,提升开发可观测性。

Agentic AI Process Observability: Discovering Behavioral Variability

  • 用过程与因果发现分析智能体执行轨迹
  • 区分有意与无意的行为波动,识别模糊设定
  • 适合需要调试复杂智能体系统的开发者

基于大语言模型(LLMs)的AI智能体正成为现代软件系统的核心组件。现有框架支持通过自然语言提示定义智能体的角色、目标和工具。然而,同一输入下智能体行为具有非确定性,亟需可靠的调试与可观测工具。本文探索将过程发现与因果发现应用于智能体执行轨迹,以增强开发者对行为变异性的监控与理解。同时结合基于LLM的静态分析技术,区分预期与非预期的行为差异。该方法有助于开发者掌控不断演进的规范,并识别出需要更精确定义的功能部分。

原文摘要 · Abstract (English)

AI agents that leverage Large Language Models (LLMs) are increasingly becoming core building blocks of modern software systems. A wide range of frameworks is now available to support the specification of such applications. These frameworks enable the definition of agent setups using natural language prompting, which specifies the roles, goals, and tools assigned to the various agents involved. Within such setups, agent behavior is non-deterministic for any given input, highlighting the critical need for robust debugging and observability tools. In this work, we explore the use of process and causal discovery applied to agent execution trajectories as a means of enhancing developer observability. This approach aids in monitoring and understanding the emergent variability in agent behavior. Additionally, we complement this with LLM-based static analysis techniques to distinguish between intended and unintended behavioral variability. We argue that such instrumentation is essential for giving developers greater control over evolving specifications and for identifying aspects of functionality that may require more precise and explicit definitions.

智能体可观测性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。