发现并量化了多智能体大模型长期互动中的行为退化问题。
Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions
- 提出三类退化现象:语义漂移、协同漂移和行为漂移。
- 设计ASI指标,从12个维度量化行为稳定性,精度下降超30%。
- 适合关注AI系统长期可靠性与安全性的研发人员阅读。
多智能体大语言模型系统在复杂任务分解与协作求解中展现出强大能力,但其长期行为稳定性尚未得到充分研究。本文提出‘智能体漂移’概念,指在长时间交互序列中,智能体行为、决策质量及跨代理一致性逐步退化的现象。我们构建了理论框架,识别出三种表现形式:语义漂移(偏离原始意图)、协同漂移(共识机制失效)和行为漂移(出现非预期策略)。提出智能体稳定性指数(ASI),一个涵盖响应一致性、工具使用模式、推理路径稳定性及代理间一致率等十二个维度的复合度量体系。通过仿真分析与理论建模,证明未受控的漂移会导致任务完成准确率显著下降,人工干预需求上升。提出三种缓解策略:周期性记忆固化、感知漂移的路由协议与自适应行为锚定。理论分析表明,这些方法可大幅减少漂移相关错误,同时保持系统吞吐量。本研究为生产环境中智能体系统的监控、测量与治理提供了基础方法,对企业部署可靠性和人工智能安全研究具有重要意义。
原文摘要 · Abstract (English)
Multi-agent Large Language Model (LLM) systems have emerged as powerful architectures for complex task decomposition and collaborative problem-solving. However, their long-term behavioral stability remains largely unexamined. This study introduces the concept of agent drift, defined as the progressive degradation of agent behavior, decision quality, and inter-agent coherence over extended interaction sequences. We present a comprehensive theoretical framework for understanding drift phenomena, proposing three distinct manifestations: semantic drift (progressive deviation from original intent), coordination drift (breakdown in multi-agent consensus mechanisms), and behavioral drift (emergence of unintended strategies). We introduce the Agent Stability Index (ASI), a novel composite metric framework for quantifying drift across twelve dimensions, including response consistency, tool usage patterns, reasoning pathway stability, and inter-agent agreement rates. Through simulation-based analysis and theoretical modeling, we demonstrate how unchecked agent drift can lead to substantial reductions in task completion accuracy and increased human intervention requirements. We propose three mitigation strategies: episodic memory consolidation, drift-aware routing protocols, and adaptive behavioral anchoring. Theoretical analysis suggests these approaches can significantly reduce drift-related errors while maintaining system throughput. This work establishes a foundational methodology for monitoring, measuring, and mitigating agent drift in production agentic AI systems, with direct implications for enterprise deployment reliability and AI safety research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。