arXiv:2511.20663cs.MAcs.AI2025-11被引 2

提出量化多智能体系统认知恢复延迟的新指标MTTR-A。

MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems

  • 引入MTTR-A指标,衡量智能体系统在推理偏离后的恢复时间。
  • 实验证明不同恢复策略下,平均恢复时间可测且存在差异。
  • 适合关注大模型系统可靠性与稳定性研究的开发者。

基于大语言模型的多智能体系统(MAS)的可靠性正日益受认知故障而非基础设施故障的制约。现有可观测工具虽能描述故障,却无法量化分布式推理在失去一致性后恢复所需时间。本文提出MTTR-A(智能体系统平均恢复时间),一个运行时可靠性指标,用于测量多智能体系统中认知恢复的延迟。该指标将经典可靠度理论应用于智能体编排,捕捉检测推理漂移并恢复一致运行所需的时间。我们进一步定义了互补指标,包括MTBF和归一化恢复率(NRR),并建立理论边界,将恢复延迟与长期认知可用性关联。通过基于LangGraph的基准测试,模拟推理漂移与反射式恢复,实验验证了多种反射策略下的可测量恢复行为。本工作为分布式智能体系统的运行时认知可靠性建立了量化基础。

原文摘要 · Abstract (English)

Reliability in multi-agent systems (MAS) built on large language models is increasingly limited by cognitive failures rather than infrastructure faults. Existing observability tools describe failures but do not quantify how quickly distributed reasoning recovers once coherence is lost. We introduce MTTR-A (Mean Time-to-Recovery for Agentic Systems), a runtime reliability metric that measures cognitive recovery latency in MAS. MTTR-A adapts classical dependability theory to agentic orchestration, capturing the time required to detect reasoning drift and restore coherent operation. We further define complementary metrics, including MTBF and a normalized recovery ratio (NRR), and establish theoretical bounds linking recovery latency to long-run cognitive uptime. Using a LangGraph-based benchmark with simulated drift and reflex recovery, we empirically demonstrate measurable recovery behavior across multiple reflex strategies. This work establishes a quantitative foundation for runtime cognitive dependability in distributed agentic systems.

多智能体可靠性大模型评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。