arXiv:2511.04032cs.AI2025-11被引 4

首次系统检测多智能体AI中的隐性故障,提供可复现的评测数据集。

Detecting Silent Failures in Multi-Agentic AI Trajectories

  • 构建轨迹异常检测新任务,捕捉非确定性导致的漂移与循环问题。
  • 创建两个基准数据集,分别含4275和894条轨迹,标注准确率达98%。
  • 适合关注AI可靠性、系统监控与安全评估的研究者使用。

多智能体AI系统由大语言模型(LLMs)驱动,具有内在非确定性,易发生漂移、循环和输出缺失等难以察觉的隐性故障。本文提出在智能体轨迹中进行异常检测的新任务,并设计数据集构建流程,涵盖用户行为、智能体非确定性及LLM变异。基于该流程,我们构建并标注了两个基准数据集,包含4,275和894条多智能体轨迹。在这些数据集上对异常检测方法进行基准测试,发现监督学习(XGBoost)与半监督学习(SVDD)方法表现相当,准确率分别高达98%和96%。本工作首次系统研究了多智能体AI中的异常检测,提供了数据集、基准与研究洞见,为未来研究提供支持。

原文摘要 · Abstract (English)

Multi-Agentic AI systems, powered by large language models (LLMs), are inherently non-deterministic and prone to silent failures such as drift, cycles, and missing details in outputs, which are difficult to detect. We introduce the task of anomaly detection in agentic trajectories to identify these failures and present a dataset curation pipeline that captures user behavior, agent non-determinism, and LLM variation. Using this pipeline, we curate and label two benchmark datasets comprising \textbf{4,275 and 894} trajectories from Multi-Agentic AI systems. Benchmarking anomaly detection methods on these datasets, we show that supervised (XGBoost) and semi-supervised (SVDD) approaches perform comparably, achieving accuracies up to 98% and 96%, respectively. This work provides the first systematic study of anomaly detection in Multi-Agentic AI systems, offering datasets, benchmarks, and insights to guide future research.

多智能体异常检测LLM可靠性系统监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。