构建多数据集基准,评估大模型在微服务故障诊断中的推理能力。
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

- 基于因果证据的三维度评估:定位、识别、推理是否合理。
- 覆盖500+专家标注故障案例,涵盖资源、网络、应用等多类故障。
- 经6000+团队竞赛验证,适合作为真实环境下的智能运维评测标准。
基于大模型的智能体正重塑微服务运维为AgentOps,但现有基准多仅关注最终答案,忽略故障诊断中的系统性推理过程。本文提出AIOps2025和RCA100两个大规模数据集,采用推理过程评估范式,从定位(故障位置)、识别(故障类型)和理由(推理是否基于证据)三方面评估智能体诊断能力。数据集包含超过500个由领域专家标注的故障案例,覆盖HipsterShop与OpenTelemetry Demo Store两个典型微服务系统,涵盖资源、网络、运行时、中间件/数据库及应用逻辑等五类故障场景,并提供细粒度因果证据以支持学习与评估。数据集经大规模竞赛验证,累计超6000支团队参与,不仅具备专家标注质量,更构成可信赖的竞争性基准,适用于真实微服务环境中智能体故障诊断能力的评测。数据集已公开于https://www.aiops.cn/gitlab/aiops-live-benchmark/agenticopseval。
原文摘要 · Abstract (English)
LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data. However, existing benchmarks remain largely outcome-oriented: they score only the final answer and fail to assess the systematic reasoning process in failure diagnosis. We address this gap by introducing two large-scale datasets (AIOps2025 and RCA100) under a reasoning-process evaluation paradigm that assesses agentic diagnostic capability along three dimensions: Localization (where the fault occurs), Identification (what type of fault it is), and Reason (whether the reasoning trace is grounded in relevant evidence). Together, the two datasets comprise over 500 expert-labeled failure cases across two representative microservice systems (HipsterShop and the OpenTelemetry Demo Store). They cover diverse fault scenarios across resource, network, runtime, middleware/database, and application-logic categories and provide fine-grained causal evidence to support agent learning and reasoning-process evaluation. Beyond scale and coverage, the datasets have been carefully labelled by domain experts and validated through large-scale competitions, supporting more than 6,000 participating teams. This makes them not only expert-labeled diagnostic datasets, but also competition-validated benchmarks for evaluating agentic failure diagnosis in real-world microservice environments. Datasets are available at https://www.aiops.cn/gitlab/aiops-live-benchmark/agenticopseval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。