评测AI在真实通信网络中排查故障的能力,发现其诊断证据不足。
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

- 构建专家标注的电信故障排查基准CTBench,聚焦根因分析与路径恢复。
- 顶尖AI代理能识别路径终点但难以准确找出故障根因,尤其在链路层等场景表现差。
- 即使答案正确,多数代理无法提供操作所需的证据支撑,不适合实际运维使用。
AI代理正被考虑用于自动化网络运维,工程师需诊断故障、优化配置以提升服务并降低成本,同时受限于严格约束。然而现有评估未能真实反映网络特征,也未在部分可观测的多厂商、多设备、多协议环境中测试代理能力。本文提出公开基准CTBench,用于评估代理是否具备专业电信故障排查能力。该基准聚焦根因分析与路径恢复,任务由专家构建并标注丰富元数据(包括黄金证据步骤)。采用专家定义的指标,评估最终答案及诊断证据。实验表明,当前先进代理在路径恢复任务中可良好识别端点,但在根因分析上普遍表现不佳,尤其在接口状态、链路层、服务管理等故障上。更重要的是,即使代理给出合理或正确的最终答案,往往缺乏实际运维所需的证据支持。结果还显示,路径恢复通常资源开销更高,但更大资源投入并不必然带来更好诊断效果。
原文摘要 · Abstract (English)
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。