真实网络故障诊断中,错误工单如何影响AI助手判断?
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

- 构建200个场景的基准测试,涵盖真故障与误报等复杂情况
- 发现三款AI代理在错误工单下诊断准确率骤降,甚至误判正常状态为故障
- 强调工单表述方式比内容真假更影响诊断结果,适合运维AI研究者
基于大模型的故障诊断代理日益增多,但现有基准仅评估其在准确工单和有故障条件下的表现,这在现实中很少成立。我们提出FaulT-Bench,包含200个跨八种网络拓扑的排障场景,其中五种来自公开实践实验室复现,覆盖真实故障、误报、设备归属错误及根因误判。为分离工单措辞的影响,我们将72个虚假前提工单重写为五种不同报告人角色,逐项改变报告人信心和可验证细节,保持网络状态不变。自动化测试框架在Kathará中部署每个场景,通过NIKA工具接口让代理交互,并使用大模型评判器从结果、修复方案和推理质量三方面评分。评估SADE、ReAct和Claude Code发现:三者在准确工单上表现接近饱和且对误导有一定鲁棒性,但在网络健康且工单错误时性能急剧下降,倾向于将无故障状态误判为根因,而非承认无问题。角色重写实验表明,工单表述方式比内容真假影响更大:自信的错误报告处理效果接近准确报告,而模糊不详的报告导致性能显著下滑。三款代理失败模式各异,从持续误诊到无响应,成本差异明显。这些结果凸显了FaulT-Bench作为真实网络排障中可靠推理代理开发基准的重要性。
原文摘要 · Abstract (English)
LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench, a benchmark of 200 troubleshooting scenarios across eight network topologies, five reimplemented from public practitioner labs, spanning genuine faults, false fault reports, incorrect device attribution, and incorrect root-cause claims. To isolate how ticket wording affects diagnosis, we further rewrite 72 false-premise tickets into five reporter personas that vary reporter confidence and verifiable detail one factor at a time, holding the network state fixed. Our automated harness deploys each scenario in Kathará, lets agents interact through the NIKA tool interface, and scores free-text diagnoses with an LLM judge across outcome, fix, and reasoning quality. Evaluating SADE, ReAct, and Claude Code, we find all three are near-saturated on accurate tickets and robust to misdirection, yet degrade sharply when the network is healthy and the ticket is wrong, probing until a benign condition can be promoted to a root cause rather than concluding nothing is wrong. Persona rewrites show that how a ticket is written matters more than what it claims: a confidently wrong report is handled about as well as an accurate one, while a vague, underspecified report degrades performance sharply. The three agents also fail differently, from constant over-diagnosis to unanswered runs, at very different cost. These results position FaulT-Bench as a benchmark for developing agentic systems that can reason reliably over the noisy, unreliable tickets of real-world network troubleshooting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。