arXiv:2603.22978cs.AI2026-03

用文字重构故障树,让大模型能对话式排查系统故障。

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

  • 将图像故障树转为文本表示,支持大模型直接理解
  • 构建含3130条记录、平均每条40.75轮对话的评测基准
  • 模拟用户错误并测试模型任务追踪与纠错能力

在复杂系统维护中,故障树用于定位问题并提供针对性解决方案。为使存储为图像的故障树可被大语言模型直接处理,从而辅助故障追踪与分析,我们提出一种新的故障树文本表示方法。基于此,我们构建了一个强调复杂环境下鲁棒交互的多轮对话评测基准,用于评估模型在故障定位方面的表现,该基准包含3130个条目,平均每个条目40.75轮对话。我们训练了一个端到端模型以生成模糊信息来反映用户行为,并引入长距离回滚与恢复机制,模拟用户出错场景,从而评估模型在任务追踪与错误恢复方面的综合能力。实验表明,Gemini 2.5 Pro表现最佳。

原文摘要 · Abstract (English)

In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by large language models, which can assist in tracking and analyzing malfunctions, we propose a novel textual representation of fault trees. Building on it, we construct a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments, evaluating a model's ability to assist in malfunction localization, which contains $3130$ entries and $40.75$ turns per entry on average. We train an end-to-end model to generate vague information to reflect user behavior and introduce long-range rollback and recovery procedures to simulate user error scenarios, enabling assessment of a model's integrated capabilities in task tracking and error recovery, and Gemini 2.5 pro archives the best performance.

故障诊断大模型评测对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。