让AI像专家一样一步步排查企业IT故障,准确率超七成。
DQA: Diagnostic Question Answering for IT Support
- 用持续诊断状态跟踪问题,按根本原因聚合历史案例
- 准确率78.7%比基线高一倍,平均对话轮次从8.4降为3.9
- 适合需要高效、系统性解决复杂故障的IT支持场景
企业IT支持本质上是诊断过程:有效解决需通过模糊用户报告迭代收集证据以定位根本原因。尽管检索增强生成(RAG)能通过历史案例提供依据,但标准多轮RAG系统缺乏显式诊断状态,难以在多轮交互中累积证据并处理竞争性假设。我们提出DQA,一种维护持久诊断状态的问答框架,将检索到的案例按根本原因而非单个文档进行聚合。DQA结合对话查询重写、检索聚合和状态条件响应生成,在企业延迟与上下文约束下支持系统化排查。我们在150个匿名企业IT支持场景上采用回放协议评估,三次独立运行平均显示,DQA在轨迹级成功标准下达到78.7%成功率,远超多轮RAG基线的41.3%,同时平均对话轮次从8.4降至3.9。
原文摘要 · Abstract (English)
Enterprise IT support interactions are fundamentally diagnostic: effective resolution requires iterative evidence gathering from ambiguous user reports to identify an underlying root cause. While retrieval-augmented generation (RAG) provides grounding through historical cases, standard multi-turn RAG systems lack explicit diagnostic state and therefore struggle to accumulate evidence and resolve competing hypotheses across turns. We introduce DQA, a diagnostic question-answering framework that maintains persistent diagnostic state and aggregates retrieved cases at the level of root causes rather than individual documents. DQA combines conversational query rewriting, retrieval aggregation, and state-conditioned response generation to support systematic troubleshooting under enterprise latency and context constraints. We evaluate DQA on 150 anonymized enterprise IT support scenarios using a replay-based protocol. Averaged over three independent runs, DQA achieves a 78.7% success rate under a trajectory-level success criterion, compared to 41.3% for a multi-turn RAG baseline, while reducing average turns from 8.4 to 3.9.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。