arXiv:2607.22797cs.LGcs.AI2026-07

让轴承故障诊断结果可物理验证,防止AI胡说八道。

Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis

  • 输出结构化证据:故障类别、特征频率、冲击时间定位
  • 预测频率误差仅6Hz,且能检测误判(AUROC达0.97)
  • 用语言模型生成报告时禁用虚构内容,可信度提升

在安全关键机械系统中部署AI诊断需确保可信性:预测结果能否在执行前与物理现实核验。当前智能诊断系统存在双重缺陷:标准输出仅为分类标签和置信度,无法与独立物理知识比对;而生成式语言模型用于维护报告时,更可能引入幻觉内容。以轴承故障诊断为场景,本文提出诊断证据网络(DENet),一种编码器无关的多任务框架,将输出扩展为结构化证据记录:故障分类、可与理论值比对的特征频率、以及原始波形上可检视的瞬态冲击时间定位。在四个编码器和三个公开数据集上,该证据输出未显著影响精度,1,024点段落中频率误差约6 Hz,此时谱估计本不适用。关键在于,预测频率与理论值的偏差构成无需标签、推理时可用的验证信号,在高置信度区域仍具区分力(AUROC达0.970和0.871)。此外,采用QLoRA微调的语言模型被限制仅翻译不生成内容,使未经证实声明率从10-12%降至2%,彻底消除虚构数值。

原文摘要 · Abstract (English)

Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers fail this standard in two ways. Their standard output, a class label with a softmax confidence score, is an internal statistic of the classifier, offering nothing checkable against independent physical knowledge; and the growing use of generative language models in maintenance reporting adds a second risk: hallucinated content entering reports on which decisions rest. Taking bearing fault diagnosis as the testbed, this work addresses both problems from the output side. The proposed Diagnostic Evidence Network (DENet) is an encoder-agnostic multi-task framework extending the output to a structured evidence record: the classification, a predicted characteristic frequency comparable against the theoretical value determined by bearing geometry and shaft speed, and a temporal localization of transient impulses inspectable on the raw waveform. Across four encoders and three public datasets, this evidence incurs no statistically significant accuracy cost, with a frequency error of about 6 Hz on 1,024-point segments where spectral estimation is structurally inapplicable. Centrally, the deviation between predicted and theoretical frequency constitutes a label-free, inference-time validation signal: it detects misclassifications with AUROC values of 0.970 and 0.871, and remains discriminative in the high-confidence regime where confidence-derived detectors are blind. Finally, a QLoRA-adapted language model is constrained to translate, but never generate, diagnostic content, reducing unsupported-claim rates from 10-12% to 2% and eliminating fabricated quantities.

故障诊断可解释性大模型应用物理验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。