arXiv:2508.02490cs.AI2025-08被引 3

为健康与故障管理领域打造专用大模型评估框架。

PHM-Bench: A Domain-Specific Benchmarking Framework for Systematic Evaluation of Large Models in Prognostics and Health Management

  • 构建三维度评估体系,覆盖能力、任务与生命周期。
  • 支持故障诊断、剩余寿命预测等多任务精准评测。
  • 适合工业界推进通用模型向专业模型转化。

随着生成式人工智能的快速发展,大型语言模型(LLMs)在工业领域应用日益广泛,为健康与故障管理(PHM)带来新机遇,有助于解决开发成本高、部署周期长、泛化能力弱等问题。然而,现有评估方法在结构完整性、维度全面性和评估粒度方面仍显不足,制约了LLMs在PHM领域的深度集成。为此,本文提出PHM-Bench——一个面向PHM的大模型评估基准框架。该框架基于基础能力、核心任务和全生命周期的三维结构,针对PHM系统工程特点设计,涵盖知识理解、算法生成、任务优化等多层次指标,适配条件监测、故障诊断、剩余使用寿命(RUL)预测与维护决策等典型任务。通过自建案例集与公开工业数据集,实现对通用模型与领域特定模型的多维评测。PHM-Bench为大模型在PHM中的大规模评估奠定方法基础,推动通用模型向专业化模型演进。

原文摘要 · Abstract (English)

With the rapid advancement of generative artificial intelligence, large language models (LLMs) are increasingly adopted in industrial domains, offering new opportunities for Prognostics and Health Management (PHM). These models help address challenges such as high development costs, long deployment cycles, and limited generalizability. However, despite the growing synergy between PHM and LLMs, existing evaluation methodologies often fall short in structural completeness, dimensional comprehensiveness, and evaluation granularity. This hampers the in-depth integration of LLMs into the PHM domain. To address these limitations, this study proposes PHM-Bench, a novel three-dimensional evaluation framework for PHM-oriented large models. Grounded in the triadic structure of fundamental capability, core task, and entire lifecycle, PHM-Bench is tailored to the unique demands of PHM system engineering. It defines multi-level evaluation metrics spanning knowledge comprehension, algorithmic generation, and task optimization. These metrics align with typical PHM tasks, including condition monitoring, fault diagnosis, RUL prediction, and maintenance decision-making. Utilizing both curated case sets and publicly available industrial datasets, our study enables multi-dimensional evaluation of general-purpose and domain-specific models across diverse PHM tasks. PHM-Bench establishes a methodological foundation for large-scale assessment of LLMs in PHM and offers a critical benchmark to guide the transition from general-purpose to PHM-specialized models.

PHM大模型评估故障诊断生命周期

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。