arXiv:2504.01373cs.LG2025-04被引 5

UniFault用统一方法处理轴承数据,实现跨场景故障诊断的高泛化能力。

UniFault: A Fault Diagnosis Foundation Model from Bearing Data

  • 将多变量数据转为标准单变量序列,融合跨域时间特征增强泛化。
  • 在超过690万样本上预训练,少样本下性能超越现有模型。
  • 适合工业界需要跨设备、少标注数据的智能维护场景。

机器故障诊断(FD)是预测性维护的关键任务,可实现早期故障检测并防止意外停机。尽管重要,现有FD模型通常针对特定工况,跨数据集泛化能力有限。基础模型(FM)在视觉和语言领域展现出强大泛化能力,即使少量数据也能实现少样本或零样本学习。但将其应用于FD面临独特挑战:相比图像和文本的大规模一致数据集,FD数据集通常较小且异构,采样频率和通道数差异大,难以设计通用架构以有效处理多样化数据并保持鲁棒特征提取。本文提出UniFault,一种面向故障诊断的基础模型,系统解决上述问题。模型包含两个关键创新的数据统一流程:一是将多变量输入转化为标准化单变量序列;二是提出新颖的跨域时间融合策略,缓解分布偏移,提升样本多样性和数量,增强模型在不同条件下的泛化能力。UniFault在超过690万条涵盖多种FD数据集的样本上进行预训练,实现卓越的少样本性能。大量实验证明,UniFault在真实世界FD数据集上达到当前最优表现,为故障诊断模型树立新基准,推动更可扩展、更鲁棒的预测性维护解决方案发展。

原文摘要 · Abstract (English)

Machine fault diagnosis (FD) is a critical task for predictive maintenance, enabling early fault detection and preventing unexpected failures. Despite its importance, existing FD models are operation-specific with limited generalization across diverse datasets. Foundation models (FM) have demonstrated remarkable potential in both visual and language domains, achieving impressive generalization capabilities even with minimal data through few-shot or zero-shot learning. However, translating these advances to FD presents unique hurdles. Unlike the large-scale, cohesive datasets available for images and text, FD datasets are typically smaller and more heterogeneous, with significant variations in sampling frequencies and the number of channels across different systems and applications. This heterogeneity complicates the design of a universal architecture capable of effectively processing such diverse data while maintaining robust feature extraction and learning capabilities. In this paper, we introduce UniFault, a foundation model for fault diagnosis that systematically addresses these issues. Specifically, the model incorporates a comprehensive data harmonization pipeline featuring two key innovations. First, a unification scheme transforms multivariate inputs into standardized univariate sequences. Second, a novel cross-domain temporal fusion strategy mitigates distribution shifts and enriches sample diversity and count, improving the model generalization across varying conditions. UniFault is pretrained on over 6.9 million samples spanning diverse FD datasets, enabling superior few-shot performance. Extensive experiments on real-world FD datasets demonstrate that UniFault achieves state-of-the-art performance, setting a new benchmark for fault diagnosis models and paving the way for more scalable and robust predictive maintenance solutions.

故障诊断基础模型轴承监测少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。