生成特定故障概率的轴承振动信号,解决中间概率样本稀缺问题。
Generating Bearing Vibration Signals at User-Specified Fault Probabilities Using PR-GAN and Counterfactual Methods

- 用集成分类器作为可微分概率基准,指导生成目标概率信号。
- 反事实方法误差仅0.005-0.008,成功率100%,改动更小。
- 适合需要精确决策边界分析的工业故障诊断研究者。
轴承振动数据集中,多数样本的故障概率预测值接近0或1,而中间概率(灰区)样本稀少。这类边界样本对维护决策和研究决策边界至关重要。为解决该问题,本文提出并比较两种方法:生成故障概率恰好为0.25、0.50或0.75的振动信号。采用异构集成分类器平均输出作为固定、可微的概率基准。第一种为基于训练的PR-GAN方法,扩展WGAN-GP,通过残差生成器修改真实信号,推动分类器输出向目标概率逼近;第二种为免训练的逐样本反事实(CF)方法,直接优化输入信号以达到目标概率,同时保持与原信号接近。在CWRU和Paderborn数据集上评估,使用均方绝对概率误差、时域总变差及频域对数功率谱密度差异。所有设置下,CF方法平均概率误差0.005–0.008,保留样本成功率1.000;而PR-GAN误差为0.046–0.059,成功率0.501–0.680。因此CF更可靠且改动更小,但PR-GAN在多数场景下运行时间更低。
原文摘要 · Abstract (English)
In bearing vibration datasets, most samples receive predicted fault probabilities close to 0 or 1, while samples with intermediate (gray-zone) probabilities are rare. Such borderline samples are important because they reflect conditions in which maintenance decisions may require additional inspection or a conservative response and are useful for studying decision boundaries. To address this scarcity, this paper proposes and compares two approaches that generate vibration signals whose predicted fault probability matches a target probability of 0.25, 0.50, or 0.75. We use the average output of a heterogeneous ensemble classifier with different architectures and random initializations as a fixed, gradient-accessible probability oracle. The first, training-based approach, Probability-Regularized Generative Adversarial Network (PR-GAN), extends Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP) and edits a real signal through a residual generator while pushing the classifier output toward the target probability. The second is a training-free, per-sample Wachter-style counterfactual (CF) procedure that directly optimizes each input signal to reach the target probability while remaining close to the source signal. We evaluate both methods on the Case Western Reserve University (CWRU) and Paderborn bearing datasets using mean absolute target-probability error, time-domain total variation, and frequency-domain log power spectral density (log-PSD) differences. Across all settings, CF reaches the target with a mean absolute probability error of 0.005-0.008 and a within-tolerance success rate of 1.000 on retained samples, whereas PR-GAN's mean error is 0.046-0.059 with success rates between 0.501 and 0.680. CF therefore steers the probability more reliably and requires smaller average L1 changes, whereas PR-GAN has a lower reported runtime in most settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。