arXiv:2602.22297cs.LGcs.AI2026-02中稿 · be published in AA…被引 4

用对抗逆强化学习自动学故障评分,无需人工标注标签。

Learning Rewards, Not Labels: Adversarial Inverse Reinforcement Learning for Machinery Fault Detection

  • 从正常运行数据中学习奖励函数,替代人工设计
  • 在三个数据集上对正常样本评分低、故障样本评分高
  • 适合工业场景下无标签故障检测,提升早期预警能力

强化学习(RL)为机械设备故障检测(MFD)提供了巨大潜力。然而,现有基于RL的MFD方法大多未能充分发挥其序列决策优势,常将问题简化为上下文无关的猜谜游戏(上下文老虎机)。为此,本文将MFD建模为离线逆强化学习问题,让智能体直接从健康运行序列中学习奖励动态,从而避免手动奖励设计和故障标签依赖。框架采用对抗逆强化学习训练一个判别器,用于区分正常(专家)与策略生成的状态转移。判别器所学得的奖励作为异常分数,反映偏离正常行为的程度。在三个运行至失效基准数据集(HUMS2023、IMS、XJTU-SY)上评估,模型对正常样本赋予低异常分,对故障样本赋予高异常分,实现早期且稳健的故障检测。通过将强化学习的序列推理与故障检测的时间结构对齐,本工作为数据驱动工业环境中基于强化学习的诊断开辟了新路径。

原文摘要 · Abstract (English)

Reinforcement learning (RL) offers significant promise for machinery fault detection (MFD). However, most existing RL-based MFD approaches do not fully exploit RL's sequential decision-making strengths, often treating MFD as a simple guessing game (Contextual Bandits). To bridge this gap, we formulate MFD as an offline inverse reinforcement learning problem, where the agent learns the reward dynamics directly from healthy operational sequences, thereby bypassing the need for manual reward engineering and fault labels. Our framework employs Adversarial Inverse Reinforcement Learning to train a discriminator that distinguishes between normal (expert) and policy-generated transitions. The discriminator's learned reward serves as an anomaly score, indicating deviations from normal operating behaviour. When evaluated on three run-to-failure benchmark datasets (HUMS2023, IMS, and XJTU-SY), the model consistently assigns low anomaly scores to normal samples and high scores to faulty ones, enabling early and robust fault detection. By aligning RL's sequential reasoning with MFD's temporal structure, this work opens a path toward RL-based diagnostics in data-driven industrial settings.

故障检测逆强化学习无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。