arXiv:2603.25955cs.LG2026-03

公开真实车辆发动机异常检测数据集,助力工业级故障预警模型研发。

EngineAD: A Real-World Vehicle Engine Anomaly Detection Dataset

  • 从25辆商用车收集6个月高分辨率传感器数据,构建多变量真实故障数据集。
  • 基于300步时间窗与8个主成分的预处理,实现故障模式精准识别。
  • 实验证明经典方法在该任务中表现优于深度学习,适合实际部署。

安全关键领域(如交通)的异常检测进展受限于缺乏大规模真实世界基准。为此,我们提出EngineAD,一个新型多变量数据集,包含25辆商用汽车在六个月内采集的高分辨率传感器遥测数据。不同于合成数据集,EngineAD包含专家标注的真实运行数据,可区分正常状态与早期发动机故障的细微迹象。我们将数据预处理为300个时间步长的8个主成分片段,并使用九种不同的单类异常检测模型建立初步基准。实验揭示车队间性能差异显著,凸显跨车辆泛化的挑战。同时,研究结果支持最新文献观点:在该分段评估中,简单经典方法(如K-Means和One-Class SVM)常表现出色,甚至优于深度学习方法。通过公开发布EngineAD,我们旨在为汽车行业提供一个真实、具有挑战性的资源,用于开发鲁棒且可现场部署的异常检测与预测解决方案。

原文摘要 · Abstract (English)

The progress of Anomaly Detection (AD) in safety-critical domains, such as transportation, is severely constrained by the lack of large-scale, real-world benchmarks. To address this, we introduce EngineAD, a novel, multivariate dataset comprising high-resolution sensor telemetry collected from a fleet of 25 commercial vehicles over a six-month period. Unlike synthetic datasets, EngineAD features authentic operational data labeled with expert annotations, distinguishing normal states from subtle indicators of incipient engine faults. We preprocess the data into $300$-timestep segments of $8$ principal components and establish an initial benchmark using nine diverse one-class anomaly detection models. Our experiments reveal significant performance variability across the vehicle fleet, underscoring the challenge of cross-vehicle generalization. Furthermore, our findings corroborate recent literature, showing that simple classical methods (e.g., K-Means and One-Class SVM) are often highly competitive with, or superior to, deep learning approaches in this segment-based evaluation. By publicly releasing EngineAD, we aim to provide a realistic, challenging resource for developing robust and field-deployable anomaly detection and anomaly prediction solutions for the automotive industry.

异常检测车辆故障真实数据集工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。