提出巡航阶段标准评估基准,提升发动机剩余寿命预测可比性。
CruiseBench: A Real-Flight-Aligned N-CMAPSS Benchmark for Engine RUL Prediction

- 从N-CMAPSS中提取巡航阶段数据,固定评估流程提升可复现性。
- TSMixer模型在基准上表现最优,平均RMSE为3.4±1.71。
- 适合需要可控对比的RUL模型研究者使用,尤其关注迁移学习。
剩余使用寿命(RUL)预测是制定维护计划的核心,能估计发动机安全运行时间。N-CMAPSS通过真实飞行轨迹模拟航空发动机的退化过程,并保留完整的飞行时间序列,而非仅周期级快照,增强了现实性。但这也导致数据量大增,退化信号与飞行工况变化交织,使预处理选择复杂,难以直接比较不同RUL模型性能。为此,本文提出CruiseBench,一个基于N-CMAPSS的巡航阶段RUL评估基准。引入CPM-N-CMAPSS(巡航期掩码),利用恒定高度法识别九个子数据集中的巡航周期区间。CruiseBench采用固定协议:以场景描述符和实测传感器为输入,排除虚拟传感器、健康参数及辅助元数据,保留原始分辨率窗口,并施加数据集级的RUL上限。实验采用LSTM、GRU、TCN和TSMixer模型,提供基线结果。在CruiseBench-eta5-W256-S10设置下,TSMixer达到最低平均RMSE(3.4±1.71)和萨克纳分数((2.50±2.99)×10⁴)。消融实验表明,飞行阶段选择、时间降采样方法和RUL上限阈值显著影响结果。该基准提供可复现的子基准,便于控制条件下的模型对比,而CPM-N-CMAPSS则为未来迁移学习与域适应研究提供阶段特异性数据基础。
原文摘要 · Abstract (English)
Remaining useful life (RUL) prediction estimates how long an engine can continue safe operation and is central to maintenance planning. N-CMAPSS extends C-MAPSS by simulating run-to-failure aero-engine trajectories using recorded real-flight profiles and retaining complete within-flight time series rather than cycle-level snapshots. However, this added realism reduces evaluation control because full-flight records increase data volume and entangle degradation cues with operating-regime variation, complicating preprocessing choices and direct comparisons of RUL modeling performance. To mitigate this issue, this paper proposes CruiseBench, a cruise-stage RUL benchmark derived from N-CMAPSS. It introduces CPM-N-CMAPSS (Cruising-Period Mask for N-CMAPSS), a mask artifact that stores cycle-local cruising intervals identified by the common-altitude method for the nine accessible subdatasets. CruiseBench applies a fixed protocol to the masked rows, using scenario descriptors and measured sensors as inputs while excluding virtual sensors, health parameters, and auxiliary metadata from the feature tensor, preserving native-resolution windows, and applying dataset-wise RUL caps. Experiments with LSTM, GRU, TCN, and TSMixer provide baseline results for this setting. Under CruiseBench-eta5-W256-S10, TSMixer obtains the lowest average RMSE, $3.4\pm1.71$, and Saxena score, $(2.50\pm2.99)\times 10^{4}$. Ablation studies show that flight-stage selection, temporal downscaling method, and RUL-cap threshold affect reported results. With its fixed cruise-stage protocol, CruiseBench provides a reproducible sub-benchmark for controlled RUL model comparison and CPM-N-CMAPSS provides a stage-specific data foundation for future transfer-learning and domain-adaptation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。