小数据下用神经架构搜索优化冷却器寿命预测模型
Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction
- 用小数据专用编码器+神经架构搜索,控制模型复杂度
- 在低温设备数据上实现高精度寿命预测,训练成本更低
- 适合工业界数据少、需可解释性的场景
大规模预训练时间序列模型虽性能强,但依赖大量多样数据,而工业与科学领域常缺乏此类数据。为此,我们提出小数据表示模型家族(FSD-RM)范式,作为替代方案。不依赖大规模预训练,而是采用容量可控的表示学习,选用适合小数据场景且可解释性强的编码器(CNN1D、LSTM、GRU、Transformer),在多变量遥测数据上无监督训练,并集成到两阶段下游寿命预测流程中。为系统评估数据受限下的架构权衡,我们采用维度感知神经架构搜索(NAS),联合优化模型容量与输入维度。在冷冻器遥测数据上的实验表明,该方法在保持竞争力预测性能的同时,显著降低训练成本与模型复杂度。贡献在于将成熟表示学习技术整合进由NAS驱动、专为小数据设计的框架中,明确参数设置与设计选择。结果表明,在施加适当归纳偏置与容量控制的前提下,无需大规模预训练也可实现有效表示学习。
原文摘要 · Abstract (English)
Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack. We therefore propose the FSD-RM (Family of Small-Data Representation Models) paradigm as a practical alternative for limited, domain-specific telemetry. Rather than relying on large-scale pretraining, we focus on capacity-controlled representation learning using established encoder architectures (CNN1D, LSTM, GRU, Transformer), selected for their suitability in small-data settings and interpretability. These encoders are trained unsupervised on multivariate telemetry data and integrated into a two-stage pipeline for downstream lifetime prediction. To systematically examine architectural trade-offs under data constraints, we employ \textbf{dimension-aware neural architecture search (NAS)} to jointly optimize model capacity and input dimensionality. Experiments on cryocooler telemetry show that the proposed approach achieves competitive predictive performance while reducing training cost and model complexity. The contribution lies in combining established representation learning techniques within a coherent, NAS-driven framework tailored to small-data regimes, with explicitly defined parameter settings and design choices. The results indicate that effective representation learning can be achieved without large-scale pretraining when appropriate inductive bias and capacity control are applied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。