通过强制神经坍缩,提升时间序列少样本预训练效果。
rETF-semiSL: Semi-Supervised Learning for Neural Collapse in Temporal Data
- 用旋转等角紧框架和伪标签实现少样本半监督预训练。
- 在三个数据集上显著优于传统预训练方法,尤其适配LSTM/Transformer。
- 结合生成任务与新序列增强策略,增强时序特征分离性。
时间序列深度神经网络需捕捉复杂动态模式以有效表示动态数据。自监督与半监督学习在大规模模型预训练中表现优异,微调后分类性能常优于从零开始训练。然而,预训练任务选择常依赖经验,其向下游分类的迁移能力不可保证。为此,本文提出一种新型半监督预训练策略,强制潜在表示满足最优训练神经分类器中的神经坍缩现象。采用旋转等角紧框架分类器与伪标签,用少量标注样本预训练深度编码器。为进一步捕捉时序动态并强化嵌入可分性,融合生成式预训练任务,并设计新颖的序列增强策略。在三个多变量时间序列分类数据集上,该方法在LSTM、Transformer及状态空间模型上均显著优于先前预训练方法。结果表明,将预训练目标与理论驱动的嵌入几何对齐具有显著优势。
原文摘要 · Abstract (English)
Deep neural networks for time series must capture complex temporal patterns, to effectively represent dynamic data. Self- and semi-supervised learning methods show promising results in pre-training large models, which -- when finetuned for classification -- often outperform their counterparts trained from scratch. Still, the choice of pretext training tasks is often heuristic and their transferability to downstream classification is not granted, thus we propose a novel semi-supervised pre-training strategy to enforce latent representations that satisfy the Neural Collapse phenomenon observed in optimally trained neural classifiers. We use a rotational equiangular tight frame-classifier and pseudo-labeling to pre-train deep encoders with few labeled samples. Furthermore, to effectively capture temporal dynamics while enforcing embedding separability, we integrate generative pretext tasks with our method, and we define a novel sequential augmentation strategy. We show that our method significantly outperforms previous pretext tasks when applied to LSTMs, transformers, and state-space models on three multivariate time series classification datasets. These results highlight the benefit of aligning pre-training objectives with theoretically grounded embedding geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。