用对比学习提升Mamba对关键时间点的识别能力
Repetitive Contrastive Learning Enhances Mamba's Selectivity in Time Series Prediction
- 通过重复对比学习强化Mamba对关键时间步的选择性
- 在多个数据集上显著提升预测性能,达新最优
- 适合关注时序建模中注意力机制优化的研究者
长序列预测是时间序列建模的核心挑战。尽管基于Mamba的模型因其序列选择能力表现出色,但仍存在对关键时间步关注不足、噪声抑制不充分的问题,根源在于其选择能力有限。为此,本文提出重复对比学习(RCL),一种面向标记级别的对比预训练框架,旨在增强Mamba的选通能力。RCL对单个Mamba模块进行预训练,提升其选择性,并将参数迁移至多种骨干模型中,从而改善时序预测性能。该方法通过加入高斯噪声的序列增强,实施跨序列与内序列对比学习,使Mamba模块更聚焦于信息丰富的时序点,同时忽略噪声。大量实验表明,RCL持续提升骨干模型性能,超越现有方法,达到当前最优水平。此外,我们提出了两种量化Mamba选择能力的指标,为RCL带来的改进提供了理论、定性和定量支持。
原文摘要 · Abstract (English)
Long sequence prediction is a key challenge in time series forecasting. While Mamba-based models have shown strong performance due to their sequence selection capabilities, they still struggle with insufficient focus on critical time steps and incomplete noise suppression, caused by limited selective abilities. To address this, we introduce Repetitive Contrastive Learning (RCL), a token-level contrastive pretraining framework aimed at enhancing Mamba's selective capabilities. RCL pretrains a single Mamba block to strengthen its selective abilities and then transfers these pretrained parameters to initialize Mamba blocks in various backbone models, improving their temporal prediction performance. RCL uses sequence augmentation with Gaussian noise and applies inter-sequence and intra-sequence contrastive learning to help the Mamba module prioritize information-rich time steps while ignoring noisy ones. Extensive experiments show that RCL consistently boosts the performance of backbone models, surpassing existing methods and achieving state-of-the-art results. Additionally, we propose two metrics to quantify Mamba's selective capabilities, providing theoretical, qualitative, and quantitative evidence for the improvements brought by RCL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。