提出对比准确率评估时间序列预训练数据质量,无需标签即可选优。
Measuring Pre-training Data Quality without Labels for Time Series Foundation Models
- 用对比学习构建表示空间,通过对比准确率衡量数据质量。
- 对比准确率与下游任务平均准确率呈正相关,验证其有效性。
- 适合无标签场景下筛选高质量时间序列预训练数据集。
近年来,时间序列基础模型在跨下游任务泛化方面受到关注。强大基础模型的关键在于多样化的预训练数据集,但时间序列分类的数据收集尤为困难。本文研究了基于对比学习的基础模型性能随预训练数据的变化情况。提出一种新指标——对比准确率,用于评估基础模型所学表示空间的质量。实验表明,该指标与模型在多个下游任务上的准确率呈正相关。这说明对比准确率可作为筛选提升预训练效果、增强模型泛化能力的时间序列数据集的标准。
原文摘要 · Abstract (English)
Recently, there has been a growing interest in time series foundation models that generalize across different downstream tasks. A key to strong foundation models is a diverse pre-training dataset, which is particularly challenging to collect for time series classification. In this work, we explore the performance of a contrastive-learning-based foundation model as a function of the data used for pre-training. We introduce contrastive accuracy, a new measure to evaluate the quality of the representation space learned by the foundation model. Our experiments reveal the positive correlation between the proposed measure and the accuracy of the model on a collection of downstream tasks. This suggests that the contrastive accuracy can serve as a criterion to search for time series datasets that can enhance the pre-training and improve thereby the foundation model's generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。