用沃瑟斯坦距离衡量时间序列数据集相似性,助力模型选型与性能预测。
Measuring Time-Series Dataset Similarity using Wasserstein Distance
- 将时间序列数据集视为多元正态分布,用沃瑟斯坦距离度量其相似性。
- 在分布外和迁移学习中,该方法与推理损失相关性超0.60,效果显著。
- 适合研究时间序列基础模型、模型评估与数据筛选的学者使用。
时间序列基础模型的研究兴起,推动了对时间序列数据集(不)相似性度量的需求。该度量方法可辅助模型选择、微调与可视化等任务。本文提出一种基于分布的方法,利用沃瑟斯坦距离衡量时间序列数据集的相似性。将时间序列数据集视为潜在的多元正态分布(MVN)实例,两个数据集的相似性即为其对应MVN间的沃瑟斯坦距离。大量实验与可视化表明,该方法能有效识别相似数据集,并在分布外与迁移学习评估中辅助估计基础模型的推理性能,所提度量与推理损失的相关性高于0.60。
原文摘要 · Abstract (English)
The emergence of time-series foundation model research elevates the growing need to measure the (dis)similarity of time-series datasets. A time-series dataset similarity measure aids research in multiple ways, including model selection, finetuning, and visualization. In this paper, we propose a distribution-based method to measure time-series dataset similarity by leveraging the Wasserstein distance. We consider a time-series dataset an empirical instantiation of an underlying multivariate normal distribution (MVN). The similarity between two time-series datasets is thus computed as the Wasserstein distance between their corresponding MVNs. Comprehensive experiments and visualization show the effectiveness of our approach. Specifically, we show how the Wasserstein distance helps identify similar time-series datasets and facilitates inference performance estimation of foundation models in both out-of-distribution and transfer learning evaluation, with high correlations between our proposed measure and the inference loss (>0.60).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。