统一工具箱,让时间序列数据相似性对比更高效可复现。
TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity

- 构建统一框架,支持多种相似性方法对比
- 可灵活扩展新数据集与下游任务
- 适配数据集级与序列级相似性评估
人工智能的快速发展推动了时间序列分析的研究,尤其在预测、分类和生成任务中。近年来,基础模型受益于时间序列数据集相似性,在微调时选择源数据集方面具有关键作用。然而,现有相似性评测工具实现分散,难以扩展。为此,我们提出统一框架——时间序列数据集相似性工具箱(TSDS-Toolbox)。该工具箱支持:(1)系统化、可复现的时间序列数据集相似性方法对比;(2)用户灵活添加自定义数据集、相似性方法及下游时间序列任务;(3)通过集成的时间序列数据集压缩器,一致评估数据集级与序列级相似性方法。在多种实验设置下进行了全面验证,证明其有效性。该工具箱已公开可用。
原文摘要 · Abstract (English)
The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis, particularly in forecasting, classification, and generation tasks. Recent models, especially foundation models, benefit from time-series dataset similarity due to its significant role in source dataset selection for fine-tuning. However, many existing implementations for benchmarking time-series dataset similarity methods are fragmented and difficult to extend. To address this, we present a unified framework, the Time-Series Dataset Similarity Toolbox (TSDS-Toolbox). Our work enables (1) systematic and reproducible comparisons of time-series dataset similarity methods; (2) flexible extensibility for users to add customized datasets, similarity methods, and downstream time-series tasks; and (3) consistent evaluation of both dataset-level and series-level similarity methods through integrated time-series dataset reducers. The effectiveness of TSDS-Toolbox is validated through comprehensive experiments under diverse experimental settings. Our toolbox is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。