arXiv:2410.10687cs.CLcs.AI2024-10

借鉴NLP经验构建多变量时间序列基准数据集

Building a Multivariate Time Series Benchmarking Datasets Inspired by Natural Language Processing (NLP)

  • 仿照NLP基准数据集构建流程,整合多样且具挑战性的时序数据
  • 强调领域相关性与数据复杂度,提升模型训练代表性
  • 支持多任务学习,助力时序模型性能提升

时间序列分析在多个领域日益重要,高效模型的开发依赖高质量基准数据集。受自然语言处理(NLP)基准数据集推动预训练模型发展的启发,本文提出一种面向时间序列分析的综合性基准数据集构建方法。通过研究NLP中基准数据集的构建策略,并针对时间序列数据的独特挑战进行适配,系统探讨了如何筛选具有多样性、代表性及挑战性的数据集,强调领域相关性和数据复杂度的重要性。同时,研究基于该基准数据集的多任务学习策略,以提升时间序列模型性能。本工作旨在借鉴NLP成功经验,推动时间序列建模技术的发展。

原文摘要 · Abstract (English)

Time series analysis has become increasingly important in various domains, and developing effective models relies heavily on high-quality benchmark datasets. Inspired by the success of Natural Language Processing (NLP) benchmark datasets in advancing pre-trained models, we propose a new approach to create a comprehensive benchmark dataset for time series analysis. This paper explores the methodologies used in NLP benchmark dataset creation and adapts them to the unique challenges of time series data. We discuss the process of curating diverse, representative, and challenging time series datasets, highlighting the importance of domain relevance and data complexity. Additionally, we investigate multi-task learning strategies that leverage the benchmark dataset to enhance the performance of time series models. This research contributes to the broader goal of advancing the state-of-the-art in time series modeling by adopting successful strategies from the NLP domain.

时间序列基准数据集多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。