通过同步因果对提升多维时间序列预测的泛化能力
Causal Time-Series Synchronization for Multi-Dimensional Forecasting
- 将多维时序数据拆解为因果变量对,利用滞后依赖关系建模
- 在工业数字孪生场景中,预测精度与跨任务泛化能力显著提升
- 适合需要跨领域、多变量时间序列建模的研究者和工程师
流程工业对数字孪生的高要求催生了需具备跨任务、跨领域泛化能力的建模方法,尤其面对不同数据维度与分布偏移的挑战。尽管自然语言处理和计算机视觉已成功应用自监督预训练,但工业数字孪生中的多维时序数据因存在滞后的因果依赖、复杂因果结构及变量数量不一等问题,使此类方法尚未被充分探索。本文提出一种新型通道依赖式预训练策略,通过识别数据驱动的高滞后因果关系,并将因果对同步以生成用于预训练的样本,从而构建通用模型。实验表明,该方法在通道依赖预测中显著优于传统训练方式,提升了预测准确率与泛化性能。
原文摘要 · Abstract (English)
The process industry's high expectations for Digital Twins require modeling approaches that can generalize across tasks and diverse domains with potentially different data dimensions and distributional shifts i.e., Foundational Models. Despite success in natural language processing and computer vision, transfer learning with (self-) supervised signals for pre-training general-purpose models is largely unexplored in the context of Digital Twins in the process industry due to challenges posed by multi-dimensional time-series data, lagged cause-effect dependencies, complex causal structures, and varying number of (exogenous) variables. We propose a novel channel-dependent pre-training strategy that leverages synchronized cause-effect pairs to overcome these challenges by breaking down the multi-dimensional time-series data into pairs of cause-effect variables. Our approach focuses on: (i) identifying highly lagged causal relationships using data-driven methods, (ii) synchronizing cause-effect pairs to generate training samples for channel-dependent pre-training, and (iii) evaluating the effectiveness of this approach in channel-dependent forecasting. Our experimental results demonstrate significant improvements in forecasting accuracy and generalization capability compared to traditional training methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。