arXiv:2605.25943cs.LG2026-05

用三模态融合提升非平稳时间序列预测的形状保真度

STaT: Resolving Shape Distortion in Non-Stationary Time Series via Tri-Modal Synergy

论文配图:STaT: Resolving Shape Distortion in Non-Stationary Time Series via Tri-Modal Synergy
图 1 · 摘自论文原文
  • 符号、时序、文本三模态协同建模,捕捉结构变化点
  • 相比传统方法,形状失真降低8.5%,精度提升最高8.9%
  • 适合需要精准捕捉波动趋势的金融、气象等场景

近期时间序列预测研究常将文本与视觉模态融入数值模型以应对非平稳环境。尽管数值表现良好,现有方法往往陷入困境:过度追求平均误差最小化会导致预测过于平滑,忽略关键波动。为此,我们提出STaT——一种符号-时序-文本对齐的多模态架构,无缝融合三种协同模态。符号模态将连续时间序列转为离散标记,精准识别结构模式与转折点;时序模态提取内在序列依赖;文本模态利用领域语义引导宏观趋势。在八个真实世界基准上的综合评估显示,STaT在提升传统幅度指标最高达8.9%的同时,形状失真减少最多达8.5%。

原文摘要 · Abstract (English)

Recent research in time series forecasting frequently investigates the integration of textual and visual modalities with numerical models to better navigate non-stationary environments. Despite delivering solid numerical results, existing multi-modal approaches usually encounter a dilemma: prioritizing the minimization of average errors can result in excessively smooth forecasts that overlook essential fluctuations. To resolve this limitation, we introduce STaT, an innovative multimodal architecture for Symbolic-Temporal-Textual Alignment, which seamlessly unites three synergistic modalities. Specifically, the symbolic modality converts continuous time series into discrete tokens, facilitating the accurate identification of structural patterns and turning points; the temporal modality extracts inherent sequential dependencies; and the textual modality leverages domain semantics to steer the macroscopic forecasting trends. Comprehensive evaluations on eight real-world benchmarks indicate that STaT delivers exceptional performance, enhancing conventional magnitude indicators by up to 8.9% while simultaneously decreasing shape distortion by up to 8.5%.

时间序列多模态形状保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。