arXiv:2510.21795cs.CVcs.AI2025-10被引 4

提出HIBA架构,让时间序列模型在零样本下也能高效捕捉多尺度依赖。

Xihe: Scalable Zero-Shot Time Series Learner Via Hierarchical Interleaved Block Attention

  • 用分层块内块间稀疏注意力,同时处理局部和全局时序模式。
  • 9.5M小模型超越多数现有模型,15亿大模型创零样本新纪录。
  • 适合需要高泛化能力的跨数据集时序分析任务。

时间序列基础模型(TSFMs)虽借助语言模型架构取得进展,但直接迁移导致难以有效捕捉时间序列固有的多尺度时序依赖,尤其在数据分布差异大、采样策略不同时表现受限。为此,本文提出分层交错块注意力(HIBA),通过块内稀疏注意力实现局部信息交互,块间注意力捕捉全局时序模式动态演化。基于HIBA架构,我们构建了可扩展的Xihe系列模型,参数量从9.5M到1.5B不等。在GIFT-Eval基准测试中,最紧凑的Xihe-tiny(9.5M)优于多数现有TSFMs,而最大规模的Xihe-max(1.5B)在零样本任务上达到新纪录,显著超越此前最优结果。整个参数范围内持续优异的表现,充分证明HIBA架构在泛化能力和设计上的优越性。

原文摘要 · Abstract (English)

The rapid advancement of time series foundation models (TSFMs) has been propelled by migrating architectures from language models. While existing TSFMs demonstrate impressive performance, their direct adoption of cross-domain architectures constrains effective capture of multiscale temporal dependencies inherent to time series data. This limitation becomes particularly pronounced during zero-shot transfer across datasets with divergent underlying patterns and sampling strategies. To address these challenges, we propose Hierarchical Interleaved Block Attention (HIBA) which employs hierarchical inter- and intra-block sparse attention to effectively capture multi-scale dependencies. Intra-block attention facilitates local information exchange, and inter-block attention operates across blocks to capture global temporal pattern interaction and dynamic evolution. Leveraging the HIBA architecture, we introduce Xihe, a scalable TSFM family spanning from an ultra-efficient 9.5M parameter configuration to high-capacity 1.5B variant. Evaluated on the comprehensive GIFT-Eval benchmark, our most compact Xihe-tiny model (9.5M) surpasses the majority of contemporary TSFMs, demonstrating remarkable parameter efficiency. More impressively, Xihe-max (1.5B) establishes new state-of-the-art zero-shot performance, surpassing previous best results by a substantial margin. This consistent performance excellence across the entire parameter spectrum provides compelling evidence for the exceptional generalization capabilities and architectural superiority of HIBA.

时间序列零样本学习注意力机制模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。