arXiv:2510.20119cs.LG2025-10

时间序列模型需从不变性设计出发,而非盲目堆数据。

There is No "apple" in Timeseries: Rethinking TSFM through the Lens of Invariance

  • 提出基于不变性原理构建时间序列数据集的新思路
  • 指出现有模型性能瓶颈源于缺乏概念性数据覆盖
  • 适合研究时序建模与基础模型的学者参考

时间序列基础模型(TSFMs)数量激增,但轻量级监督基线甚至经典模型常能与之比肩。我们认为这一差距源于将自然语言处理或计算机视觉的流水线照搬至时间序列领域。在语言和视觉中,海量网络语料密集捕捉人类概念(如存在大量苹果图像和文本)。而时间序列数据本应补充图像与文本模态,却不存在包含‘苹果’概念的数据集。因此,‘爬取一切在线数据’的范式在时间序列上失效。我们主张,进步必须从机会主义聚合转向原则性设计:构建系统覆盖保持时序语义的不变性空间的数据集。为此,时间序列不变性本体应基于第一性原理建立。唯有通过不变性覆盖实现表征完备,时间序列基础模型才能具备通用化、推理及真正涌现行为所需的对齐结构。

原文摘要 · Abstract (English)

Timeseries foundation models (TSFMs) have multiplied, yet lightweight supervised baselines and even classical models often match them. We argue this gap stems from the naive importation of NLP or CV pipelines. In language and vision, large web-scale corpora densely capture human concepts i.e. there are countless images and text of apples. In contrast, timeseries data is built to complement the image and text modalities. There are no timeseries dataset that contains the concept apple. As a result, the scrape-everything-online paradigm fails for TS. We posit that progress demands a shift from opportunistic aggregation to principled design: constructing datasets that systematically span the space of invariance that preserve temporal semantics. To this end, we suggest that the ontology of timeseries invariances should be built based on first principles. Only by ensuring representational completeness through invariance coverage can TSFMs achieve the aligned structure necessary for generalisation, reasoning, and truly emergent behaviour.

时间序列基础模型不变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。