新基准TimesX提升多模态时序预测评估可靠性
Rethinking Multimodal Time-Series Forecasting Evaluation

- 构建包含真实世界数据与丰富文本上下文的多模态时序基准
- 零样本方法在旧基准表现好,但在新基准上普遍失效
- 结合文本信息的简单集成方法反而优于复杂模型
我们提出一个新的上下文丰富、多模态时间序列预测基准TimesX。该基准通过自动化数据生成流程,包含多个领域的真实世界时间序列及多样化的文本上下文,解决了现有基准的三大问题:(1)数据规模小且为合成数据导致泛化能力差;(2)文本上下文类型有限;(3)评估中难以避免数据泄露。我们在TimesX上对零样本多模态预测方法进行了全面实证研究。结果表明,许多在旧基准表现优异的方法在TimesX上表现不佳。相反,利用丰富文本上下文的简单集成方法在TimesX上超越了强基线模型。
原文摘要 · Abstract (English)
We introduce a new context-enriched, multimodal time series forecasting benchmark, TimesX. TimesX contains a wide selection of high-quality real-world time series with diverse domains and textual contexts obtained from an automated data generation pipeline, which helps address three main issues of existing multimodal forecasting benchmarks: (1) poor generalization due to the small scale and synthetic nature of benchmark data, (2) very limited types of textual contexts in the benchmarks, and (3) an inability to mitigate data leakage in evaluation. We conduct a thorough empirical study of zero-shot multimodal forecasting approaches on TimesX. Our results suggest that many approaches that perform well on existing benchmarks may fail on TimesX. In contrast, simple ensemble methods that leverage rich textual context accompanying time-series can outperform strong baselines on TimesX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。