arXiv:2506.21611cs.CLcs.AI2025-06被引 7

多模态提升时间序列预测,但效果取决于模型与数据条件。

When Does Multimodality Lead to Better Time Series Forecasting?

  • 对比对齐与提示两种多模态方法,分析其适用场景。
  • 在文本提供互补信息且数据充足时,多模态显著提升预测效果。
  • 高容量文本模型+弱时序模型+合适对齐策略最受益。

近年来,将文本信息融入基础模型以进行时间序列预测日益受到关注。然而,这种多模态融合是否始终带来提升,以及在何种条件下有效仍不明确。本文在涵盖健康、环境、经济等7个领域的16个预测任务上系统研究了这一问题。评估了两种主流多模态范式:基于对齐的方法(对齐时序与文本表征)和基于提示的方法(直接用大语言模型生成预测)。结果表明,多模态的收益高度依赖具体条件。尽管某些场景中报告了性能提升,但并非在所有数据集或模型上均成立。通过解耦模型架构属性与数据特征的影响,我们获得跨领域通用的定性洞见:模型层面,当文本模型容量高、时序模型较弱且对齐策略合适时,文本信息最有帮助;数据层面,当训练数据充足且文本提供时序数据未覆盖的互补预测信号时,性能提升更可能出现。本研究为理解多模态何时能助力预测提供了严谨、量化的依据,揭示其优势并非普遍,也常与直觉不符。

原文摘要 · Abstract (English)

Recently, there has been growing interest in incorporating textual information into foundation models for time series forecasting. However, it remains unclear whether and under what conditions such multimodal integration consistently yields gains. We systematically investigate these questions across a diverse benchmark of 16 forecasting tasks spanning 7 domains, including health, environment, and economics. We evaluate two popular multimodal forecasting paradigms: aligning-based methods, which align time series and text representations; and prompting-based methods, which directly prompt large language models for forecasting. Our findings reveal that the benefits of multimodality are highly condition-dependent. While we confirm reported gains in some settings, these improvements are not universal across datasets or models. To move beyond empirical observations, we disentangle the effects of model architectural properties and data characteristics, drawing data-agnostic insights that generalize across domains. Our findings highlight that on the modeling side, incorporating text information is most helpful given (1) high-capacity text models, (2) comparatively weaker time series models, and (3) appropriate aligning strategies. On the data side, performance gains are more likely when (4) sufficient training data is available and (5) the text offers complementary predictive signal beyond what is already captured from the time series alone. Our study offers a rigorous, quantitative foundation for understanding when multimodality can be expected to aid forecasting tasks, and reveals that its benefits are neither universal nor always aligned with intuition.

多模态时间序列预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。