arXiv:2410.18959cs.LGcs.AI2024-10被引 83

构建多模态时间序列预测基准,验证文本上下文对准确预测的关键作用。

Context is Key: A Benchmark for Forecasting with Essential Textual Information

  • 设计含多种文本上下文的时间序列预测任务,要求模型融合数值与文本信息。
  • 基于大模型的提示方法在基准上表现最优,超越传统统计与基础模型。
  • 揭示大模型在理解关键文本信息上的潜力与局限,适合决策类应用研究者。

预测是众多领域决策中的关键任务。尽管历史数值数据提供起点,却无法传递可靠预测所需的完整上下文。人类预测者常依赖背景知识与约束条件,这些信息可通过自然语言高效传达。然而,尽管大语言模型(LLM)在预测方面取得进展,其有效整合文本信息的能力仍不明确。为此,我们提出「上下文即关键」(Context is Key, CiK)基准,将数值数据与精心设计的多样化文本上下文配对,要求模型融合两种模态;关键在于,每个任务均需理解文本上下文才能成功解决。我们评估了多种方法,包括统计模型、时间序列基础模型及基于LLM的预测器,并提出一种简单但有效的LLM提示方法,在该基准上优于所有其他测试方法。实验凸显了引入上下文信息的重要性,展示了使用基于LLM的预测模型时令人意外的表现,也揭示了其关键缺陷。该基准旨在推动多模态预测发展,促进既准确又面向各类技术背景决策者的模型进步。基准可视化地址:https://servicenow.github.io/context-is-key-forecasting/v0/。

原文摘要 · Abstract (English)

Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge and constraints, which can efficiently be communicated through natural language. However, in spite of recent progress with LLM-based forecasters, their ability to effectively integrate this textual information remains an open question. To address this, we introduce "Context is Key" (CiK), a time-series forecasting benchmark that pairs numerical data with diverse types of carefully crafted textual context, requiring models to integrate both modalities; crucially, every task in CiK requires understanding textual context to be solved successfully. We evaluate a range of approaches, including statistical models, time series foundation models, and LLM-based forecasters, and propose a simple yet effective LLM prompting method that outperforms all other tested methods on our benchmark. Our experiments highlight the importance of incorporating contextual information, demonstrate surprising performance when using LLM-based forecasting models, and also reveal some of their critical shortcomings. This benchmark aims to advance multimodal forecasting by promoting models that are both accurate and accessible to decision-makers with varied technical expertise. The benchmark can be visualized at https://servicenow.github.io/context-is-key-forecasting/v0/.

时间序列多模态大模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。