用情景引导的多模态预测新基准,让模型学会根据假设场景调整预测。
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
- 引入专家设计的合理或反事实情景作为上下文,测试模型能否据此调整预测。
- 相比传统单模态方法,该基准能更真实反映人类在决策中如何结合假设与历史数据。
- 适合研究多模态预测、具身推理和决策支持系统的学者使用。
时间序列预测对实际决策至关重要,但现有方法多为单模态,依赖历史模式外推。尽管大语言模型(LLMs)展现出多模态预测潜力,但现有基准大多提供回溯性或不匹配的原始上下文,难以判断模型是否真正利用了文本输入。实践中,人类专家常结合假设性‘如果’情景与历史证据进行预测,同一观测在不同情景下可能产生不同结果。受此启发,我们提出 What If TSF(WIT),一个面向情景引导的多模态预测的基准,用于评估模型能否基于上下文文本(尤其是未来情景)进行条件化预测。通过提供专家构建的合理或反事实情景,WIT 构建了一个严谨的测试环境,推动多模态预测的发展。该基准已开源:https://github.com/jinkwan1115/WhatIfTSF。
原文摘要 · Abstract (English)
Time series forecasting is critical to real-world decision making, yet most existing approaches remain unimodal and rely on extrapolating historical patterns. While recent progress in large language models (LLMs) highlights the potential for multimodal forecasting, existing benchmarks largely provide retrospective or misaligned raw context, making it unclear whether such models meaningfully leverage textual inputs. In practice, human experts incorporate what-if scenarios with historical evidence, often producing distinct forecasts from the same observations under different scenarios. Inspired by this, we introduce What If TSF (WIT), a multimodal forecasting benchmark designed to evaluate whether models can condition their forecasts on contextual text, especially future scenarios. By providing expert-crafted plausible or counterfactual scenarios, WIT offers a rigorous testbed for scenario-guided multimodal forecasting. The benchmark is available at https://github.com/jinkwan1115/WhatIfTSF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。