构建多模态时间序列基准,评估模型理解新闻与数据趋势关系的能力
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
- 设计金融与天气领域配对的文本与时间序列数据集
- 发现主流大模型在长期依赖和因果推理上表现不足
- 适合研究跨模态推理、时间序列分析的学者使用
理解文本新闻与时间序列演变之间的关系是应用数据科学中的关键但研究不足的挑战。尽管多模态学习日益流行,现有数据集在评估跨模态推理和复杂问答任务方面仍显不足,而这正是捕捉叙事信息与时间模式复杂交互的关键。为此,我们提出多模态时间序列基准(MTBench),一个大规模基准,用于评估大语言模型(LLMs)在金融与天气领域的时序与文本理解能力。MTBench包含配对的时间序列与文本数据,如新闻与对应股价变化、天气报告与历史温度记录。不同于仅关注单一模态的现有基准,MTBench提供了一个全面测试平台,使模型能够联合推理结构化数值趋势与非结构化文本叙述。其丰富性支持多种任务,包括时间序列预测、语义与技术趋势分析、新闻驱动的问答(QA),旨在考察模型捕捉时间依赖、从文本中提取关键洞察及融合跨模态信息的能力。我们在MTBench上评估了先进大模型,发现当前模型在捕捉长期依赖、解释金融与天气趋势中的因果关系、有效融合多模态信息方面仍存在显著挑战。
原文摘要 · Abstract (English)
Understanding the relationship between textual news and time-series evolution is a critical yet under-explored challenge in applied data science. While multimodal learning has gained traction, existing multimodal time-series datasets fall short in evaluating cross-modal reasoning and complex question answering, which are essential for capturing complex interactions between narrative information and temporal patterns. To bridge this gap, we introduce Multimodal Time Series Benchmark (MTBench), a large-scale benchmark designed to evaluate large language models (LLMs) on time series and text understanding across financial and weather domains. MTbench comprises paired time series and textual data, including financial news with corresponding stock price movements and weather reports aligned with historical temperature records. Unlike existing benchmarks that focus on isolated modalities, MTbench provides a comprehensive testbed for models to jointly reason over structured numerical trends and unstructured textual narratives. The richness of MTbench enables formulation of diverse tasks that require a deep understanding of both text and time-series data, including time-series forecasting, semantic and technical trend analysis, and news-driven question answering (QA). These tasks target the model's ability to capture temporal dependencies, extract key insights from textual context, and integrate cross-modal information. We evaluate state-of-the-art LLMs on MTbench, analyzing their effectiveness in modeling the complex relationships between news narratives and temporal patterns. Our findings reveal significant challenges in current models, including difficulties in capturing long-term dependencies, interpreting causality in financial and weather trends, and effectively fusing multimodal information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。