重构时间序列推理的模型、任务与评估,才能真正理解时序数据。
Achieving Time Series Reasoning Requires Rethinking Model Design, Tasks Formulation, and Evaluation
- 重新设计模型,融合时序特性而非套用NLP方法
- 突破传统预测分类,面向决策相关任务
- 强调可解释性与真实场景鲁棒性,非仅基准得分
理解时间序列数据是众多现实应用的基础。近期研究尝试使用多模态大语言模型(MLLM)结合上下文信息提升时序理解能力,该领域从2023年的7篇论文激增至2025年的超过580篇,但现有方法在真实场景中表现仍差。我们分析了2025年20篇有影响力的工作,涵盖模型设计、任务定义和评估方式,发现关键缺陷:方法多沿用NLP技术,忽视时序核心特性;任务局限于传统预测与分类;评估偏重基准性能,忽略鲁棒性、可解释性及决策相关性。我们主张需同步重构模型设计、任务设定与评估体系。本文定义了时间序列推理,梳理挑战与未来方向,并呼吁构建统一框架,实现真实应用中的鲁棒、可解释、决策导向的推理能力。相关资源见https://github.com/Eleanorkong/Awesome-Time-Series-Reasoning。
原文摘要 · Abstract (English)
Understanding time series data is fundamental to many real-world applications. Recent work explores multimodal large language models (MLLMs) to enhance time series understanding with contextual information beyond numerical signals. This area has grown from 7 papers in 2023 to over 580 in 2025, yet existing methods struggle in real-world settings. We analyze 20 influential works from 2025 across model design, task formulation, and evaluation, and identify critical gaps: methods adapt NLP techniques with limited attention to core time series properties; tasks remain restricted to traditional prediction and classification; and evaluations emphasize benchmarks over robustness, interpretability, or decision relevance. We argue that achieving time series reasoning requires rethinking model design, task formulation, and evaluation together. We define time series reasoning, outline challenges and future directions, and call on researchers to develop unified frameworks for robust, interpretable, and decision-relevant reasoning in real-world applications. The material is available at https://github.com/Eleanorkong/Awesome-Time-Series-Reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。