arXiv:2503.01875cs.CLcs.AI2025-03ACL被引 92

让大模型用自然语言处理多任务时间序列数据,支持推理与问答。

Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement

  • 构建统一框架,用自然语言提问实现时间序列多任务分析。
  • 推出含20万问答对的TSQA数据集,覆盖环境、交通等多元场景。
  • 持续预训练主流大模型,显著提升其时间序列推理能力。

时间序列数据在金融、医疗、能源等领域至关重要。然而现有方法与数据集大多局限于预测或异常检测等单一任务。为此,我们提出时间序列多任务问答框架Time-MQA,支持自然语言查询下的数值分析与开放性推理问答。核心是TSQA数据集,包含约20万条来自环境、交通等多样时间序列的问答对,覆盖不同长度,促进模型鲁棒性发展。我们进一步通过在TSQA上持续预训练Mistral 7B、Llama-3 8B和Qwen-2.5 7B等大模型,显著增强其时间序列理解与推理能力,实现更高级、直观的时序数据交互。完整数据集、模型、评估问卷等均已开源。

原文摘要 · Abstract (English)

Time series data are foundational in finance, healthcare, and energy domains. However, most existing methods and datasets remain focused on a narrow spectrum of tasks, such as forecasting or anomaly detection. To bridge this gap, we introduce Time Series Multi-Task Question Answering (Time-MQA), a unified framework that enables natural language queries across multiple time series tasks - numerical analytical tasks and open-ended question answering with reasoning. Central to Time-MQA is the TSQA dataset, a large-scale dataset containing $\sim$200k question-answer pairs derived from diverse time series spanning environment, traffic, etc. This comprehensive resource covers various time series lengths and promotes robust model development. We further demonstrate how continually pre-training large language models (Mistral 7B, Llama-3 8B, and Qwen-2.5 7B) on the TSQA dataset enhanced time series reasoning capabilities, moving beyond mere numeric tasks and enabling more advanced and intuitive interactions with temporal data. The complete TSQA dataset, models, user study questionnaires for evaluation, and other related materials have been open-sourced.

时间序列大模型问答系统多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。