arXiv:2509.24803cs.LGcs.AI2025-09中稿 · the 14th Internati…被引 25

首个支持复杂时间序列推理的统一模型,提升因果发现与预测准确率。

TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models

  • 构建四类原子任务框架,涵盖感知、外推与决策三重推理能力。
  • 在2.3万条高质量数据上训练,因果发现准确率达64.0%(超GPT-4.1)。
  • 适合需要时间序列深度理解的金融、医疗等场景研究者使用。

近期多模态时间序列学习推动了从基础模式分析向高级理解与推理的范式转变。然而,现有数据集仍停留在表面对齐与问答层面,缺乏真正需要时间序列推理的任务设计与高质量数据支持,限制了实用时间序列推理模型(TSRM)的发展。为此,我们提出时间序列推理套件(TSR-Suite),首次系统化定义了四项原子任务,覆盖三大核心推理能力:(1)感知——通过情景理解与因果发现;(2)外推——基于事件感知的预测;(3)决策——结合感知与外推的权衡判断。该套件包含超过2.3万样本,其中2.3千条经人工引导的分层标注流程精心构建。在此基础上,我们推出首个统一推理模型TimeOmni-1,采用多阶段训练,融合多样化任务场景、新型奖励函数与定制优化策略。实验表明,TimeOmni-1在所有任务中均展现强大分布外泛化能力,有效响应率达高。其因果发现准确率提升至64.0%(相较GPT-4.1的35.9%),事件感知预测任务的有效响应率较GPT-4.1提高超6%。

原文摘要 · Abstract (English)

Recent advances in multimodal time series learning underscore a paradigm shift from analytics centered on basic patterns toward advanced time series understanding and reasoning. However, existing multimodal time series datasets mostly remain at the level of surface alignment and question answering, without reaching the depth of genuine reasoning. The absence of well-defined tasks that genuinely require time series reasoning, along with the scarcity of high-quality data, has limited progress in building practical time series reasoning models (TSRMs). To this end, we introduce Time Series Reasoning Suite (TSR-Suite), which formalizes four atomic tasks that span three fundamental capabilities for reasoning with time series: (1) perception, acquired through scenario understanding and causality discovery; (2) extrapolation, realized via event-aware forecasting; and (3) decision-making, developed through deliberation over perception and extrapolation. TSR-Suite is the first comprehensive time series reasoning suite that supports not only thorough evaluation but also the data pipeline and training of TSRMs. It contains more than 23K samples, of which 2.3K are carefully curated through a human-guided hierarchical annotation process. Building on this foundation, we introduce TimeOmni-1, the first unified reasoning model designed to address diverse real-world problems demanding time series reasoning. The model is trained in multiple stages, integrating a mixture of task scenarios, novel reward functions, and tailored optimizations. Experiments show that TimeOmni-1 delivers strong out-of-distribution generalization across all tasks and achieves a high rate of valid responses. It significantly improves causality discovery accuracy (64.0% vs. 35.9% with GPT-4.1) and raises the valid response rate by over 6% compared to GPT-4.1 on the event-aware forecasting task.

时间序列推理模型因果发现多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。