arXiv:2602.07830cs.AI2026-02被引 1

用可验证的思维数据训练小模型,让其在时间序列推理上媲美大模型。

Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning

  • 构建多模态时间序列文本数据集,支持过程可验证的思维链标注。
  • 按难度与任务类型分层调度数据,提升训练效率。
  • 设计细粒度多目标奖励机制,让3B-4B小模型超越大型闭源模型。

时间序列是众多应用领域的常见数据形式,实现对各类时间序列任务的合理求解一直是长期目标。近期大语言模型(LLMs)通过强化学习(RL)解锁的推理能力,为长链式思维(CoT)任务提供了新机遇。然而,将LLM推理应用于时间序列仍处于初期阶段,受限于缺乏精心构造的时间序列CoT训练数据、数据效率不足以及未针对此类数据定制的强化学习算法。本文提出VeriTime框架,通过数据合成、数据调度和强化学习训练,定制化提升LLM在时间序列推理中的表现。首先,设计了一种数据合成流程,构建具有过程可验证标注的时序-文本多模态数据集;其次,提出一种基于难度层级与任务分类体系的数据调度机制;第三,开发两阶段强化微调方法,采用细粒度多目标奖励,利用可验证的过程级CoT数据。大量实验表明,VeriTime显著提升了不同时间序列推理任务中LLM的性能。特别地,使小型3B、4B模型在推理能力上达到甚至超过更大规模的专有大模型水平。

原文摘要 · Abstract (English)

Time series is a pervasive data type across various application domains, rendering the reasonable solving of diverse time series tasks a long-standing goal. Recent advances in large language models (LLMs), especially their reasoning abilities unlocked through reinforcement learning (RL), have opened new opportunities for tackling tasks with long Chain-of-Thought (CoT) reasoning. However, leveraging LLM reasoning for time series remains in its infancy, hindered by the absence of carefully curated time series CoT data for training, limited data efficiency caused by underexplored data scheduling, and the lack of RL algorithms tailored for exploiting such time series CoT data. In this paper, we introduce VeriTime, a framework that tailors LLMs for time series reasoning through data synthesis, data scheduling, and RL training. First, we propose a data synthesis pipeline that constructs a TS-text multimodal dataset with process-verifiable annotations. Second, we design a data scheduling mechanism that arranges training samples according to a principled hierarchy of difficulty and task taxonomy. Third, we develop a two-stage reinforcement finetuning featuring fine-grained, multi-objective rewards that leverage verifiable process-level CoT data. Extensive experiments show that VeriTime substantially boosts LLM performance across diverse time series reasoning tasks. Notably, it enables compact 3B, 4B models to achieve reasoning capabilities on par with or exceeding those of larger proprietary LLMs.

时间序列推理生成强化学习小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。