arXiv:2510.01116cs.LG2025-10被引 5

用强化学习让大模型学会时间序列分析的一步步推理。

Eliciting Chain-of-Thought Reasoning for Time Series Analysis using Reinforcement Learning

  • 通过可验证奖励机制训练大模型生成分步推理过程。
  • 在多种时间序列任务上显著提升模型表现,关键指标提升达23%以上。
  • 适合需要复杂时序逻辑推理的研究者与工业应用者。

复杂数值时间序列分析通常需要超越现有模型能力的多步推理。医疗诊断和天气预报等任务要求进行反事实分析、逻辑推导、知识应用及多模态上下文整合等序列推理,而现有时间序列模型无法明确执行。尽管近期研究显示大语言模型(LLMs)可通过强化学习(RL)实现复杂链式思维(CoT)推理,但主要集中在数学与编程领域,对时间序列任务表现仍不佳。我们提出首个基于强化学习的框架COUNTS,使大模型在多样化时间序列任务中实现显式的链式思维推理。该方法采用残差向量量化变分自编码器(Residual Vector-Quantized VAE)生成高保真离散标记,无缝融入预训练大模型词汇表。COUNTS经历两阶段训练:首先在时间序列分析任务上进行监督微调以掌握新表示,随后在可验证问题上使用分组相对策略优化(Group Relative Policy Optimization),结合提示策略引导模型在输出最终答案前生成明确的推理步骤。实验表明,这种基于强化学习的中间链式思维方法显著提升了大模型在各类时间序列分析任务中的性能,为复杂时序数据推理开辟了新路径。

原文摘要 · Abstract (English)

Complex numerical time series analysis often demands multi-step reasoning capabilities beyond current models' reach. Tasks like medical diagnosis and weather forecasting require sequential reasoning processes - including counterfactual analysis, logical deduction, knowledge application, and multi-modal contextual integration - that existing time series models cannot explicitly perform. While recent research has shown large language models (LLMs) can achieve sophisticated Chain-of-Thought (CoT) reasoning through reinforcement learning (RL), these advances have primarily focused on mathematical and coding domains, with LLMs still demonstrating poor performance on time series tasks. We introduce Chain Of thought for Understanding Numerical Time Series (COUNTS), the first framework that trains LLMs to perform CoT reasoning across diverse time series tasks using RL with verifiable rewards. Our approach employs a Residual Vector-Quantized VAE to create high-fidelity discrete tokens that seamlessly integrate into a pre-trained LLM's vocabulary. COUNTS undergoes a two-stage training process: first, supervised fine-tuning on time series analysis tasks to master our novel representations, followed by Group Relative Policy Optimization training on verifiable problems using prompting strategies that encourage explicit reasoning steps before producing final answers. Our experiments demonstrate that this RL-driven approach with intermediate CoT reasoning significantly enhances LLM performance across various time series analysis tasks, opening new possibilities for complex temporal data reasoning.

链式思维时间序列强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。