arXiv:2503.01895cs.LGcs.AI2025-03被引 15

首次系统评估推理策略在零样本时间序列预测中的效果

Evaluating System 1 vs. 2 Reasoning Approaches for Zero-Shot Time Series Forecasting: A Benchmark and Insights

  • 构建ReC4TS基准,测试多种推理方法在零样本预测中的表现
  • 自洽性推理是效果最好的测试时策略,多模态任务受益更明显
  • 提供带推理轨迹的TimeThinking数据集和可扩展的推理框架

推理能力对解决复杂任务至关重要。随着大语言模型(LLM)的发展,出现了多种推理策略,包括测试时增强(如思维链)和训练后优化(如DeepSeek-R1)。尽管这些策略在语言或视觉任务中表现优异,其在时间序列预测(TSF),尤其是零样本TSF中的适用性和影响仍不明确。本文提出首个系统评估主流推理策略在零样本TSF中有效性的基准ReC4TS。该基准覆盖八个领域、单模态与多模态、短时与长时预测任务。关键发现:(1) 自洽性为最有效的测试时推理策略;(2) 组相对策略优化更适合训练后激励推理能力;(3) 多模态TSF比单模态更受益于推理策略。此外,ReC4TS构建了两个开创性基础:(1) 新数据集TimeThinking,包含多个先进LLM生成的推理轨迹;(2) 基于自洽性推理的简单测试时缩放律,在基础TSF模型上验证。所有数据与代码公开于https://github.com/AdityaLab/OpenTimeR

原文摘要 · Abstract (English)

Reasoning ability is crucial for solving challenging tasks. With the advancement of foundation models, such as the emergence of large language models (LLMs), a wide range of reasoning strategies has been proposed, including test-time enhancements, such as Chain-ofThought, and post-training optimizations, as used in DeepSeek-R1. While these reasoning strategies have demonstrated effectiveness across various challenging language or vision tasks, their applicability and impact on time-series forecasting (TSF), particularly the challenging zero-shot TSF, remain largely unexplored. In particular, it is unclear whether zero-shot TSF benefits from reasoning and, if so, what types of reasoning strategies are most effective. To bridge this gap, we propose ReC4TS, the first benchmark that systematically evaluates the effectiveness of popular reasoning strategies when applied to zero-shot TSF tasks. ReC4TS conducts comprehensive evaluations across datasets spanning eight domains, covering both unimodal and multimodal with short-term and longterm forecasting tasks. More importantly, ReC4TS provides key insights: (1) Self-consistency emerges as the most effective test-time reasoning strategy; (2) Group-relative policy optimization emerges as a more suitable approach for incentivizing reasoning ability during post-training; (3) Multimodal TSF benefits more from reasoning strategies compared to unimodal TSF. Beyond these insights, ReC4TS establishes two pioneering starting blocks to support future zero-shot TSF reasoning research: (1) A novel dataset, TimeThinking, containing forecasting samples annotated with reasoning trajectories from multiple advanced LLMs, and (2) A new and simple test-time scaling-law validated on foundational TSF models enabled by self-consistency reasoning strategy. All data and code are publicly accessible at: https://github.com/AdityaLab/OpenTimeR

时间序列预测推理能力零样本学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。