梳理大模型在时间序列推理中的三种思维结构及应用前景。
A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
- 按推理路径分三类:一步直推、线性链式、分支结构。
- 揭示不同结构在可信度与鲁棒性上的优劣表现。
- 适合关注可解释性、决策支持与长期动态系统的研究者。
时间序列推理将时间视为首要维度,直接将中间证据融入答案。本综述定义问题并按推理拓扑将文献分为三类:单步直接推理、带显式中间步骤的线性链推理、以及探索-修正-聚合的分支结构推理。该拓扑与四大核心目标交叉:传统时间序列分析、解释与理解、因果推断与决策、时间序列生成。通过紧凑标签集涵盖分解与验证、集成、工具使用、知识访问、多模态、智能体循环、大模型对齐等机制。跨领域回顾方法与系统,揭示各类拓扑的能力边界与失效场景,并整理了支持研究与部署的精选数据集、基准测试与资源(https://github.com/blacksnail789521/Time-Series-Reasoning-Survey)。强调评估需保持证据可见且时序对齐,提出匹配拓扑与不确定性、以可观测物为基、应对流式与漂移的规划策略,以及将成本与延迟视为设计预算。强调推理结构需在可接地性与自修正能力、计算成本与可复现性间取得平衡;未来进展依赖于将推理质量与实际效用挂钩的基准,以及在漂移感知、流式、长周期设置下权衡成本与风险的闭环测试平台。整体方向从狭义准确转向规模化可靠性,推动系统不仅能分析,还能理解、解释并基于可追溯证据可信地行动于动态世界。
原文摘要 · Abstract (English)
Time series reasoning treats time as a first-class axis and incorporates intermediate evidence directly into the answer. This survey defines the problem and organizes the literature by reasoning topology with three families: direct reasoning in one step, linear chain reasoning with explicit intermediates, and branch-structured reasoning that explores, revises, and aggregates. The topology is crossed with the main objectives of the field, including traditional time series analysis, explanation and understanding, causal inference and decision making, and time series generation, while a compact tag set spans these axes and captures decomposition and verification, ensembling, tool use, knowledge access, multimodality, agent loops, and LLM alignment regimes. Methods and systems are reviewed across domains, showing what each topology enables and where it breaks down in faithfulness or robustness, along with curated datasets, benchmarks, and resources that support study and deployment (https://github.com/blacksnail789521/Time-Series-Reasoning-Survey). Evaluation practices that keep evidence visible and temporally aligned are highlighted, and guidance is distilled on matching topology to uncertainty, grounding with observable artifacts, planning for shift and streaming, and treating cost and latency as design budgets. We emphasize that reasoning structures must balance capacity for grounding and self-correction against computational cost and reproducibility, while future progress will likely depend on benchmarks that tie reasoning quality to utility and on closed-loop testbeds that trade off cost and risk under shift-aware, streaming, and long-horizon settings. Taken together, these directions mark a shift from narrow accuracy toward reliability at scale, enabling systems that not only analyze but also understand, explain, and act on dynamic worlds with traceable evidence and credible outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。