arXiv:2601.03248cs.CL2026-01ACL被引 11

让大模型学会结合时空数据与文本进行推理,提升关键系统决策能力。

STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning

  • 用空间感知强化学习训练大模型,显式融合时间序列、图结构和文本。
  • 在四个任务上平均提升17%~135%准确率,仅需0.004倍专有模型成本。
  • 适合交通、电网、疫情等需要时空推理的高风险决策场景。

时间序列中的时空推理需要显式整合时间动态、空间依赖和文本上下文,对交通网络、电力系统、疾病传播等高风险决策系统至关重要。然而当前研究多侧重预测精度,忽视推理能力。为此,我们构建了包含病因推理、实体识别、相关性推理和上下文预测四类任务的ST-Bench基准,基于网络随机微分方程(SDE)的多智能体数据生成流程构建。进一步提出STReasoner,通过空间感知强化学习(S-GRPO)算法,使大模型能有效融合时序数据、图结构与文本信息实现可解释推理。实验表明,STReasoner在各项任务中平均准确率提升17%至135%,计算成本仅为专有模型的0.004倍,并在真实世界数据上表现稳健。

原文摘要 · Abstract (English)

Spatio-temporal reasoning in time series involves the explicit synthesis of temporal dynamics, spatial dependencies, and textual context. This capability is vital for high-stakes decision-making in systems such as traffic networks, power grids, and disease propagation. However, the field remains underdeveloped because most existing works prioritize predictive accuracy over reasoning. To address the gap, we introduce ST-Bench, a benchmark consisting of four core tasks, including etiological reasoning, entity identification, correlation reasoning, and in-context forecasting, developed via a network SDE-based multi-agent data synthesis pipeline. We then propose STReasoner, which empowers LLM to integrate time series, graph structure, and text for explicit reasoning. To promote spatially grounded logic, we introduce S-GRPO, a reinforcement learning algorithm that rewards performance gains specifically attributable to spatial information. Experiments show that STReasoner achieves average accuracy gains between 17% and 135% at only 0.004X the cost of proprietary models and generalizes robustly to real-world data.

时空推理大模型强化学习时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。