从众多变量中自动选最相关者,提升时空预测精度与效率。
Spatio-temporal Multivariate Time Series Forecast with Chosen Variables
- 联合优化变量选择与模型参数,动态筛选关键变量输入。
- 在5个真实数据集上,预测误差显著低于现有方法。
- 适合传感器数量受限的交通、空气质量等预测场景。
时空多变量时间序列预测(STMF)利用空间分布的n个变量在近期历史时段的时间序列,预测其在近未来时段的值,在交通流量、空气污染等传感预测中有重要应用。现有研究面对传感器数量m远小于监测点数n的预算限制时,通常假设输入变量预先确定,但如何最优选择m个输入变量以提升预测精度却未被研究。本文首次提出STMF中变量选择的新问题:从n个变量中选出m个最优输入以最大化预测准确率。为此提出统一框架,联合进行变量选择与模型优化,包含三项创新技术:(1) 基于分位数掩码的变量-参数剪枝,逐步剔除低信息量变量与注意力参数;(2) 优先变量-参数重播机制,通过低损失历史样本保留知识以增强模型稳定性;(3) 动态外推机制,利用可学习的空间嵌入和邻接信息,将输入变量信息传播至所有其他变量。在五个真实数据集上的实验表明,该方法在预测精度与计算效率上均显著优于现有基线。
原文摘要 · Abstract (English)
Spatio-Temporal Multivariate time series Forecast (STMF) uses the time series of $n$ spatially distributed variables in a period of recent past to forecast their values in a period of near future. It has important applications in spatio-temporal sensing forecast such as road traffic prediction and air pollution prediction. Recent papers have addressed a practical problem of missing variables in the model input, which arises in the sensing applications where the number $m$ of sensors is far less than the number $n$ of locations to be monitored, due to budget constraints. We observe that the state of the art assumes that the $m$ variables (i.e., locations with sensors) in the model input are pre-determined and the important problem of how to choose the $m$ variables in the input has never been studied. This paper fills the gap by studying a new problem of STMF with chosen variables, which optimally selects $m$-out-of-$n$ variables for the model input in order to maximize the forecast accuracy. We propose a unified framework that jointly performs variable selection and model optimization for both forecast accuracy and model efficiency. It consists of three novel technical components: (1) masked variable-parameter pruning, which progressively prunes less informative variables and attention parameters through quantile-based masking; (2) prioritized variable-parameter replay, which replays low-loss past samples to preserve learned knowledge for model stability; (3) dynamic extrapolation mechanism, which propagates information from variables selected for the input to all other variables via learnable spatial embeddings and adjacency information. Experiments on five real-world datasets show that our work significantly outperforms the state-of-the-art baselines in both accuracy and efficiency, demonstrating the effectiveness of joint variable selection and model optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。