将地理统计学先验注入注意力机制,提升时空预测精度与不确定性校准。
Spatially-informed transformers: Injecting geostatistical covariance biases into self-attention for spatio-temporal forecasting
- 在自注意力中引入可学习协方差核,嵌入空间距离先验。
- 端到端学习真实空间衰减参数,实现深度变差分析。
- 适合需要精准概率预测的交通、气象等时空建模任务。
高维时空过程建模面临经典地统计学的概率严谨性与深度学习灵活表示能力之间的根本矛盾。高斯过程虽具理论一致性与精确不确定性量化,但计算开销过大,难以应用于大规模传感器网络。而现代变换器擅长序列建模,却缺乏几何归纳偏置,将空间传感器视为无序标记,无法理解距离。本文提出空间感知变换器,通过可学习协方差核直接将地统计学归纳偏置注入自注意力机制。通过将注意力结构形式分解为平稳物理先验与非平稳数据驱动残差,施加软拓扑约束,偏好空间邻近交互,同时保留复杂动态建模能力。实验表明,网络可端到端恢复真实空间衰减参数,实现“深度变差分析”。在合成高斯随机场与真实交通基准上的广泛实验验证了其优于现有图神经网络的表现。严格统计验证显示,该方法不仅预测更准确,且概率预测校准良好,有效弥合了物理感知建模与数据驱动学习的鸿沟。
原文摘要 · Abstract (English)
The modeling of high-dimensional spatio-temporal processes presents a fundamental dichotomy between the probabilistic rigor of classical geostatistics and the flexible, high-capacity representations of deep learning. While Gaussian processes offer theoretical consistency and exact uncertainty quantification, their prohibitive computational scaling renders them impractical for massive sensor networks. Conversely, modern transformer architectures excel at sequence modeling but inherently lack a geometric inductive bias, treating spatial sensors as permutation-invariant tokens without a native understanding of distance. In this work, we propose a spatially-informed transformer, a hybrid architecture that injects a geostatistical inductive bias directly into the self-attention mechanism via a learnable covariance kernel. By formally decomposing the attention structure into a stationary physical prior and a non-stationary data-driven residual, we impose a soft topological constraint that favors spatially proximal interactions while retaining the capacity to model complex dynamics. We demonstrate the phenomenon of ``Deep Variography'', where the network successfully recovers the true spatial decay parameters of the underlying process end-to-end via backpropagation. Extensive experiments on synthetic Gaussian random fields and real-world traffic benchmarks confirm that our method outperforms state-of-the-art graph neural networks. Furthermore, rigorous statistical validation confirms that the proposed method delivers not only superior predictive accuracy but also well-calibrated probabilistic forecasts, effectively bridging the gap between physics-aware modeling and data-driven learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。