用轻量化结构提升气候数据建模效率与可解释性
FAConvLSTM: Factorized-Attention ConvLSTM for Efficient Feature Extraction in Multivariate Climate Data
- 分解门控计算,引入稀疏注意力与多尺度分支
- 比标准ConvLSTM降低73%计算量,保持高精度
- 适合需要高效建模长程气候关联的研究者
从高分辨率多变量地球观测数据中学习具有物理意义的时空表征极具挑战,受限于强局部动力、长程遥相关、多尺度交互及非平稳性。尽管2D ConvLSTM是常用基线,但其密集卷积门控带来高计算开销,且严格局部感受野限制了长程空间结构和解耦气候动态的建模能力。为此,我们提出FAConvLSTM,一种可直接替换ConvLSTM2D的因子化注意力卷积循环神经网络层,同时提升效率、空间表达能力和物理可解释性。该方法通过轻量级[1×1]瓶颈和共享深度可分离空间混合分解递归门计算,显著降低通道复杂度,同时保留递归动态。多尺度空洞深度分支与挤压-激励重校准机制实现跨尺度物理过程的高效建模,窥视孔连接增强时间精度。为在不增加全局注意力成本下捕捉遥相关尺度依赖,引入轻量级轴向空间注意力机制,并稀疏应用于时间维度。专用子空间头进一步生成每时刻紧凑嵌入,通过固定季节位置编码的时序自注意力进行优化。在多变量时空气候数据上的实验表明,相比标准ConvLSTM,FAConvLSTM能生成更稳定、可解释、鲁棒的潜在表征,同时显著降低计算开销。
原文摘要 · Abstract (English)
Learning physically meaningful spatiotemporal representations from high-resolution multivariate Earth observation data is challenging due to strong local dynamics, long-range teleconnections, multi-scale interactions, and nonstationarity. While ConvLSTM2D is a commonly used baseline, its dense convolutional gating incurs high computational cost and its strictly local receptive fields limit the modeling of long-range spatial structure and disentangled climate dynamics. To address these limitations, we propose FAConvLSTM, a Factorized-Attention ConvLSTM layer designed as a drop-in replacement for ConvLSTM2D that simultaneously improves efficiency, spatial expressiveness, and physical interpretability. FAConvLSTM factorizes recurrent gate computations using lightweight [1 times 1] bottlenecks and shared depthwise spatial mixing, substantially reducing channel complexity while preserving recurrent dynamics. Multi-scale dilated depthwise branches and squeeze-and-excitation recalibration enable efficient modeling of interacting physical processes across spatial scales, while peephole connections enhance temporal precision. To capture teleconnection-scale dependencies without incurring global attention cost, FAConvLSTM incorporates a lightweight axial spatial attention mechanism applied sparsely in time. A dedicated subspace head further produces compact per timestep embeddings refined through temporal self-attention with fixed seasonal positional encoding. Experiments on multivariate spatiotemporal climate data shows superiority demonstrating that FAConvLSTM yields more stable, interpretable, and robust latent representations than standard ConvLSTM, while significantly reducing computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。