arXiv:2509.25393cs.CVcs.AI2025-09被引 3

融合动态与静态数据,用新模型大幅提升高分辨率地面沉降预测精度

Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction

  • 设计多模态时空注意力机制,统一处理位移与物理先验数据
  • 在EGMS数据集上,长程预测均方根误差降低一个数量级
  • 适合需要高精度地质灾害预警的研究者和工程应用

高分辨率地面沉降预测因动态复杂且非线性而极具挑战。传统方法如ConvLSTM难以捕捉长程依赖,而现有工作更根本的局限在于单一模态数据范式。为此,我们提出多模态时空变换器(MM-STT),融合动态形变数据与静态物理先验。其核心创新是联合时空注意力机制,统一处理多模态特征。在公开的EGMS数据集上,MM-STT达到新基准,相较所有基线(包括STGCN、STAEformer等先进方法)的长程预测均方根误差降低一个数量级。结果表明,此类问题中,模型深层多模态融合能力是实现突破性性能的关键。

原文摘要 · Abstract (English)

Forecasting high-resolution land subsidence is a critical yet challenging task due to its complex, non-linear dynamics. While standard architectures like ConvLSTM often fail to model long-range dependencies, we argue that a more fundamental limitation of prior work lies in the uni-modal data paradigm. To address this, we propose the Multi-Modal Spatio-Temporal Transformer (MM-STT), a novel framework that fuses dynamic displacement data with static physical priors. Its core innovation is a joint spatio-temporal attention mechanism that processes all multi-modal features in a unified manner. On the public EGMS dataset, MM-STT establishes a new state-of-the-art, reducing the long-range forecast RMSE by an order of magnitude compared to all baselines, including SOTA methods like STGCN and STAEformer. Our results demonstrate that for this class of problems, an architecture's inherent capacity for deep multi-modal fusion is paramount for achieving transformative performance.

地面沉降多模态融合时空建模深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。