arXiv:2509.08003cs.CV2025-09被引 7

用多模态注意力与频域感知网络提升城市洪涝预测准确率

An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities

  • 融合视觉与文本特征的动态跨模态门控注意力机制
  • 频域增强的通道与空间注意力使模型在三数据集上达93.33%以上F1分数
  • 适合关注智能防灾、可解释性模型的科研与工程人员

在气候变化加剧的背景下,城市洪涝已成为可持续城市发展的重要挑战,威胁生命、基础设施与生态系统。传统检测方法依赖单一模态数据和静态规则系统,难以捕捉洪涝事件中固有的动态非线性关系。现有注意力机制与集成学习方法在层级细化、跨模态特征融合及噪声环境适应性方面存在局限,导致分类性能不佳。为此,本文提出XFloodNet框架,通过三项创新组件实现城市洪涝分类的重构:(1)分层跨模态门控注意力机制,动态对齐视觉与文本特征,实现精准多粒度交互并消除上下文歧义;(2)异构卷积自适应多尺度注意力模块,利用频域增强的通道注意力与频调制的空间注意力,在谱域与空域中提取并优先筛选关键洪涝特征;(3)级联卷积变换器特征精炼技术,通过自适应缩放与级联操作融合层级特征,保障鲁棒且抗噪的洪涝检测能力。在Chennai Floods、Rhine18 Floods与Harz17 Floods三个基准数据集上,XFloodNet分别取得93.33%、82.24%与88.60%的F1分数,显著超越现有方法。

原文摘要 · Abstract (English)

In an era of escalating climate change, urban flooding has emerged as a critical challenge for sustainable cities, threatening lives, infrastructure, and ecosystems. Traditional flood detection methods are constrained by their reliance on unimodal data and static rule-based systems, which fail to capture the dynamic, non-linear relationships inherent in flood events. Furthermore, existing attention mechanisms and ensemble learning approaches exhibit limitations in hierarchical refinement, cross-modal feature integration, and adaptability to noisy or unstructured environments, resulting in suboptimal flood classification performance. To address these challenges, we present XFloodNet, a novel framework that redefines urban flood classification through advanced deep-learning techniques. XFloodNet integrates three novel components: (1) a Hierarchical Cross-Modal Gated Attention mechanism that dynamically aligns visual and textual features, enabling precise multi-granularity interactions and resolving contextual ambiguities; (2) a Heterogeneous Convolutional Adaptive Multi-Scale Attention module, which leverages frequency-enhanced channel attention and frequency-modulated spatial attention to extract and prioritize discriminative flood-related features across spectral and spatial domains; and (3) a Cascading Convolutional Transformer Feature Refinement technique that harmonizes hierarchical features through adaptive scaling and cascading operations, ensuring robust and noise-resistant flood detection. We evaluate our proposed method on three benchmark datasets, such as Chennai Floods, Rhine18 Floods, and Harz17 Floods, XFloodNet achieves state-of-the-art F1-scores of 93.33%, 82.24%, and 88.60%, respectively, surpassing existing methods by significant margins.

洪水预测多模态可解释性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。