arXiv:2607.16338cs.CVcs.AI2026-07

双主干多尺度融合网络提升城市遥感图像分类精度

DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification

论文配图:DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification
图 1 · 摘自论文原文
  • 用两个预训练主干提取多层次特征,通过残差传播增强跨尺度交互
  • 在AID数据集上达到97.46%平均准确率,优于多数现有方法
  • 适合需要高精度遥感场景分类的研究者和应用开发者

本文提出DMFNet,一种用于遥感场景分类的双主干多尺度特征融合框架,包含残差特征传播与空间注意力机制。现有方法在捕捉多尺度特征交互及从高类内差异、类间相似的复杂空域场景中学习鲁棒表征方面面临挑战。为此,该框架采用两个预训练主干网络提取多样化的层次化特征表示,并引入具有残差特征传播的多尺度特征融合机制,以增强多分辨率层级间的特征交互。此外,设计空间注意力模块以突出多目标场景中有信息量的空间区域。还采用两阶段训练策略:先冻结主干,再选择性微调,确保优化稳定并提升泛化能力。在基准AID数据集上的实验表明,DMFNet实现了97.46% ± 0.14%的平均准确率。消融分析进一步验证了各组件协同作用的重要性。

原文摘要 · Abstract (English)

This article presents DMFNet, a dual-backbone multiscale feature fusion framework with residual feature propagation and spatial attention for remote sensing scene classification. Existing approaches often face challenges in effectively capturing multiscale feature interactions and learning robust feature representations from complex aerial scenes with high intra-class variability and inter-class similarity. To address these limitations, the proposed framework employs two pretrained backbone networks to extract diverse hierarchical feature representations. A multiscale feature fusion mechanism with residual feature propagation is introduced to enhance feature interaction across multiple resolution levels. In addition, a spatial attention module is introduced to emphasize informative spatial regions in multi-object scenes. Further, a two-stage training strategy consisting of backbone freezing followed by selective fine-tuning is adopted to ensure stable optimization and improved generalization. Experiments conducted on the benchmark AID dataset demonstrate that the DMFNet achieves an average accuracy of 97.46\% $\pm$ 0.14\%. Ablative analysis further show the importance of various components in unison.

遥感分类多尺度融合空间注意力双主干

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。