提出新型注意力网络,提升航拍图像分类准确率
Revisiting Aerial Scene Classification on the AID Benchmark
- 设计融合多尺度特征的注意力增强网络
- 在AID数据集上达到91.72%分类准确率
- 适合遥感图像分析与深度学习研究者参考
航拍图像在城市规划与环境保护中至关重要,其包含建筑、森林、山地及空地等多种结构,具有高度异质性,因而构建稳健的场景分类模型仍具挑战。本文系统回顾了多种机器学习方法在航拍图像分类中的应用,涵盖手工特征(如SIFT、LBP)、传统CNN(如VGG、GoogLeNet)以及先进深度混合网络。在此基础上,提出Aerial-Y-Net,一种结合空间注意力机制与多尺度特征融合的CNN模型,可更好地捕捉航拍图像复杂特性。在AID数据集上的实验表明,该模型达到91.72%的准确率,优于多个基准架构。
原文摘要 · Abstract (English)
Aerial images play a vital role in urban planning and environmental preservation, as they consist of various structures, representing different types of buildings, forests, mountains, and unoccupied lands. Due to its heterogeneous nature, developing robust models for scene classification remains a challenge. In this study, we conduct a literature review of various machine learning methods for aerial image classification. Our survey covers a range of approaches from handcrafted features (e.g., SIFT, LBP) to traditional CNNs (e.g., VGG, GoogLeNet), and advanced deep hybrid networks. In this connection, we have also designed Aerial-Y-Net, a spatial attention-enhanced CNN with multi-scale feature fusion mechanism, which acts as an attention-based model and helps us to better understand the complexities of aerial images. Evaluated on the AID dataset, our model achieves 91.72% accuracy, outperforming several baseline architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。