用双流协同特征融合提升矿区场景分类准确率
Dual-Stream Global-Local Feature Collaborative Representation Network for Scene Classification of Mining Area
- 分全球与局部双分支提取特征,协同优化语义表示
- 模型整体准确率达83.63%,优于现有方法
- 适合需要精细识别矿区地貌的遥感应用
矿区场景分类为地质环境监测与资源开发规划提供精准基础数据。本研究融合多源数据构建多模态矿区地表覆盖场景分类数据集。矿区分类面临空间布局复杂、多尺度特征显著的挑战。通过提取全局与局部特征,可全面反映空间分布,更准确捕捉矿区整体特性。提出一种双分支融合模型,利用协同表示将全局特征分解为一组关键语义向量。模型包含三个核心组件:(1) 多尺度全局Transformer分支,利用大尺度特征生成小尺度特征的全局通道注意力,有效捕捉多尺度特征关系;(2) 局部增强协同表示分支,通过局部特征与重构的关键语义集优化注意力权重,确保矿区局部上下文与细节特征有效融合,提升对细粒度空间变化的敏感性;(3) 双分支深度特征融合模块,融合两分支互补特征,引入更多场景信息,强化模型对复杂矿区景观的区分与分类能力。最后采用多损失计算,平衡各模块集成效果。模型总体准确率为83.63%,优于其他对比模型,并在所有其他评估指标上表现最佳。
原文摘要 · Abstract (English)
Scene classification of mining areas provides accurate foundational data for geological environment monitoring and resource development planning. This study fuses multi-source data to construct a multi-modal mine land cover scene classification dataset. A significant challenge in mining area classification lies in the complex spatial layout and multi-scale characteristics. By extracting global and local features, it becomes possible to comprehensively reflect the spatial distribution, thereby enabling a more accurate capture of the holistic characteristics of mining scenes. We propose a dual-branch fusion model utilizing collaborative representation to decompose global features into a set of key semantic vectors. This model comprises three key components:(1) Multi-scale Global Transformer Branch: It leverages adjacent large-scale features to generate global channel attention features for small-scale features, effectively capturing the multi-scale feature relationships. (2) Local Enhancement Collaborative Representation Branch: It refines the attention weights by leveraging local features and reconstructed key semantic sets, ensuring that the local context and detailed characteristics of the mining area are effectively integrated. This enhances the model's sensitivity to fine-grained spatial variations. (3) Dual-Branch Deep Feature Fusion Module: It fuses the complementary features of the two branches to incorporate more scene information. This fusion strengthens the model's ability to distinguish and classify complex mining landscapes. Finally, this study employs multi-loss computation to ensure a balanced integration of the modules. The overall accuracy of this model is 83.63%, which outperforms other comparative models. Additionally, it achieves the best performance across all other evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。