比较三种融合策略,提升多模态遥感图像对非洲草原生态特征的识别能力。
Multimodal Fusion Strategies for Mapping Biophysical Landscape Features
- 对比早期融合、晚期融合与专家混合模型的多模态数据整合方式
- 晚期融合整体表现最佳,AUC达0.698;不同方法对不同地物识别效果差异显著
- 适合生态监测、遥感分类研究者参考,尤其关注特定生物物理特征识别
多模态航空数据被用于监测自然系统,机器学习可显著加速此类影像中景观特征的分类,助力生态保护。然而,如何在深度学习模型中融合多种模态仍待深入探索。为此,我们基于三模态空间对齐正射影像数据集(热成像、RGB、LiDAR),研究了三种融合策略:早期融合、晚期融合与专家混合模型。目标是识别非洲稀树草原生态系统中三种生态相关的生物物理特征:犀牛粪堆、白蚁丘和水体。三种策略在模态融合时机及权重生成方式上存在差异:早期融合在所有阶段合并数据,晚期融合在决策阶段融合,而专家混合模型则根据输入动态生成每类的模态权重。总体上,三种方法宏观平均性能相近,晚期融合达到0.698的AUC;但各类别表现差异明显:早期融合对粪堆和水体召回率最高,专家混合模型对白蚁丘召回率最优。
原文摘要 · Abstract (English)
Multimodal aerial data are used to monitor natural systems, and machine learning can significantly accelerate the classification of landscape features within such imagery to benefit ecology and conservation. It remains under-explored, however, how these multiple modalities ought to be fused in a deep learning model. As a step towards filling this gap, we study three strategies (Early fusion, Late fusion, and Mixture of Experts) for fusing thermal, RGB, and LiDAR imagery using a dataset of spatially-aligned orthomosaics in these three modalities. In particular, we aim to map three ecologically-relevant biophysical landscape features in African savanna ecosystems: rhino middens, termite mounds, and water. The three fusion strategies differ in whether the modalities are fused early or late, and if late, whether the model learns fixed weights per modality for each class or generates weights for each class adaptively, based on the input. Overall, the three methods have similar macro-averaged performance with Late fusion achieving an AUC of 0.698, but their per-class performance varies strongly, with Early fusion achieving the best recall for middens and water and Mixture of Experts achieving the best recall for mounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。