针对多光谱地物分类中的光谱偏移问题,提出高效微调新方法。
Generalizable Multispectral Land Cover Classification via Frequency-Aware Mixture of Low-Rank Token Experts
- 引入频域感知的低秩专家混合模块,动态调整特征
- 跨传感器与跨地理场景下性能显著超越现有方法
- 适合遥感图像泛化任务,尤其适用于参数受限场景
本文提出Land-MoE,一种用于多光谱地物分类(MLCC)的新方法。由于传感器差异和地理条件不同导致的光谱偏移是该领域的主要挑战。现有方法多依赖小规模模型进行域适应与泛化,性能有限。Land-MoE通过分层插入频域感知的低秩专家混合模块(MoLTE)和频域滤波器(FAF),以参数高效方式微调视觉基础模型(VFMs)。MoLTE利用不同秩的令牌生成多样化特征调整,增强对光谱偏移的鲁棒性;FAF在频域对优化特征进行调制,有效捕捉与语义强相关的频段信息,同时抑制无关噪声。在跨传感器与跨地理场景的MLCC任务上,实验表明其性能大幅优于现有方法,并在RGB遥感图像的域泛化语义分割任务中达到最新水平。
原文摘要 · Abstract (English)
We introduce Land-MoE, a novel approach for multispectral land cover classification (MLCC). Spectral shift, which emerges from disparities in sensors and geospatial conditions, poses a significant challenge in this domain. Existing methods predominantly rely on domain adaptation and generalization strategies, often utilizing small-scale models that exhibit limited performance. In contrast, Land-MoE addresses these issues by hierarchically inserting a Frequency-aware Mixture of Low-rank Token Experts, to fine-tune Vision Foundation Models (VFMs) in a parameter-efficient manner. Specifically, Land-MoE comprises two key modules: the mixture of low-rank token experts (MoLTE) and frequency-aware filters (FAF). MoLTE leverages rank-differentiated tokens to generate diverse feature adjustments for individual instances within multispectral images. By dynamically combining learnable low-rank token experts of varying ranks, it enhances the robustness against spectral shifts. Meanwhile, FAF conducts frequency-domain modulation on the refined features. This process enables the model to effectively capture frequency band information that is strongly correlated with semantic essence, while simultaneously suppressing frequency noise irrelevant to the task. Comprehensive experiments on MLCC tasks involving cross-sensor and cross-geospatial setups demonstrate that Land-MoE outperforms existing methods by a large margin. Additionally, the proposed approach has also achieved state-of-the-art performance in domain generalization semantic segmentation tasks of RGB remote sensing images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。