arXiv:2609.05351cs.CV2026-09

用不到600万参数实现多源遥感图像高效建模,适合资源有限场景。

MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation

论文配图:MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation
图 1 · 摘自论文原文
  • 采用专家混合结构与传感器适配器,分治处理不同遥感数据
  • 在64像素任务中达64.42%交并比,224像素上准确率90.56%
  • 仅300多万参数却支持多模态融合与跨传感器对齐,适合部署

近期地球观测表征学习常依赖更大模型处理异构传感器和缺失数据。我们提出MEOX(多模态地球观测专家),一个参数量为293.9万的编码器和总参数311.5万的多模态掩码自编码器。通过传感器特定适配器、显式有效性信号和共享稀疏专家模块,在局部块级融合前保留模态特异性处理。四个元数据标记随后伴随单一空间序列通过十四层编码器。共享专家投影结合私有低秩残差限制参数增长,旋转注意力支持预训练外的空间网格。模型在122.8万MMEarth64样本上进行模态平衡掩码重建与结构化传感器丢弃预训练。冻结迁移在六个GEO-Bench任务上评估,分别在64和224像素下测试。64像素下可可豆分割达到64.42%平均交并比,224像素下EuroSAT分类平均准确率达90.56%,优于现有CSMoE结果。BigEarthNet微调达72.95%微平均精确率。路由诊断揭示专家参与度、空间依赖性、模态关联与功能贡献。独立世界覆盖探测显示元数据带来0.64个百分点提升,检索任务分离同传感器语义与跨传感器对齐。结果表明:小规模参数下仍具强泛化能力与任务迁移性能。

原文摘要 · Abstract (English)

Recent advances in Earth Observation representation learning accommodate heterogeneous sensors and missing observations, often through larger architectures. We present MEOX (Multimodal Earth Observation with eXperts), a multimodal masked autoencoder with a 2.939 million-parameter encoder and 3.115 million parameters in total. Sensor-specific adapters, explicit validity signals, and a shared sparse-expert block preserve modality-dependent processing before a learned patch-wise fusion. Four metadata tokens then accompany a single spatial sequence through fourteen further encoder blocks. Shared expert projections with private low-rank residuals constrain parameter growth, while rotary attention supports downstream spatial grids different from pretraining. The model is pretrained on 1.228 million MMEarth64 samples using modality-balanced masked reconstruction and structured sensor dropout. Frozen transfer is evaluated on six GEO-Bench tasks at both 64 and 224 pixels. The model reaches 64.42% mean intersection-over-union on cashew segmentation at 64 pixels and 90.56% average accuracy on EuroSAT at 224 pixels, exceeding the corresponding reported CSMoE results. BigEarthNet finetuning reaches 72.95% micro-average precision. Routing diagnostics distinguish expert participation, spatial dependence, modality association, and functional contribution. A held-out WorldCover probe measures a 0.64-percentage-point benefit from metadata, while retrieval separates same-sensor semantics from cross-sensor alignment. These results demonstrate sensor-flexible representation learning and strong task transfer using a compact parameter budget.

遥感建模专家混合多模态轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。